Prosecution Insights
Last updated: August 17, 2026
Application No. 18/510,199

METHOD AND SERVER FOR SEARCHING FOR OPTIMAL NEURAL NETWORK ARCHITECTURE BASED ON CHANNEL CONCATENATION

Non-Final OA §101§103
Filed
Nov 15, 2023
Priority
Dec 20, 2022 — RE 10-2022-0179159
Examiner
TRAN, UYEN-NHU PHAM
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Electronics and Telecommunications Research Institute
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-55.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
7 currently pending
Career history
7
Total Applications
across all art units

Statute-Specific Performance

§101
28.0%
-12.0% vs TC avg
§103
56.0%
+16.0% vs TC avg
§112
16.0%
-24.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This office action is in response to submission of application on 11/15/2023. Claims 1-19 are presented for examination. Claim Objections Claims 1, 6, 7, 12, 13, and 17 are objected to because of the following informalities: Claims 1, 12, and 13 recite the limitation “and additionally extending an output feature map that is results of…” There appears to be a typo, “that results of”. Appropriate corrections are required. Claims 6, 7, and 17 recites the limitation “the input feature map” There is insufficient antecedent basis for this limitation in the claims. Claim 1 and 13 only introduces “input feature map candidate group” and is unclear if it’s the same as “the input feature map”. Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Is the claim to a process, machine, manufacture, or composition of matter? Claims 1-11 are directed to a method; claim 12 is directed to a method; and claims 13-19 are directed to a machine; therefore, all claims are directed to one of the four statutory categories. Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon? Claim 1 recites limitations of: adjusting spatial size information of an input feature map candidate group so that the spatial size information of the input feature map candidate group corresponds to spatial size information of an output feature map; - mathematical concept (relationships, formulas or equations, calculations) of resizing information of an input feature map performing a channel-based concatenation operation on the input feature map candidate group; - mathematical concept (relationships, formulas or equations, calculations) of a concatenation operation Claim 12 recites limitations of: setting spatial size information of an output feature map; - mathematical concept (relationships, formulas or equations, calculations) of defining a parameter adjusting spatial size information of an input feature map candidate group so that the spatial size information of the input feature map candidate group corresponds to spatial size information of an output feature map; - mathematical concept (relationships, formulas or equations, calculations) of resizing information of an input feature map performing a channel-based concatenation operation on the input feature map candidate group; - mathematical concept (relationships, formulas or equations, calculations) of resizing information of an input feature map Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application? Claim 1 recites additional elements of: and additionally extending an output feature map that is results of the channel-concatenated operation to the input feature map candidate group. – extending an output feature map merely amounts to recalculate and readjust which is insignificant extra-solution activity. See MPEP § 2106.05(g), item (3), which identifies necessary recalculating and readjusting as an example of extra-solution activity. Claim 12 recites additional elements of: additionally extending an output feature map that is results of the channel-concatenated operation to the input feature map candidate group; - extending an output feature map merely amounts to recalculate and readjust which is insignificant extra-solution activity. See MPEP § 2106.05(g), item (3), which identifies necessary recalculating and readjusting as an example of extra-solution activity. and performing learning for searching for a one-shot neural network on a super-net comprising each node corresponding to the additionally extended input feature map candidate group. - this element constitutes “mere instructions to apply an exception.” (MPEP § 2106.05(f)). The additional elements do not integrate the abstract idea into a practical application. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? The additional elements are: Claim 1 recites additional elements of: and additionally extending an output feature map that is results of the channel-concatenated operation to the input feature map candidate group. – extending an output feature map merely amounts to recalculate and readjust which is insignificant extra-solution activity. See MPEP § 2106.05(g). Recalculating and Readjusting is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(ii). Claim 12 recites additional elements of: additionally extending an output feature map that is results of the channel-concatenated operation to the input feature map candidate group; - extending an output feature map merely amounts to recalculate and readjust which is insignificant extra-solution activity. See MPEP § 2106.05(g). Recalculating and Readjusting is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(ii). and performing learning for searching for a one-shot neural network on a super-net comprising each node corresponding to the additionally extended input feature map candidate group. - this element constitutes “mere instructions to apply an exception.” (MPEP § 2106.05(f)). Independent claim 13 recites the same relevant limitations as claim 12 and similar analysis applies. Claim 13 recites the additional elements of, A server for searching for optimal neural network architecture based on channel concatenation, the server comprising: memory in which a program for searching for optimal neural network architecture based on channel concatenation is stored; and a processor configured to - components recited at a high level are construed as generic computer components used to implement the abstract idea. See MPEP 2106.05(f)(2). It does not integrate the abstract idea into a practical application. Nor do they amount to significantly more. Therefore, the independent claim is not patent eligible. The above analysis similarly applies to the dependent claims. Dependent claim 2 recites, setting the spatial size information of the output feature map; - mathematical concept (relationships, formulas or equations, calculations) of defining a parameter and determining an input feature map to be included in the input feature map candidate group based on the spatial size information of the output feature map. - mental process (observation, evaluation, judgement) as a human mind is able to determine an input feature map Dependent claim 3 recite, setting the spatial size information of the output feature map so that the spatial size information of the output feature map corresponds to spatial size information of an output feature map in a corresponding node according to a path aggregation network (PAN) path sequence. - mathematical concept (relationships, formulas or equations, calculations) of using PAN sequence Dependent claim 4 and 15 recites, when a first node according to the PAN path sequence is a node connected to a backbone, - getting the first node merely amounts to storing and retrieving data which is insignificant extra-solution activity. See MPEP § 2106.05(g). Recalculating and Readjusting is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(iii). setting an input feature map having identical resolution from the backbone as an essential input feature map; - mathematical concept (relationships, formulas or equations, calculations) of defining a parameter and setting an input feature map having different resolution from the backbone as a candidate input feature map, wherein when an output feature map from another node is present, an output feature map from the another node is set as the candidate input feature map. - mathematical concept (relationships, formulas or equations, calculations) of defining a parameter Dependent claim 5 and 16 recites, when a second node according to the PAN path sequence is a node connected to a first node not a backbone, getting the second node merely amounts to storing and retrieving data which is insignificant extra-solution activity. See MPEP § 2106.05(g). Recalculating and Readjusting is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(iii). setting an input feature map having identical resolution from the first node as an essential input feature map; - mathematical concept (relationships, formulas or equations, calculations) of defining a parameter and setting an input feature map from the backbone as a candidate input feature map, wherein when an output feature map from another node except the first node is present, an output feature map from the another node is set as the candidate input feature map. - mathematical concept (relationships, formulas or equations, calculations) of defining a parameter Dependent claim 6 recites, adjusting the spatial size information of the input feature map to be identical with the spatial size information of the output feature map by applying a convolution product operation having a stride size of 2 or more and a 1×1 convolution product operation when spatial size information of an input feature map included in the input feature map candidate group is greater than the spatial size information of the output feature map. – adjusting the spatial size information merely amounts to recalculate and readjust which is insignificant extra-solution activity. See MPEP § 2106.05(g). Recalculating and Readjusting is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(ii). Dependent claim 7 recites, adjusting the spatial size information of the input feature map to be identical with the spatial size information of the output feature map by applying an up-sampling operation and a 1×1 convolution product operation when spatial size information of an input feature map included in the input feature map candidate group is smaller than the spatial size information of the output feature map. - adjusting the spatial size information merely amounts to recalculate and readjust which is insignificant extra-solution activity. See MPEP § 2106.05(g). Recalculating and Readjusting is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(ii). Dependent claim 8 and 18 recites, applying a structure parameter and a softmax function to each of candidate input feature maps except an essential input feature map, in the input feature map candidate group. - mathematical concept (relationships, formulas or equations, calculations) of using the softmax function Dependent claim 9 recites, performing learning for searching for a one-shot neural network on a super-net comprising each node corresponding to the additionally extended input feature map candidate group. - mathematical concept (relationships, formulas or equations, calculations) performing a learning for searching is optimization Dependent claim 10 and 19 recites, removing an input from a node having a structure parameter that satisfies a predetermined condition in the super-net on which the learning has been completed. – removing an input merely amounts to recalculate and readjust which is insignificant extra-solution activity. See MPEP § 2106.05(g). Recalculating and Readjusting is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(ii). Dependent claim 11 recites, further comprising performing a channel scaling operation on the output feature map that is the results of the channel-concatenated operation. - mathematical concept (relationships, formulas or equations, calculations) of using a scaling operation Dependent claim 14 recites, sets the spatial size information of the output feature map so that the spatial size information of the output feature map corresponds to spatial size information of an output feature map in a corresponding node according to a path aggregation network (PAN) path sequence, - mathematical concept (relationships, formulas or equations, calculations) of using the PAN sequence and determines an input feature map to be included in the input feature map candidate group based on the spatial size information of the output feature map. - mental process (observation, evaluation, judgement) as a human mind is able to determine an input feature map Dependent claim 17 recites, wherein the processor adjusts the spatial size information of the input feature map to be identical with the spatial size information of the output feature map, by applying a convolution product operation having a stride size of 2 or more and a 1×1 convolution product operation when spatial size information of an input feature map included in the input feature map candidate group is greater than the spatial size information of the output feature map and applying an up-sampling operation and a 1×1 convolution product operation when the spatial size information of the input feature map included in the input feature map candidate group is smaller than the spatial size information of the output feature map. - mathematical concept (relationships, formulas or equations, calculations) of using up-sampling operation, convolution product operation The dependent claims do not integrate the abstract idea into a practical application, nor do they amount to significantly more than the abstract idea. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The following are the references used: Ghaisi et al. (NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection, herein Ghaisi) Wang et al. (EAutoDet: Efficient Architecture Search for Object Detection, herein Wang) Liang et al. (OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection, herein Liang) Liu et al. (DARTS: Differentiable Architecture Search, herein Liu) Claim(s) 1, 2, 8, and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ghaisi and Wang. Regarding claim 1, Ghaisi teaches, adjusting spatial size information of an input feature map candidate group so that the spatial size information of the input feature map candidate group corresponds to spatial size information of an output feature map; (Ghaisi, page 4, section 3.1, “The input feature layers are adjusted to the output resolution by nearest neighbor upsampling or max pooling if needed before applying the binary operation.” note: the input feature layers are the pool of the inputs which maps to the candidate group. The adjusted to the output resolution means their size is changed to equal the output’s size which maps to the spatial size made to correspond to the output. By nearest neighbor upsampling or max pooling is the mechanism that does the resizing.) and additionally extending an output feature map that is results of the channel-concatenated operation to the input feature map candidate group. (Ghaisi, page 4, section 3.1, “the newly-generated feature layer is appended to the list of existing input candidates and becomes a new candidate for the next merging cell.” note: the newly generated feature layer maps to the node’s output which maps to the output feature map. Appended to the list of existing input candidates is the act of adding it back into the pool which maps to extended to the candidate group. Becomes a new candidate for the next merging cell confirms the added output is then available to later nodes.) Ghaisi does not teach, performing a channel-based concatenation operation on the input feature map candidate group; Wang teaches, performing a channel-based concatenation operation on the input feature map candidate group; (Wang, page 5, section 3.3, “Nodes indicate feature maps and connect with all their predecessors in the supernet. After the search stage, only two predecessors will be selected for each node” and page 4, section 3.2, “suppose two features with expansion rates and base channels (e1, C1) and (e2,C2) are concatenated. A convolution is applied after the concatenation layer” note: the candidate group is a node’s predecessor, since nodes indicate feature maps and connect with all their predecessors in the supernet. The channel-based concatenation is the combining operation applies to those inputs, since the combination blocks are concatenated and it states that two candidate features are concatenated before a convolution is applied. This happens in the supernet and the predecessors are searched and then narrowed which maps to inside a search) It would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Ghaisi and Wang because Ghaisi searches for the best way to combine feature maps, resizes them all to a common size, and feed each result back into the candidate pool, but it combines them by adding rather than concatenating. Wang uses concatenation. Concatenation keeps each input’s channels intact instead of merging them away, letting the network learn how much to rely on each one. Regarding claim 2, The combination of Ghaisi and Wang teaches, setting the spatial size information of the output feature map; (Ghaisi, page 4, section 3.1, “Step 3. Select the output feature resolution.” note: the output feature resolution is the output’s spatial size, and select is the act of setting it) and determining an input feature map to be included in the input feature map candidate group based on the spatial size information of the output feature map. (Ghaisi, page 4, section 3.1, “Step 1. Select a feature layer hi from candidates. Step 2. Select another feature layer hj from candidates without replacement… Step 4. Select a binary op to combine hi and hj selected in Step 1 and Step 2 and generate a feature layer with the resolution selected in Step 3” note: Steps 1 and 2 perform the choosing aspect, they select a feature layer from candidates which means they pick which inputs are drawn from the candidate pool. The based on the output size maps because those chosen inputs are then combined in step 4 with the resolution selected in step 3, so the inputs are selected for the specific output resolution being built, which ties the input choice to the output size) Regarding claim 8, The combination of Ghaisi and Wang teaches, The method of claim 1, wherein the performing of the channel-based concatenation operation on the input feature map candidate group comprises applying a structure parameter and a softmax function to each of candidate input feature maps except an essential input feature map, in the input feature map candidate group. (Wang, page 6, section 3.3, “For each fusion block, we introduce architecture parameters αe and αo to denote the importance of edges and operations. Suppose ˜α=softmax(α) is the normalized weight. The fused feature zj = PNG media_image1.png 26 197 media_image1.png Greyscale where O is the candidate operation set, xi is the features of predecessors” note: Its architecture parameters a are the structure parameters, they denote the importance of edges and an edge is a candidate input feeding a node (x_i, the features of predecessors). The softmax function is applied to those parameters.) Regarding claim 11, The combination of Ghaisi and Wang teaches, The method of claim 1, further comprising performing a channel scaling operation on the output feature map that is the results of the channel-concatenated operation. (Wang, page 4, section 3.2, “we introduce the dynamic channel refinement technique based on Gumbel reparameterization technique (Gumbel 1954) to search for optimal expansion rates for layers… After sampling an expansion rate el for layer l with base channel number Cl, the output channel becomes elCl. The operation weights on layer l can be refined by preserving the first elCl filters.” note: The claim talks about adding a step that scaled the number of channels on the output feature map. Wang does this through dynamic channel refinement. It applies an expansion rate el to a layers base channel number Cl so that the output channel becomes elCL. Multiplying the base channel count by a rate to set a new channel count is channel scaling, the output’s channel dimension is enlarged or reduced by the factor el. It performs its channel refinement on the fusion-block outputs which is the concatenated feature maps.) Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ghaisi and Wang in view of Liu. Regarding claim 7, Ghaisi teaches, The method of claim 1, wherein the adjusting of the spatial size information of the input feature map candidate group so that the spatial size information of the input feature map candidate group corresponds to spatial size information of an output feature map comprises adjusting the spatial size information of the input feature map to be identical with the spatial size information of the output feature map by applying an up-sampling operation (Ghaisi, page 4, section 3.1, “The input feature layers are adjusted to the output resolution by nearest neighbor upsampling or max pooling if needed before applying the binary operation.”) Ghaisi does not teach, and a 1×1 convolution product Liu teaches, and a 1×1 convolution product (Liu, page 5, 3.1.1, “1×1 convolutions are inserted as necessary.”) It would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Ghaisi, Wang, and Liu because Ghaisi explains using upsampling to enlarge it when the input feature map is smaller than the output. But it does not use 1x1 convolution. Pairing the two is beneficial because upsampling only fixes the height and width of the feature map, leaving the channel count unchanged, and a 1x1 convolution is a way to reconcile channel depth so the resized input can be combined with others. Claim(s) 3-5, 9, 12-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ghaisi and Wang in view of Liang. Regarding claim 3, Liang teaches, The method of claim 2, wherein the setting of the spatial size information of the output feature map comprises setting the spatial size information of the output feature map so that the spatial size information of the output feature map corresponds to spatial size information of an output feature map in a corresponding node according to a path aggregation network (PAN) path sequence. (Liang, page 1, abstract, “we first introduce six heterogeneous information paths to build our search space, namely top-down, bottom-up, fusing splitting, scale-equalizing, skip-connect and none… each FPN candidate is represented by a densely-connected directed acyclic graph (each node is a feature pyramid and each edge is one of the six heterogeneous information paths).” and page 4, section 3.1, “each feature map (Ft i) is iteratively built by combining input pyramid feature map of the same level (Pi) and the higher-level output feature (Ft i+1): Ft i = Wt i ⊗(U(Ft i+1) +Pi)” note: the claim’s set the output size step requires the size to be set according to PAN where each node’s output resolution is determined by where the node sits in a top-down/bottom-up aggregation path. Liang shows its network is a graph whose nodes are feature pyramids and whose edges are information paths. Those paths are the top-down/bottom-up path aggregation paths. A node’s output being build at a specific pyramid level as it advances along the top-down path. The node’s position in the PAN sequence determines the resolution of the output it produces.) It would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Ghaisi, Wang, and Liang because Ghaisi lets its search pick any two feature maps to combine in any order, which means a huge number of possibly layouts and would make hard to read results. Liang organizes the same kind of search into a set path: top-down, then bottom-up, so that each node has a fixed place in line and produces its output at whatever resolution that place calls for. This allows for a cheaper search. Regarding claim 4, The combination of Ghaisi, Wang, and Liang teaches, when a first node according to the PAN path sequence is a node connected to a backbone, (Ghaisi, page 3, section 3.1, “We follow the design by RetinaNet [23] which uses the last layer in each group of feature layers as the inputs to the first pyramid network… We use as inputs features in 5 scales {C3,C4,C5,C6,C7} with corresponding feature stride of {8,16,32,64,128} pixels.”, note: Ghaisi’s first pyramid network have inputs come directly from the backbone, and the node receives the backbone features.) setting an input feature map having identical resolution from the backbone as an essential input feature map; (Liang, page 4, section 3.1, “each feature map (Ft i) is iteratively built by combining input pyramid feature map of the same level (Pi) and the higher-level output feature (Ft i+1): Ft i = Wt i ⊗(U(Ft i+1) +Pi),” and page 5, section 3.2.1, “We assume the intermediate nodes are fully connected with former nodes… we associate an edge importance weight to each edge.” note: Pi is the input pyramid feature of the same level, the same resolution feature from the backbone pyramid) and setting an input feature map having different resolution from the backbone as a candidate input feature map, (Ghaisi, page 3, section 3.1, “The RNN controller selects any two candidate feature layers and a binary operation to com bine them into a new feature layer, where all feature layers may have different resolution.”, note: an input feature map having different resolution from the backbone maps to all feature layers may have different resolutions) wherein when an output feature map from another node is present, an output feature map from the another node is set as the candidate input feature map. (Ghaisi, page 4, section 3.1, “In Step 5, the newly-generated feature layer is appended to the list of existing input candidates and becomes a new candidate for the next merging cell.”) Regarding claim 5, when a second node according to the PAN path sequence is a node connected to a first node not a backbone, (Liang, page 4, section 3.1, “Each feature map (Fb i ) is obtained by merging the input feature maps (Pi) of the same level, and the output feature map below it (Fb i−1): Fb i = Wb i ⊗(D(Fb i−1) +Pi),” note: in the bottom-up path, node i is fed by Fb i−1, the output of the previous node, not the backbone. That is the claim’s second node… connected to a first node not a backbone.) setting an input feature map having identical resolution from the first node as an essential input feature map; (Liang, page 4, section 3.1, Equation 2, “Fb i = Wb i ⊗(D(Fb i−1) +Pi),” and page 5, section 3.2.1, “We assume the intermediate nodes are fully connected with former nodes… we associate an edge importance weight to each edge” note: Fb i is the nodes prior node’s output, it’s a fixed term in the equation, entering every node’s construction after being brought to the node’s resolution.) and setting an input feature map from the backbone as a candidate input feature map, (Ghaisi, page 3, section 3.1, “The RNN controller selects any two candidate feature layers and a binary operation to com bine them into a new feature layer, where all feature layers may have different resolution.”, note: an input feature map having different resolution from the backbone maps to all feature layers may have different resolutions) wherein when an output feature map from another node except the first node is present, an output feature map from the another node is set as the candidate input feature map. (Ghaisi, page 4, section 3.1, “In Step 5, the newly-generated feature layer is appended to the list of existing input candidates and becomes a new candidate for the next merging cell.”) Regarding claim 9, The combination of Ghaisi, Wang, and Liang teaches, The method of claim 8, further comprising performing learning for searching for a one-shot neural network on a super-net comprising each node corresponding to the additionally extended input feature map candidate group. (Liang, page 4, section 3.2, “a super-net A, which is a fully-connected Multigraph DAG (directed acyclic graph). The node of DAG stands for feature maps (in the way of a feature pyramid), and there are six edges of different types between two nodes,” page 5, section 3.2, “The weights of super-net are fixed once this training is done (one-shot optimization)… In such DAG model, each node I ∈ {1,2,...,N} aggregates inputs from previous nodes,” note: super-net… Multigraph DAG maps to the supernet, the one-shot optimization maps to the one-shot training, and each node stands for feature maps and aggregates inputs from the previous node, therefore every node corresponds to the candidate group supplying it.) Regarding claim 12, The combination of Ghaisi, Wang, and Liang teaches, Ghaisi teaches, A method of searching for optimal neural network architecture based on channel concatenation, the method comprising: setting spatial size information of an output feature map; (Ghaisi, page 4, section 3.1, “Step 3. Select the output feature resolution.” note: the output feature resolution is the output’s spatial size, and select is the act of setting it) adjusting spatial size information of an input feature map candidate group so that the spatial size information of the input feature map candidate group corresponds to spatial size information of an output feature map; (Ghaisi, page 4, section 3.1, “The input feature layers are adjusted to the output resolution by nearest neighbor upsampling or max pooling if needed before applying the binary operation.” note: the input feature layers are the pool of the inputs which maps to the candidate group. The adjusted to the output resolution means their size is changed to equal the output’s size which maps to the spatial size made to correspond to the output. By nearest neighbor upsampling or max pooling is the mechanism that does the resizing.) additionally extending an output feature map that is results of the channel-concatenated operation to the input feature map candidate group; (Ghaisi, page 4, section 3.1, “the newly-generated feature layer is appended to the list of existing input candidates and becomes a new candidate for the next merging cell.” note: the newly generated feature layer maps to the node’s output which maps to the output feature map. Appended to the list of existing input candidates is the act of adding it back into the pool which maps to extended to the candidate group. Becomes a new candidate for the next merging cell confirms the added output is then available to later nodes.) Ghaisi does not teach, performing a channel-based concatenation operation on the input feature map candidate group; and performing learning for searching for a one-shot neural network on a super-net comprising each node corresponding to the additionally extended input feature map candidate group. Wang teaches, performing a channel-based concatenation operation on the input feature map candidate group; (Wang, page 5, section 3.3, “Nodes indicate feature maps and connect with all their predecessors in the supernet. After the search stage, only two predecessors will be selected for each node” and page 4, section 3.2, “suppose two features with expansion rates and base channels (e1, C1) and (e2,C2) are concatenated. A convolution is applied after the concatenation layer” note: the candidate group is a node’s predecessor, since nodes indicate feature maps and connect with all their predecessors in the supernet. The channel-based concatenation is the combining operation applies to those inputs, since the combination blocks are concatenated and it states that two candidate features are concatenated before a convolution is applied. This happens in the supernet and the predecessors are searched and then narrowed which maps to inside a search) Liang teaches, and performing learning for searching for a one-shot neural network on a super-net comprising each node corresponding to the additionally extended input feature map candidate group. (Liang, page 4, section 3.2, “a super-net A, which is a fully-connected Multigraph DAG (directed acyclic graph). The node of DAG stands for feature maps (in the way of a feature pyramid), and there are six edges of different types between two nodes,” page 5, section 3.2, “The weights of super-net are fixed once this training is done (one-shot optimization)… In such DAG model, each node I ∈ {1,2,...,N} aggregates inputs from previous nodes,” note: super-net… Multigraph DAG maps to the supernet, the one-shot optimization maps to the one-shot training, and each node stands for feature maps and aggregates inputs from the previous node, therefore every node corresponds to the candidate group supplying it.) It would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Ghaisi, Wang, and Liang because the channel concatenation allows stacking to keep each input’s channels intact instead of merging them away, letting the network weigh each input independently. Liang explains training one over-parameterized super-net a single time, rather than training each candidate architecture separately which makes it more cost efficient. Claim 13 is a machine claim, A server for searching for optimal neural network architecture based on channel concatenation, the server comprising: memory in which a program for searching for optimal neural network architecture based on channel concatenation is stored; and a processor configured to (Ghaisi, page 4, section 4.1, “We use the open-source implementation of RetinaNet1 for experiments… The models are trained on TPUs with 64 images in a batch”, that corresponds to method claim 12. Otherwise, they are not patentably distinguishable. Therefore, claim 13 is rejected for the same reasons as claim 12. Regarding claim 14, The combination of Ghaisi, Wang, and Liang teaches, The server of claim 13, wherein the processor sets the spatial size information of the output feature map so that the spatial size information of the output feature map corresponds to spatial size information of an output feature map in a corresponding node according to a path aggregation network (PAN) path sequence, (Liang, page 1, abstract, “we first introduce six heterogeneous information paths to build our search space, namely top-down, bottom-up, fusing splitting, scale-equalizing, skip-connect and none… each FPN candidate is represented by a densely-connected directed acyclic graph (each node is a feature pyramid and each edge is one of the six heterogeneous information paths).” and page 4, section 3.1, “each feature map (Ft i) is iteratively built by combining input pyramid feature map of the same level (Pi) and the higher-level output feature (Ft i+1): Ft i = Wt i ⊗(U(Ft i+1) +Pi)” note: the claim’s set the output size step requires the size to be set according to PAN where each node’s output resolution is determined by where the node sits in a top-down/bottom-up aggregation path. Liang shows its network is a graph whose nodes are feature pyramids and whose edges are information paths. Those paths are the top-down/bottom-up path aggregation paths. A node’s output being build at a specific pyramid level as it advances along the top-down path. The node’s position in the PAN sequence determines the resolution of the output it produces.) and determines an input feature map to be included in the input feature map candidate group based on the spatial size information of the output feature map. (Ghaisi, page 4, section 3.1, “Step 1. Select a feature layer hi from candidates. Step 2. Select another feature layer hj from candidates without replacement… Step 4. Select a binary op to combine hi and hj selected in Step 1 and Step 2 and generate a feature layer with the resolution selected in Step 3” note: Steps 1 and 2 perform the choosing aspect, they select a feature layer from candidates which means they pick which inputs are drawn from the candidate pool. The based on the output size maps because those chosen inputs are then combined in step 4 with the resolution selected in step 3, so the inputs are selected for the specific output resolution being built, which ties the input choice to the output size) Regarding claim 15, The combination of Ghaisi, Wang, and Liang teaches, wherein when a first node according to the PAN path sequence is a node connected to a backbone, (Ghaisi, page 3, section 3.1, “We follow the design by RetinaNet [23] which uses the last layer in each group of feature layers as the inputs to the first pyramid network… We use as inputs features in 5 scales {C3,C4,C5,C6,C7} with corresponding feature stride of {8,16,32,64,128} pixels.”, note: Ghaisi’s first pyramid network have inputs come directly from the backbone, and the node receives the backbone features.) the processor sets an input feature map having identical resolution from the backbone as an essential input feature map, (Liang, page 4, section 3.1, “each feature map (Ft i) is iteratively built by combining input pyramid feature map of the same level (Pi) and the higher-level output feature (Ft i+1): Ft i = Wt i ⊗(U(Ft i+1) +Pi),” and page 5, section 3.2.1, “We assume the intermediate nodes are fully connected with former nodes… we associate an edge importance weight to each edge.” note: Pi is the input pyramid feature of the same level, the same resolution feature from the backbone pyramid) and sets an input feature map having different resolution from the backbone as a candidate input feature map, (Ghaisi, page 3, section 3.1, “The RNN controller selects any two candidate feature layers and a binary operation to com bine them into a new feature layer, where all feature layers may have different resolution.”, note: an input feature map having different resolution from the backbone maps to all feature layers may have different resolutions) wherein when an output feature map from another node is present, the processor sets an output feature map from the another node as the candidate input feature map. (Ghaisi, page 4, section 3.1, “In Step 5, the newly-generated feature layer is appended to the list of existing input candidates and becomes a new candidate for the next merging cell.”) Regarding claim 16, wherein when a second node according to the PAN path sequence is a node connected to a first node not a backbone, (Liang, page 4, section 3.1, “Each feature map (Fb i ) is obtained by merging the input feature maps (Pi) of the same level, and the output feature map below it (Fb i−1): Fb i = Wb i ⊗(D(Fb i−1) +Pi),” note: in the bottom-up path, node i is fed by Fb i−1, the output of the previous node, not the backbone. That is the claim’s second node… connected to a first node not a backbone.) the processor sets an input feature map having identical resolution from the first node as an essential input feature map, (Liang, page 4, section 3.1, Equation 2, “Fb i = Wb i ⊗(D(Fb i−1) +Pi),” and page 5, section 3.2.1, “We assume the intermediate nodes are fully connected with former nodes… we associate an edge importance weight to each edge” note: Fb i is the nodes prior node’s output, it’s a fixed term in the equation, entering every node’s construction after being brought to the node’s resolution.) and sets an input feature map from the backbone as a candidate input feature map, (Ghaisi, page 3, section 3.1, “The RNN controller selects any two candidate feature layers and a binary operation to com bine them into a new feature layer, where all feature layers may have different resolution.”, note: an input feature map having different resolution from the backbone maps to all feature layers may have different resolutions) wherein when an output feature map from another node except the first node is present, the processor sets an output feature map from the another node as the candidate input feature map. (Ghaisi, page 4, section 3.1, “In Step 5, the newly-generated feature layer is appended to the list of existing input candidates and becomes a new candidate for the next merging cell.”) Regarding claim 18, The combination of Ghaisi, Wang, and Liang teaches, The server of claim 13, wherein the processor applies a structure parameter and a softmax function to each of candidate input feature maps except an essential input feature map in the input feature map candidate group. (Wang, page 6, section 3.3, “For each fusion block, we introduce architecture parameters αe and αo to denote the importance of edges and operations. Suppose ˜α=softmax(α) is the normalized weight. The fused feature zj = PNG media_image1.png 26 197 media_image1.png Greyscale where O is the candidate operation set, xi is the features of predecessors” note: Its architecture parameters a are the structure parameters, they denote the importance of edges and an edge is a candidate input feeding a node (x_i, the features of predecessors). The softmax function is applied to those parameters.) Claim(s) 6, 10, 17, 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ghaisi and Wang in view of Liang and in further view of Liu. Regarding claim 6, Ghaisi teaches, The method of claim 1, wherein the adjusting of the spatial size information of the input feature map candidate group so that the spatial size information of the input feature map candidate group corresponds to spatial size information of an output feature map comprises adjusting the spatial size information of the input feature map to be identical with the spatial size information of the output feature map by applying (Ghaisi, page 4, section 3.1, “The input feature layers are adjusted to the output resolution by nearest neighbor upsampling or max pooling if needed before applying the binary operation.”) Ghaisi does not teach, a convolution product operation having a stride size of 2 or more and a 1×1 convolution product operation Liu teaches, a 1×1 convolution product operation (Liu, page 5, 3.1.1, “1×1 convolutions are inserted as necessary.”) Liang teaches, a convolution product operation having a stride size of 2 or more (Liang, page 4, section 3.1, “where Ws 1,Ws 0,Ws −1 are 3 × 3 deformable convolution filters and the stride of Ws −1 is set to 2 to down-sample.”) It would be obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Ghaisi, Wang, Liang, and Liu because Ghaisi establishes that an oversized input has to be shrunk down to the output’s size before it could be combined, but does the shrinking by max pooling. Liang uses a convolution with its stride set to 2 to downsample and Liu inserts 1x1 convolutions as needed. Pooling could result in discarding values on a fixed rule, while a strided convolution shrinks the map using weights the network can learn and the 1x1 convolution is provided to reconcile the channel count after. Regarding claim 10, The combination of Ghaisi, Wang, Liang, and Liu teaches, The method of claim 9, wherein the performing of the learning for searching for the one-shot neural network on the super-net comprises removing an input from a node having a structure parameter that satisfies a predetermined condition in the super-net on which the learning has been completed. (Liu, page 4, section 2.4, “To form each node in the discrete architecture, we retain the top-k strongest operations (from distinct nodes) among all non-zero candidate operations collected from all the previous nodes. The strength of an operation is defined as exp(α(i,j) o ) o∈Oexp(α(i,j) o .”) Regarding claim 17, The combination of Ghaisi, Wang, Liang, and Liu teaches, adjusts the spatial size information of the input feature map to be identical with the spatial size information of the output feature map, (Ghaisi, page 4, section 3.1, “The input feature layers are adjusted to the output resolution by nearest neighbor upsampling or max pooling if needed before applying the binary operation.) by applying a convolution product operation having a stride size of 2 or more and a 1×1 convolution product operation when spatial size information of an input feature map included in the input feature map candidate group is greater than the spatial size information of the output feature map (Liang, page 4, section 3.1, “where Ws 1,Ws 0,Ws −1 are 3 × 3 deformable convolution filters and the stride of Ws −1 is set to 2 to down-sample” and Liu, page 5, section 3.1.1, “1×1 convolutions are inserted as necessary” note: the claim needs a convolution with stride greater than or equal to 2 to shrink an oversized input. Liang has a convolution filter whose stride… is set to 2 to down-sample. Liu named the second operation which is a 1x1 convolution.) and applying an up-sampling operation and a 1×1 convolution product operation when the spatial size information of the input feature map included in the input feature map candidate group is smaller than the spatial size information of the output feature map. (Ghaisi, page 4, section 3.1, “The input feature layers are adjusted to the output resolution by nearest neighbor upsampling” and Liu, page 5, section 3.1.1, “1×1 convolutions are inserted as necessary”) Regarding claim 19, The combination of Ghaisi, Wang, Liang, and Liu teaches, The server of claim 13, wherein the processor removes an input from a node having a structure parameter that satisfies a predetermined condition in a super-net on which the learning has been completed. (Liu, page 4, section 2.4, “To form each node in the discrete architecture, we retain the top-k strongest operations (from distinct nodes) among all non-zero candidate operations collected from all the previous nodes. The strength of an operation is defined as exp(α(i,j) o ) o∈Oexp(α(i,j) o .”) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to UYEN-NHU PHAM TRAN whose telephone number is (571)272-1559. The examiner can normally be reached Monday - Friday 7:30-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /U.P.T./ Examiner, Art Unit 2124 /MIRANDA M HUANG/ Supervisory Patent Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Nov 15, 2023
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month