Prosecution Insights
Last updated: October 04, 2026
Application No. 18/129,523

SEARCH METHOD AND APPARATUS

Non-Final OA §103
Filed
Mar 31, 2023
Priority
Dec 03, 2020 — CN 202011396358.2 +1 more
Examiner
CHUANG, SU-TING
Art Unit
2146
Tech Center
2100 — Computer Architecture & Software
Assignee
BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO., LTD.
OA Round
2 (Non-Final)
51%
Grant Probability
Moderate
2-3
OA Rounds
1y 0m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
58 granted / 113 resolved
-3.7% vs TC avg
Strong +39% interview lift
Without
With
+39.4%
Interview Lift
resolved cases with interview
Typical timeline
4y 6m
Avg Prosecution
19 currently pending
Career history
136
Total Applications
across all art units

Statute-Specific Performance

§101
26.5%
-13.5% vs TC avg
§103
47.6%
+7.6% vs TC avg
§102
11.2%
-28.8% vs TC avg
§112
12.4%
-27.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 113 resolved cases

Office Action

§103
DETAILED ACTION This action is in response the communications filed on 05/15/2026 in which claims 1, 3, 8, 10, 15 and 17 are amended, claims 2, 9 and 16 are canceled and claims 1, 3-8, 10-15 and 17-20 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 3-5, 7-8, 10-12, 14-15 and 17-19 rejected under 35 U.S.C. 103 as being unpatentable over Cai ("Once for all: Train one network and specialize it for efficient deployment" 20200429) in view of Roth (US 20210374502 A1, filed on 20200601) in further view of Bender ("Understanding and Simplifying One-Shot Architecture Search" 2018) In regard to claims 1, 8 and 15, Cai teaches: A search method, comprising: with a processing circuitry… with the processing circuitry… with the processing circuitry… with the processing circuitry… with the processing circuitry… (Cai, p. 7, 4 EXPERIMENTS "Training Details… The full network is trained for 180 epochs with batch size 2048 network on 32 GPUs... further fine-tune the full network. The whole training process takes around 1,200 GPU hours on V100 GPUs") obtaining… network construction information corresponding to a target task, (Cai, p. 1 Abstract "In this work, we propose to train a once-for-all (OFA) network that supports diverse architectural settings… OFA is the winning solution for the 3rd Low Power Computer Vision Challenge (LPCVC), DSP classification track and the 4th LPCVC, both classification track and detection track. [e.g. a target task]"; see the next limitation for network construction information) the network construction information comprising search space information, sample data, and a search indicator; (Cai, p. 4, Architecture space [search space information] "Our once-for-all network provides one model but supports many sub-networks of different sizes, covering four important dimensions of the convolutional neural networks (CNNs) architectures, i.e., depth, width, kernel size, and resolution... We allow each unit to use arbitrary numbers of layers (denoted as elastic depth); For each layer, we allow to use arbitrary numbers of channels (denoted as elastic width) and arbitrary kernel sizes (denoted as elastic kernel size)...") (Cai, p. 7, 4 "we first apply the progressive shrinking algorithm to train the once-for-all network on ImageNet [e.g. sample data]") (Cai, p. 2, 1 Introduction "Given the target hardware and constraint, [a search indicator] a predictor-guided architecture search... is conducted to get a specialized sub-network") PNG media_image1.png 352 448 media_image1.png Greyscale constructing… a supernetwork based on the search space information… training… the supernetwork based on the sample data, the supernetwork comprising a plurality of sub-networks; (Cai, p. 5, Progressive Shrinking "The once-for-all network comprises many sub-networks of different sizes where small sub-networks are nested in large sub-networks. [the supernetwork comprising a plurality of sub-networks]… where we start with training the largest neural network with the maximum kernel size (e.g., 7), depth (e.g., 4), and width (e.g., 6). [constructing a supernetwork based on the search space information] Next, we progressively fine-tune the network to support smaller sub-networks by gradually adding them into the sampling space...") (Cai, p. 7, 4 Experiments "we first apply the progressive shrinking algorithm to train the once-for-all network on ImageNet [training the supernetwork based on the sample data]") searching… a sub-network from the trained supernetwork based on the search indicator, to obtain a target network for performing the target task; and (Cai, p. 6, 3.4 Specialized model deployment with once-for-all network "Having trained a once-for-all network, [from the trained supernetwork] the next stage is to derive the specialized sub-network [searching a sub-network... to obtain the target network] for a given deployment scenario. The goal is to search for a neural network that satisfies the efficiency (e.g., latency, energy) constraints [based on the search indicator] on the target hardware while optimizing the accuracy...") deploying… the target network on a terminal device to perform the target task. (Cai, p. 6, 3.4 Specialized model deployment with once-for-all network "Having trained a once-for-all network, the next stage is to derive the specialized sub-network for a given deployment scenario."; p. 2, Figure 1 "Figure 1… Given a deployment scenario, a specialized sub-network is directly selected from the once-for-all network without training."; also see Figure 1, a specialized sub-network is deployed to Mobile AI or Tiny AI, i.e. the target network is deployed to a terminal device) PNG media_image2.png 572 566 media_image2.png Greyscale Cai teaches the concept of a supernetwork as an OFA network (comprising many sub-networks). Cai does not explicitly teach the term ‘supernetwork,’ but Roth teaches: constructing a supernetwork… (Roth, [0063] "neural network 104 is a supernet, which may also be referred to as a supernetwork, model architecture, and/or a neural network comprising a plurality of neural networks.") It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Cai to incorporate the teachings of Roth by using Cai's OFA as Roth's supernet and by including data adaptation process. Doing so would effectively adapt a model to a target domain and result in an optimal path that defines a locally adapted sub-network. (Roth, [0069] "once supernet 104 is trained, a sub-network s0 is found, at each client 106, 108 through supernet 104, effectively adapting a model to a target domain… during adaptation, model parameters ϕ stay fixed and only path weights are optimized... this results in an optimal path... that defines a locally adapted sub-network s0 ∈ S.") Cai and Roth do not teach, but Bender teaches: the search space information comprises branch construction forms corresponding to a plurality of modules used to construct the supernetwork, the constructing the supernetwork comprises: (Bender, p. 3, 3.1. Search Space Design "This approach is applied to a much larger model as shown in Figure 3. Following Zoph et al. (2017), our network [the supernetwork] is composed of several identical cells which are stacked on top of each other. Each cell [modules] is divided into a fixed number of choice blocks. [branch construction forms]"; see Figure 3, cell is [module]) PNG media_image3.png 396 1338 media_image3.png Greyscale extracting, from the search space information, a branch construction form corresponding to each of the modules; (Bender, p. 3, 3.1. Search Space Design "The number of choice blocks within each cell, N_choice, is a hyper-parameter of the search space. In our experiments, we set N_choice = 4. [extracting/selecting a branch construction form (choice blocks)]") constructing, for each of the modules, branches of the module according to a branch construction form corresponding to the module; (Bender, p. 3, 3.1. Search Space Design "Each choice block can consume the outputs of the two most recent cells in the network… Each choice block can select up to two operations from a menu of seven possible options…") connecting the branches of the module in parallel to construct the module; and (Bender, p. 4, Figure 3 "Diagram of the one-shot architecture used in our experiments. Solid lines indicate components that are present in every architecture"; see Figure 3, choice 1, 2 and 3 are in parallel in the solid lines) connecting the modules in series to construct the supernetwork; (Bender, p. 3, 3.1. Search Space Design "our network is composed of several identical cells which are stacked on top of each other."; see Figure 3, the leftmost diagram, all the blocks are connected from the input to output) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Cai and Roth to incorporate the teachings of Bender by including the search space as proposed in Bender. Doing so would have a search space that is large and expressive enough to capture a diverse set of interesting candidate architectures. (Bender, p. 2, 3.1. Search Space Design "the search space should be large and expressive enough to capture a diverse set of interesting candidate architectures.") Claims 8 and 15 recite substantially the same limitation as claim 1, therefore the rejection applied to claim 1 also apply to claims 8 and 15. In addition, Cai teaches: (claim 8) A search apparatus, comprising: a memory operable to store computer-readable instructions; and a processor circuitry operable to read the computer-readable instructions, the processor circuitry when executing the computer-readable instructions is configured to: (claim 15) A non-transitory machine-readable media, having instructions stored on the machine-readable media, the instructions configured to, when executed, cause a machine to: (Cai, p. 7, 4 EXPERIMENTS "Training Details… The full network is trained for 180 epochs with batch size 2048 network on 32 GPUs... further fine-tune the full network. The whole training process takes around 1,200 GPU hours on V100 GPUs"; GPU and training details inherently teach all the computer components) In regard to claims 3, 10 and 17, Cai and Roth do not teach, but Bender teaches: wherein the connecting the modules in series to construct the supernetwork comprises: connecting the modules in series and connecting respective inputs and outputs of the modules to construct the supernetwork. (Bender, p. 3, 3.1. Search Space Design "our network is composed of several identical cells which are stacked on top of each other."; see Figure 3, the leftmost diagram, all the blocks are connected from the input, stem…, cell…, output; also see prior art Zoph "Figure 2. Scalable architectures for image classification consist of two repeated motifs termed Normal Cell and Reduction Cell") The rationale for combining the teachings of Cai, Roth and Bender is the same as set forth in the rejection of claim 1. In regard to claims 4, 11 and 18, Cai teaches: the training the supernetwork based on the sample data comprises: (Cai, p. 7, 4 Experiments "we first apply the progressive shrinking algorithm to train the once-for-all network on ImageNet [training the supernetwork based on the sample data]") Cai does not teach, but Roth teaches: wherein the sample data comprises training data, the training data comprises training sample data and reference sample data corresponding to the training sample data, and (Roth, [0070] "gi is a ground truth label map [reference sample data] at a given voxel i. [training sample data]... during adaptation, model parameters ϕ stay fixed and only path weights are optimized for one epoch on a local validation set. [training data]") selecting a sub-network from the supernetwork, and inputting the training sample data into the selected sub-network for forward calculation, to obtain data outputted by the selected sub-network; and (Roth, [0069] "a Dice loss is applied as a loss function, which works well in segmentation tasks with an unbalance in an amount of foreground/background regions: min(L_Dice) = (… pigi...) (2)… pi is a predicted probability from a final sigmoid activated output layer of supernet f (X) [forward calculation, to obtain data outputted by the selected sub-network] and gi is a ground truth label map at a given voxel i. [training sample data] In at least one embodiment, once supernet 104 is trained, a sub-network s0 is found, [selecting a sub-network] at each client 106, 108 through supernet 104, effectively adapting a model to a target domain.") performing backpropagation on the selected sub-network based on the data outputted by the selected sub-network and the reference sample data. (Roth, [0069] "a Dice loss is applied as a loss function... min(L_Dice) = (… pigi...) (2)… pi is a predicted probability from a final sigmoid activated output layer of supernet f (X) [the data outputted by the selected sub-network] and gi is a ground truth label map [the reference sample data] at a given voxel i... during adaptation, model parameters ϕ stay fixed and only path weights are optimized for one epoch on a local validation set. In at least one embodiment, this results in an optimal path... that defines a locally adapted sub-network s0 ∈ S."; [0077] "parameters of new sub-networks are updated during gradient back-propagation.") The rationale for combining the teachings of Cai and Roth is the same as set forth in the rejection of claim 1. In regard to claims 5, 12 and 19, Cai and Roth do not teach, but Bender teaches: wherein the supernetwork comprises a plurality of modules connected in series, each of the plurality of modules comprises a plurality of branches connected in parallel, and the selecting a sub-network from the supernetwork comprises: (Bender, p. 3, 3.1. Search Space Design "This approach is applied to a much larger model as shown in Figure 3. Following Zoph et al. (2017), our network [the supernetwork] is composed of several identical cells which are stacked on top of each other. Each cell [modules] is divided into a fixed number of choice blocks. [branch construction forms]"; see Figure 3, cell is [module]) selecting a branch from each of the plurality of modules of the supernetwork, and connecting, in series, branches selected from the modules to form a sub-network. (Bender, p. 3, 3.1. Search Space Design "Each choice block can consume the outputs of the two most recent cells in the network. This means that each choice block can select from up to five possible inputs: two from previous cells and up to three from previous choice blocks within the same cell."; p. 4, Figure 3 "-----> Edge selected by architecture search… dashed lines indicate optional components that are part of the search space."; see the dashed lines [a branch] in the cell) The rationale for combining the teachings of Cai, Roth and Bender is the same as set forth in the rejection of claim 1. In regard to claims 7 and 14, Cai teaches: wherein the method further comprises: training the target network based on the sample data. (Cai, p. 8, Table 1 "'#25' denotes the specialized sub-networks are fine-tuned [training the target network based on data] for 25 epochs after grabbing weights from the once-for-all network"; p. 8, Comparison with NAS on Mobile Devices "We can further improve the top1 accuracy to 76.4% by fine-tuning the specialized sub-network for 25 epochs and to 76.9% by fine-tuning for 75 epochs") Claims 6, 13 and 20 rejected under 35 U.S.C. 103 as being unpatentable over Cai, Roth and Bender as applied to claims 1, 8 and 15, and in further view of Li ("Random Search and Reproducibility for Neural Architecture Search" 20190730) In regard to claims 6, 13 and 20, Cai teaches: the searching the sub-network from the trained supernetwork based on the search indicator to obtain the target network comprises: (Cai, p. 6, 3.4 Specialized model deployment with once-for-all network "Having trained a once-for-all network, [from the trained supernetwork] the next stage is to derive the specialized sub-network [searching the sub-network... to obtain the target network] for a given deployment scenario. The goal is to search for a neural network that satisfies the efficiency (e.g., latency, energy) constraints [based on the search indicator] on the target hardware while optimizing the accuracy...") … selecting, from the sub-networks, a sub-network with a best performance and satisfying the search indicator as the target network based on the performance parameter. (Cai, p. 2, Figure 1 "Left:… Given a deployment scenario, a specialized sub-network is directly selected [selecting... a sub-network as the target network] from the once-for-all network without training."; p. 6, 3.4 Specialized model deployment with once-for-all network "Having trained a once-for-all network, the next stage is to derive the specialized sub-network for a given deployment scenario. The goal is to search for a neural network that satisfies the efficiency (e.g., latency, energy) constraints [satisfying the search indicator] on the target hardware while optimizing the accuracy. [a best performance]... we randomly sample 16K sub-networks [the sub-networks] with different architectures and input image sizes... These [architecture, accuracy] pairs are used to train an accuracy predictor to predict the accuracy of a model [the performance parameter] given its architecture and input image...") Cai, Roth and Bender do not teach, but Li teaches: wherein the sample data comprises test data, and (Li, p. 7, 3 METHODOLOGY "After training the shared weights for a certain number of epochs, we use these trained shared weights to evaluate the performance of a number of randomly sampled architectures on a separate held out dataset. [test data]"; a held-out dataset is deliberately separated from the original dataset, i.e. the original sample data comprises a head-out dataset [test data]) searching for sub-networks from the trained supernetwork based on a search algorithm; (Li, p. 7, 3 METHODOLOGY "In order to combine random search [a search algorithm] with weight-sharing, we simply use randomly sampled architectures [sub-networks] to train the shared weights. Shared weights are updated by selecting a single architecture for a given minibatch... the number of architectures used to update the shared weights is equivalent to the total number of minibatch training iterations."; p. 14, 4.2.2 Impact of Meta-Hyperparameters "each version of random search with weight-sharing...") performing a performance test on the sub-networks by inputting the test data into the sub-networks obtained through search, to obtain a performance parameter for each of the sub-networks; and (Li, p. 7, 3 METHODOLOGY "After training the shared weights for a certain number of epochs, we use these trained shared weights to evaluate the performance [performing a performance test] of a number of randomly sampled architectures [the sub-networks obtained through search] on a separate held out dataset. [test data]"; p. 14, 4.2.2 Impact of Meta-Hyperparameters "In stage (1), we train the shared weights and use them to evaluate a given number of randomly sampled architectures on the test set."; p. 4, Evaluation Method "For each hyperparameter configuration considered by a search method... subsequently measuring its quality, e.g., its predictive accuracy") It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to have modified Cai, Roth and Bender to incorporate the teachings of Li by including a random search algorithm. Doing so would achieving a state-of-the-art result. (Li, p. 1, Abstract "a novel random search with weight-sharing algorithm on two standard NAS benchmarks—PTB and CIFAR-10…. random search with weight-sharing outperforms random search with early-stopping, achieving a state-of-the-art NAS result on PTB and a highly competitive result on CIFAR-10.") Response to Arguments Applicant's arguments with respect to the rejection of the claims under 35 U.S.C. 101 have been fully considered and are sufficient to overcome the rejection. The 101 rejection has been withdrawn. Applicant's arguments with respect to the rejection of the claims under 35 U.S.C. 103 have been fully considered but they are not persuasive: Argument: (p. 12) … However, as shown in FIG. 3 of Bender, the choice blocks are not choice connected in parallel because block 2 consumes output of choice block 1 and choice block 3 consumes outputs of choice block 1 and choice block 2. Therefore, Bender fails to disclose, teach, or suggest connecting the branches of the module in parallel to construct the module, as recited in amended claim 1. PNG media_image4.png 346 376 media_image4.png Greyscale Response: Figure 3 of Bender illustrates required and optional elements through solid and dashed lines, respectively. The solid lines link the choice blocks in parallel within the cell, directly corresponding to the claimed feature of connecting branches in parallel within the module. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Zoph ("Learning Transferable Architectures for Scalable Image Recognition" 20180411) teaches scalable architectures. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SU-TING CHUANG whose telephone number is (408)918-7519. The examiner can normally be reached Monday - Thursday 8-5 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Usmaan Saeed can be reached at (571) 272-4046. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /S.C./Examiner, Art Unit 2146 /USMAAN SAEED/Supervisory Patent Examiner, Art Unit 2146
Read full office action

Prosecution Timeline

Mar 31, 2023
Application Filed
Feb 19, 2026
Non-Final Rejection mailed — §103
Mar 18, 2026
Examiner Interview Summary
Mar 18, 2026
Applicant Interview (Telephonic)
May 15, 2026
Response Filed
Jul 23, 2026
Final Rejection mailed — §103
Sep 21, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12718085
GENERATING NAVIGATIONAL TARGET RECOMMENDATIONS USING PARALLEL NEURAL NETWORKS
5y 4m to grant Granted Aug 25, 2026
Patent 12711359
METHOD AND DEVICE FOR PROCESSING DATA BASED ON MULTI-LAYER PERCEPTRONS
4y 3m to grant Granted Aug 18, 2026
Patent 12645997
INDIVIDUALIZED CLASSIFICATION THRESHOLDS FOR MACHINE LEARNING MODELS
3y 3m to grant Granted Jun 02, 2026
Patent 12626164
SYSTEM AND METHOD FOR REDUCTION OF DATA TRANSMISSION BY DATA RECONSTRUCTION
4y 0m to grant Granted May 12, 2026
Patent 12626106
MACHINE LEARNING MODELS FOR BEHAVIOR UNDERSTANDING
3y 11m to grant Granted May 12, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
51%
Grant Probability
91%
With Interview (+39.4%)
4y 6m (~1y 0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 113 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month