Prosecution Insights
Last updated: October 02, 2026
Application No. 18/772,334

METHOD FOR TRAINING A CONVOLUTIONAL NEURAL NETWORK COMPRISING NODES ARRANGED IN LAYERS AND A PRUNING MASK

Non-Final OA §103
Filed
Jul 15, 2024
Priority
Jul 18, 2023 — DE 10 2023 206 789.7
Examiner
DASGUPTA, SHOURJO
Art Unit
Tech Center
Assignee
Robert Bosch GmbH
OA Round
1 (Non-Final)
65%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
303 granted / 465 resolved
+5.2% vs TC avg
Strong +39% interview lift
Without
With
+39.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
26 currently pending
Career history
491
Total Applications
across all art units

Statute-Specific Performance

§101
12.6%
-27.4% vs TC avg
§103
57.9%
+17.9% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
16.3%
-23.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 465 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Allowable Subject Matter Claim 6 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office Action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-5, 7, and 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Non-Patent Literature “Pruning Convolutional Neural Networks for Resource Efficient Inference” (“MOLCHANOV”) in view of Non-Patent Literature “Disentangled Differentiable Network Pruning” (“GAO”). Regarding claim 1, MOLCHANOV teaches A computer implemented method for training a convolutional neural network including nodes arranged in layers (pages 1-2, sections 1-2, discussing “convolutional neural networks” (CNN) that are subject to pruning and fine tuning so that the resulting CNN is sufficient in its performance (see especially section 2’s 1st paragraph), i.e., a “training” process, where the CNN as taught is composed of “layers” (page 1, section 1, 2nd paragraph) and that the layers are comprised of neurons (page 2’s Figure 1 and page 13’s section A.1), i.e., “nodes”) and a pruning mask (page 3, where it continues section 2’s introduction and just before section 2.1 begins: “… an external switch which determines if a particular feature map is included or pruned during feed-forward propagation …”), the method comprising the following steps: providing at least one set of labeled training data (page 2’s section 2, 3rd paragraph: “Consider a set of training examples D = X = {x0, x1, ..., xN}, Y = {y0, y1, ..., yN} , where x and y represent an input and a target output, respectively.”, where the known target output is akin to a supervised approach that provides labeling training data as recited); initializing the convolutional neural network including at least one convolutional layer and a pruning mask following the convolutional layer and passing the training data through the convolutional neural network (page 2’s section 2, 3rd paragraph, where it mentions the training examples (as cited to just above), where the Examiner understands that the CNN is initially trained with the aforementioned training examples prior to subsequent steps (as shown in Figure 1) of neuron importance evaluation, removal of least important neurons, further fine tuning, and so forth, such that the initial training is an initializing as recited) (and further, the “external switch” (as noted above, and mentioned in section 2) is applied to each particular feature map during feed-forward propagation, and hence would be understood to follow a convolutional layer); computing a loss function and comparing the loss function with labels of the training data to quantify the convolutional neural networks process (page 2’s section 2, 3rd paragraph: “The network’s parameters … are optimized to minimize a cost value C(D|W). The most common choice for a cost function C(·) is a negative log-likelihood function. A cost function is selected independently of pruning and depends only on the task to be solved by the original network.”, where the cost value as taught, e.g., in relation to the training of a neural network generally and a CNN as taught, is understood to be the different between the input data as processed verses the input data as labeled); and minimizing the loss function by backward propagation including determining a gradient of the loss function for determining a direction for optimizing the nodes of the convolutional neural network (page 1’s section 1, 2nd paragraph: “With the goal of speeding up inference, we prune entire feature maps so the resulting networks may be run efficiently even on embedded devices. We interleave greedy criteria-based pruning with fine-tuning by backpropagation, a computationally efficient procedure that maintains good generalization in the pruned network.”, where the fine tuning as mentioned and as indicated in page 2’s Figure 1 is understood to be directed to “minimizing … loss” and “optimizing … nodes” as recited to better promote the CNN’s accuracy, where the pruning by way of this approach is understood to involve a computation of gradients (e.g., (i) page 4, between equations 7 and 8, and (ii) on page 5, 2nd paragraph (where the gradient’s approach of zero signals sufficiency of training), and (iii) page 16, section A.7)); wherein the passing of the training data through the pruning mask includes multiplying structures si of input of the pruning mask with pruning parameters Mi, the pruning parameters Mi for the structures si being 0 or 1 (again citing to page 3, where it continues section 2’s introduction and just before section 2.1 begins: the pruning gate as taught expresses a binary indication for each feature map as to its inclusion, equivalent to a 0 or 1 expression as recited, and is essentially shown to influence the parameters W, such that a 0 would null out the parameter and hence the feature map, and a 1 would promote/express the feature map), but does not teach wherein the pruning mask is approximated by an approximation function during backward propagation. Rather the Examiner relies upon GAO to teach what MOLCHANOV lacks, see e.g., Gao’s comparable framework for fine tuning a CNN through a pruning approach (Abstract, Introduction) using a mask or its equivalent having “learnable parameters to decide whether to prune the channel” (section 3.1, 2nd paragraph, and also section 3.4), and where the parameters as mentioned are learned using an approximation function approach (e.g., section 3.2 mentioning use of sigmoid function to facilitate gradient calculations (i.e., during back propagation), in combination with section 3.3’s Smoothstep function). The references are similarly directed to network pruning to facilitate more efficient performance in contemplation of deployment to resource-constrained host devices. Hence, they are analogous. It would have been obvious to one of ordinary skill in the art to incorporate Gao’s Algorithm 1, which features aspects discussed just above, as a way to determine the mask to be used in Molchanov’s framework, with a reasonable expectation of success, for purposes of improved pruning that can take into account the importance of the subjected features in relation to one another, as discussed in Gao’s Abstract. Regarding claim 2, Molchanov in view of Gao teaches The computer implemented method according to claim 1, as discussed above, and further wherein the pruning parameters Mi are determined from a helper vector Zi, wherein the helper vector Zi is a trainable parameter of the convolutional neural network, and wherein the approximation function is a function of the helper vector Zi (Gao’s mask vector a, as mentioned in relation to Figure 1 on page 4, section 3.2, and page 8’s Algorithm 1 per the inner for-loop’s 3rd step ). The motivation for combining the references is as discussed above in relation to claim 1. Regarding claim 3, Molchanov in view of Gao teaches The computer implemented method according to claim 2, as discussed above, and further wherein the pruning parameters Mi are derived from rounding a sigmoid function, wherein the helper vector is input of the sigmoid function (Gao’s page 5, section 3.1, discussing the use of the sigmoid function and its rounding aspect, and page 7, section 3.3, 2nd full paragraph, discussing the use of the same function to determine the mask vector). The motivation for combining the references is as discussed above in relation to claim 1. Regarding claim 4, Molchanov in view of Gao teaches The computer implemented method according to claim 3, as discussed above, and further wherein the helper vector Zi is initialized with a value of 0 or close to 0 (Gao’s page 5, section 3.2 discussing the mask vector’s state selectively being 0 or 1, as subject to an importance determination, and where it reasons that defaulting to not important/0 in an initialized state prior to the active determination of importance is one of two possible approaches given the binary state of the value). The motivation for combining the references is as discussed above in relation to claim 1. Regarding claim 5, Molchanov in view of Gao teaches The computer implemented method according to claim 1, as discussed above, and further wherein a sum of all pruning parameters Mi is greater than 0 (Molchanov’s pruning gate and Gao’s mask, both of which have been discussed to selectively express or void a subjected feature, must necessarily have at least one non-zero element among all the values, if the gate/mask is to permit any expression of features arrived at through the convolution layer just prior; said another way, if all the parameters summed to zero, then they would all be zero, and the gate/mask as applied would not yield any features for outputting). The motivation for combining the references is as discussed above in relation to claim 1. Regarding claim 7, Molchanov in view of Gao teaches The computer implemented method according to claim 1, as discussed above, and further wherein the convolutional neural network is configured for processing image or audio data (Molchanov’s Abstract and Introduction section, 1st paragraph, both discussing image data as being subjected to the CNN’s application, and see also top paragraph of page 3: “Since we focus our analysis on pruning feature maps from convolutional layers, let us denote a set of image feature maps …”). The motivation for combining the references is as discussed above in relation to claim 1. Regarding claims 9-10, the claims include the same or similar limitations as discussed above in relation to claim 1, and are therefore rejected under the same rationale. Claim 9 additionally recites a non-transitory computer readable data carrier on which is stored program code to essentially perform the steps discussed above in relation to claim 1. Similarly, claim 10 recites a computer to essentially perform those same steps. The Examiner reasons that Molchanov and Gao are clearly computer-implemented frameworks, where typical computer elements such as a processors and memory are understood to be used and necessary for those frameworks to function as described, and that these typical elements read on the bolded features mentioned just above. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over MOLCHANOV in view of GAO and further in view of Non-Patent Literature “Convolutional Neural Network-Based Technique for Gaze Estimation on Mobile Devices” (“AKINYELU”). Regarding claim 8, Molchanov in view of Gao teaches computer vision and image processing and so forth, as discussed above, but does not teach The computer implemented method according to claim 1, wherein the convolutional neural network is configured as an eye tracking application. Rather, the Examiner relies upon AKINYELU to teach what Molchanov etc. otherwise lacks, see e.g., Akinyelu’s application of a comparable image-based CNN application to the problem of gaze estimation, which explicitly includes “eye tracking” in the reference’s opening paragraph. Like Molchanov, Akinyelu contemplates a CNN as applied to image data to make some desired determination/output. Hence, they are directed to similar networks categorically, and hence are analogous. It would have been obvious to one of ordinary skill in the art to extend Molchanov’s modified framework, which contemplates image, vision, and video data as its subject (per its Introduction section), to address a particular known problem/challenge falling within that same space, such as eye tracking per Akinyelu, with a reasonable expectation of success. The motivation to provide better understanding of features as detected in relation to one another, as Molchanov in view of Gao would provide, would be likewise apt to the specific problem/challenge area that Akinyelu more specifically contemplates. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHOURJO DASGUPTA whose telephone number is (571)272-7207. The examiner can normally be reached M-F 8am-5pm CST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 571 272 4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHOURJO DASGUPTA/Primary Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

Jul 15, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731146
System, Method, and Computer Program Product for Determining a Reason for a Deep Learning Model Output
4y 6m to grant Granted Sep 08, 2026
Patent 12731061
QUANTUM CIRCUIT FOR ESTIMATING MATRIX SPECTRAL SUMS
4y 1m to grant Granted Sep 08, 2026
Patent 12730554
DYNAMIC RESIZABLE MEDIA ITEM PLAYER
2y 2m to grant Granted Sep 08, 2026
Patent 12699932
System, Method, and Computer Program Product for Generating Error Rate Predictions Based on Machine Learning Using Incremental Backpropagation
4y 3m to grant Granted Aug 04, 2026
Patent 12694285
Method and Apparatus for Training a Quantized Classifier
4y 11m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+39.3%)
3y 5m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 465 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month