Prosecution Insights
Last updated: October 01, 2026
Application No. 18/355,243

Systems and Methods for Machine-Learned Models Having Convolution and Attention

Final Rejection §103
Filed
Jul 19, 2023
Priority
May 27, 2021 — provisional 63/194,077 +1 more
Examiner
COLEMAN, PAUL
Art Unit
2126
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
2 (Final)
65%
Grant Probability
Favorable
3-4
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
17 granted / 26 resolved
+10.4% vs TC avg
Strong +47% interview lift
Without
With
+47.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
15 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
31.3%
-8.7% vs TC avg
§103
47.0%
+7.0% vs TC avg
§102
4.2%
-35.8% vs TC avg
§112
16.9%
-23.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 26 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of Claims The present application is being examined under the claims filed July 20, 2026. The status of the claims are as follows: Claims 23-42 are pending; Claims 23 and 40 have been amended; Claims 1-22 remain cancelled. Response to Amendment The Office action is in response to Applicant’s communication filed July 20, 2026 in response to the Office action mailed April 20, 2026. Applicant’s remarks and amendments to the claims have been considered with the results set forth below. Response to Arguments Applicant's arguments filed July 20, 2026 have been fully considered but they are not persuasive. Regarding Rejections under §103 Applicant argues that d’Ascoli does not disclose: “wherein the static convolution kernel includes a plurality of parameters indexed by a relative spatial position and the static convolutional kernel employs a global receptive field”, and further asserts that Srinivas and Liu were cited to teach other features of claim 23. Applicant therefore concludes that the cited references do not render amended claim 23 obvious. Applicant’s argument is not persuasive because it addresses whether any one reference individually teaches the entirety of the newly added limitation, whereas the rejection is based on the combined teachings of Srinivas, d’Ascoli, and Liu. Srinivas teaches the global receptive field Srinivas teaches a Bottleneck Transformer architecture in which spatial convolutional layers are replaced with multi-head self-attention layers that implement: “global (all2all) self-attention over a 2D featuremap” (Srinivas, pg. 4, §3. Method) Srinivas further illustrates that the attention logits include the sum: “ q k T + q r T ” (Srinivas, pg. 4, §3. Method) where q , k , and r respectively represent query, key, and position encodings. Srinivas specifies that the position encodings are relative-distance encodings and that the content-content term q k T and content-position term q r T are combined in the same global self-attention operation. (Srinivas, pg. 4, Figure 4 and accompanying text). Accordingly, Srinivas teaches applying a position-dependent term across the global, all-to-all spatial receptive field of the two-dimensional feature map. d'Ascoli teaches a convolutional, relative-position term summed with adaptive attention d’Ascoli teaches positional self-attention in which the attention score includes an adaptive content term and a position-only term: A i j h = s o f t m a x   ( Q i h K j h T + v p o s h T r i j ) (d’Ascoli, pg. 4, Eq. 4) d’Ascoli states that v p o s h is a trainable embedding and that the relative positional encoding r i j depends only on the distance between pixels i and j , represented by the two-dimensional relative displacement δ i j . (d’Ascoli, pgs. 3-4, §2 Background, Eq. (4) and accompanying text). d’Ascoli further teaches that a positional self-attention layer using learnable relative positional encodings can express a convolutional layer by configuring respective attention heads for the positional offsets of a convolutional kernel. (d’Ascoli, pg. 4, § Self-attention as a generalized convolution, Eq. 5). Thus, d’Ascoli teaches an input-content-independent, relative-position component having convolutional character and combined with the adaptive, input-dependent content-attention component. d’Ascoli also teaches both claimed locations of the summation relative to SoftMax. Equation (4) adds the content and positional terms before applying SoftMax. Equation (7) separately applies SoftMax to the content and positional terms and then sums the resulting matrices: “ A i j h : = 1 - σ λ h   s o f t m a x Q i h K j h T + σ λ h s o f t m a x ( v p o s h T r i j ) ” (d’Ascoli, pg. 5, Eq. 7) Accordingly, d’Ascoli teaches applying the sum of a convolutional relative-position term and adaptive attention either prior to or subsequent to SoftMax normalization. Liu teaches a plurality of learned parameters indexed by relative spatial position Liu teaches adding a relative-position bias matrix B to the adaptive content-attention score: “ A t t e n t i o n Q ,   K ,   V = S o f t M a x ( Q K T d + B ) V ” (Liu, pg. 5, Eq. 4) Liu further teaches parameterizing the relative-position bias using a smaller learned matrix: “ B ^ ∈ R 2 M - 1 x ( 2 M - 1 ) ” (Liu, pg. 5, § Relative position bias), and expressly states that the values of B are taken from B ^ according to the relative positions of the respective patches. Liu therefore teaches a plurality of learned parameters indexed by relative spatial position. Liu is relied upon for the indexed relative-position parameterization, not for the global receptive-field portion of the limitation. Srinivas supplies the global, all-to-all receptive field. Combined teachings It would have been obvious to a person of ordinary skill in the art to implement the relative-position convolutional term taught by d’Ascoli in the global all-to-all attention operation taught by Srinivas and to parameterize that relative-position term using the learned relative-position parameter table taught by Liu. Srinivas expressly combines relative-position information with adaptive content attention in a global all-to-all attention layer. (Srinivas, pg. 4, Figure 4 and accompanying text). d'Ascoli teaches that the position-only term may be configured as a convolutional kernel and summed with adaptive content attention before or after SoftMax. (d’Ascoli, pgs. 3-5, Eq. 4, 5, and 7). Liu teaches the known implementation of the relative-position term as a plurality of learned values indexed according to relative spatial displacement. (Liu, pg. 5, §3.2, Eq. 4). The resulting combination processes the input data by applying the sum of: an adaptive, input-dependent content-attention matrix; and a static convolutional relative-position term comprising learned parameters selected according to relative spatial position, within Srinivas’s global all-to-all spatial receptive field. Accordingly, the combination of Srinivas, d’Ascoli, and Liu teaches or renders obvious: “wherein the static convolution kernel includes a plurality of parameters indexed by a relative spatial position and the static convolutional kernel employs a global receptive field”, as recited in amended claim 23. The amendment therefore does not overcome the rejection under 35 U.S.C. § 103. Claim 40 Independent claim 40 recites limitations materially corresponding to those of claim 23 in system form and includes the same newly added static-kernel limitation. Applicant relies on the same arguments presented for claim 23 and identifies no separate deficiency concerning the system limitations. For the reasons discussed above, Srinivas, d’Ascoli, and Liu likewise teach or render obvious the corresponding limitations of claim 40. The rejection of claim 40 under 35 U.S.C. § 103 is therefore maintained. Dependent claims Applicant generally states that the dependent claims may include independently patentable subject matter, but does not identify a specific dependent-claim limitation that is absent from the cited references or identify a specific error in the mappings provided in the Office action. A general assertion of patentability does not rebut the particular teachings and reasoning set forth for the dependent claims. Accordingly, the rejections of claims 24-39 and 41-42 are maintained for the same reasons previously stated and as clarified herein. Claim Objections Claims 23 and 40 are objected to because of the following informalities: In claims 23 and 40, the recitation “a statice convolution kernel” appears to contain a typographical error. Applicant is required to amend “statice” to “static”. Appropriate correction is required. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 23-28, and 31-42 are rejected under 35 U.S.C. 103 as being unpatentable over AravindSrinivas (Bottleneck Transformers for Visual Recognition) in view of Stephane d'Ascoli (ConViT: Improving Vision Transformers with Soft Convolutional Inductive Biases) and in further view of Ze Liu (Swin Transformer: Hierarchical Vision Transformer using Shifted Windows). Regarding claim 23, Srinivas in view of d’Ascoli, and further in view of Liu, teach a computer-implemented method comprising: “obtaining, by a computing system, input data;” – Srinivas teaches this limitation. Srinivas teaches a visual-recognition architecture that receives image data for image classification, object detection, and instance segmentation: “We present BoTNet, a conceptually simple yet powerful backbone architecture that incorporates self-attention for multiple computer vision tasks including image classification, object detection and instance segmentation.” (Srinivas, pg. 1, § Abstract) Thus, Srinivas teaches obtaining, by a computing system implementing BoTNet, input image data to be processed for a visual-recognition task. “providing the input data to a neural network, implemented by the computing system, that includes ” – Srinivas teaches this limitation in part, namely, providing the input data to a neural network implemented by the computing system and including an adaptive attention matrix. Srinivas teaches BoTNet as a neural-network backbone that combines convolutional processing and multi-head self-attention: “Our proposed architecture BoTNet is a hybrid model that uses both convolutions and self-attention.” (Srinivas, pg. 2, Figure 2 accompanying text) Srinivas further teaches replacing spatial convolution layers of a ResNet neural network with multi-head self-attention layers: “convolutions with MHSA layers Replace only the final three bottleneck blocks of a ResNet with BoT blocks without any other changes. Or in other words, take a ResNet and only replace the final three 3 x 3 convolutions with MHSA layers” (Srinivas, pg. 2, § 1. Introduction) The adaptive attention matrix is represented by the content-dependent attention term derived from the query and key feature projections of the input data. “wherein the neural network processes the input data by applying a sum of ” – Srinivas teaches this limitation in part. Srinivas teaches that its multi-head self-attention layer calculates attention logits as a sum of a content-content term and a relative-position term: “The attention logits are q k T + q r T where q ,   k ,   r represent query, key and position encodings respectively” (Srinivas, pg. 4, § 3. Method) Figure 4 depicts the summed attention logits q k T + q r T being provided to the SoftMax operation. Thus, Srinivas teaches summing the adaptive, content-dependent attention term q k T prior to performing SoftMax normalization. Srinivas does not characterize the relative-position term q r T as the claimed static convolution kernel. “wherein ” – Srinivas teaches this limitation in part. Srinivas teaches employing a global receptive field and expressly teaches that the multi-head self-attention layers implement global, all-to-all attention over the two-dimensional feature map: “Use global (all2all) self-attention to process and aggregate the information contained in the featuremaps captured by convolutions.” (Srinivas, pg. 2, § 1. Introduction) Srinivas further states: “BoTNet by design is simple: replace the final three spatial (3 x 3) convolutions in a ResNet with Multi-Head Self-Attention (MHSA) layers that implement global (all2all) self-attention over a 2D featuremap” (Srinivas, pg. 4, § 3. Method) Thus, Srinivas teaches an attention operation in which each spatial position may attend to the other spatial positions of the two-dimensional feature map, thereby employing a global receptive field. However, Srinivas does not expressly teach that a static convolution kernel comprising a plurality of parameters indexed by relative spatial position employs that global receptive field. “and receiving, by the computing system, a prediction that is based on the neural network processing the input data.” – Srinivas teaches this limitation. Srinivas teaches using the BoTNet neural network to produce predictions for image classification, object detection, and instance segmentation: “We present BoTNet, a conceptually simple yet powerful backbone architecture that incorporates self-attention for multiple computer vision tasks including image classification, object detection and instance segmentation.” (Srinivas, pg. 1, § Abstract) Srinivas further reports classification, bounding-box, and segmentation outputs produced from the neural network: “BoTNet achieves 44.4% Mask AP and 49.7% Box AP on the COCO Instance Segmentation benchmark …” (Srinivas, pg. 1, § Abstract) and: “… models that achieve a strong performance of 84.7% top-1 accuracy on the ImageNet benchmark …” (Srinivas, pg. 1, § Abstract) Thus, Srinivas teaches receiving a prediction based on the neural network’s processing of the input data. Srinivas does not expressly teach these limitations and/or portions of: “a static convolution kernel”, “applying a sum of a static convolution kernel and the adaptive attention matrix”, “wherein the static convolution kernel includes a plurality of parameters indexed by a relative spatial position”, and that the claimed convolutional kernel, as distinguished from the attention operation generally, “employs a global receptive field”. d'Ascoli, however, teaches these limitations and/or portions of: “a static convolution kernel” and “applying a sum of a static convolution kernel and the adaptive attention matrix” – d’Ascoli teaches positional self-attention having an adaptive content term Q i h K j h T and a position-only term v p o s h T r i j : “Each attention head uses a trainable embedding v p o s h ∈ R D p o s , and the relative positional encodings r i j ∈ R D p o s   only depend on the distance between pixels i and j, denoted by a two-dimensional vector δ i j .” (d’Ascoli, pg. 4, § 2 Background) The content term Q i h K j h T is adapted because it is calculated from the input-dependent query and key representations. In contrast, the position-only term v p o s h T r i j is determined by model parameters and the relative spatial displacement between positions, rather than the content of the input at those positions. d'Ascoli further teaches that the position-only attention component can implement a convolutional kernel: “Positional self-attention layers can be initialized as convolutional layers.” (d’Ascoli, pg. 4, Figure 3 caption) d'Ascoli further explains: “Thus, the PSA layer can achieve a strictly convolutional attention map by setting the centers of attention Δ h to each of the possible positional offsets of a N h   X   N h convolutional kernel” (d’Ascoli, pg. 4, § Self-attention as a generalized convolution) Thus, d’Ascoli teaches that the input-content-independent position term constitutes a static convolutional position component. d'Ascoli teaches applying the sum before SoftMax in Equation (4), because the content term Q i h K j h T and convolutional positional term v p o s h T r i j are summed inside the SoftMax operation. d'Ascoli also teaches the alternative post-SoftMax arrangement. d'Ascoli states: “To avoid this, GPSA layers sum the content and positional terms after the softmax, with their relative importances governed by a learnable gating parameter λ h .” (d’Ascoli, pg. 5, § Positional gating) d'Ascoli provides the corresponding formulation: “ A i j h : = 1 - σ λ h   s o f t m a x Q i h K j h T + σ λ h s o f t m a x ( v p o s h T r i j ) ” (d’Ascoli, pg. 5, Eq. 7) Thus, d’Ascoli teaches applying a sum of a static convolutional position term and an adaptive attention matrix either prior to or subsequent to performing SoftMax normalization. “the static convolutional kernel employs a global receptive field” – d’Ascoli teaches that the soft convolutional positional bias is not restricted to a hard local receptive field. d’Ascoli explains that restricting attention to only neighboring patches loses long-range information: “This led some authors to restrict the attention to a subset of patches around the query patch [19], at the cost of losing long-range information.” (d’Ascoli, pg. 4, § Adaptive attention span) d'Ascoli instead permits each attention head to escape its initialized locality: “We initialize the GPSA layers to mimic the locality of convolutional layers, then give each attention head the freedom to escape locality by adjusting a gating parameter regulating the attention paid to position versus content information.” (d’Ascoli, pg. 1, § Abstract) When d’Ascoli’s static convolutional position term is incorporated into Srinivas’s expressly global, all-to-all attention operation, the static convolutional kernel is evaluated within the global receptive field taught by Srinivas. d'Ascoli does not expressly disclose implementing the convolutional position term as an indexed table containing a plurality of learned relative-position parameters in the particular manner recited in claim 23. Thus, neither d’Ascoli nor Srinivas teach these remaining limitations and/or portions of: “wherein the static convolution kernel includes a plurality of parameters indexed by a relative spatial position” Liu, however, teaches these remaining limitations and/or portions of: “wherein the static convolution kernel includes a plurality of parameters indexed by a relative spatial position” – Liu teaches adding a relative-position bias B to the adaptive content-attention score: “ A t t e n t i o n Q ,   K ,   V = S o f t M a x ( Q K T d + B ) V ” (Liu, pg. 5, Eq. 4) Liu further teaches: “Since the relative position along each axis lies in the range - M + 1 ,     M - 1 , we parameterize a smaller-sized bias matrix B ^ ∈ R 2 M - 1 x ( 2 M - 1 ) , and values in B are taken from B ^ ” (Liu, pg. 5, § Relative position bias) Liu expressly characterizes this relative-position bias as learned: “The learnt relative position bias in pre-training can be also used to initialize a model for fine-tuning with a different window size through bi-cubic interpolation” (Liu, pg. 5, § Relative position bias) Thus, Liu teaches a matrix containing a plurality of learned, input-independent parameters, wherein the parameter used for a pair of spatial positions is selected or indexed according to the relative spatial displacement between those positions. In the proposed combination, Liu’s learned relative-position parameter table supplies the parameters of the convolutional position term taught by d’Ascoli, while Srinivas supplies the global, all-to-all receptive field. It would have been obvious to a person of ordinary skill in the art before the effective filing date to implement the relative-position component of Srinivas’s global multi-head self-attention using d’Ascoli’s convolutional position-only attention term. Srinivas teaches using global self-attention to model non-local spatial dependencies. (Srinivas, pg. 1, § 1. Introduction). d'Ascoli teaches adding a soft convolutional inductive bias to self-attention to improve sample efficiency while allowing the attention heads to escape locality when appropriate. (d’Ascoli, pgs. 1-2, § Abstract and § Contribution). A person of ordinary skill therefore would have been motivated to use d’Ascoli’s convolutional position-only term in Srinivas’s global attention mechanisms to obtain the positional and sample-efficiency benefits of convolutional bias while retaining global dependency modeling, predictably resulting in global adaptive attention supplemented by an input-independent convolutional position term. It would have been further obvious to parameterize d’Ascoli’s convolutional relative-position term using Liu’s learned relative-position bias table. d'Ascoli teaches that the convolutional position term depends on the relative displacement between spatial positions. (d’Ascoli, pg. 3, Eq. 4). Liu teaches storing learned position-bias values in a compact matrix and selecting the applicable value according to the relative displacement between query and key positions. (Liu, pg. 5, § 3.2). A person of ordinary skill therefore would have been motivated to use Liu’s learned relative-position table to provide an efficient, trainable, and position-indexed implementation of d’Ascoli’s convolutional position term, predictably resulting in selection of a static parameter according to the relative spatial position of the corresponding input features. Accordingly, the combination of Srinivas, d’Ascoli, and Liu renders claim 23 obvious under 35 U.S.C. § 103. Regarding claim 24, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 23, wherein “the input data comprises an input tensor having one or more dimensions.” – Srinivas teaches this limitation. Shrinivas teaches input image data provided to a visual-recognition neural network. In particular, Srinivas teaches processing images as feature maps within a neural-network architecture for visual-recognition tasks including image classification, object detection, and instance segmentation: “We present BoTNet, a conceptually simple yet powerful backbone architecture that incorporates self-attention for multiple computer vision tasks including image classification, object detection and instance segmentation.” (Shrinivas, pg. 1, § Abstract) “BoTNet by design is simple: replace the final three spatial (3x3) convolutions in a ResNet with Multi-Head Self-Attention (MHSA) layers that implement global (all2all) self-attention over a 2D featuremap” (Shrinivas, pg. 4, § 3. Method) Srinivas further teaches that its self-attention layer operates on a 2D featuremap, stating: “(MHSA) layer used in the BoT block … attention is all2all performed on a 2D featuremap” (Srinivas, pg. 4, § 3. Method) Srinivas also teaches the self-attention layer using tensors / matrices over spatial dimensions, including: “HxWxd” (Shrinivas, pg. 4, § 3. Method) It would have been obvious to a person of ordinary skill in the art to apply the hybrid convolution/attention neural-network method of Srinivas, as modified by d’Ascoli and Liu for the reasons set forth with respect to claim 23, to input data comprising an input tensor having one or more dimensions, because Shrinivas itself teaches processing image data as a 2D featuremap with dimensions H x W x d in its MHSA layer for visual-recognition tasks. A POSITA would have understood such image/feature-map data to be a tensor representation and would have found it obvious to characterize the input in that conventional form when implementing the known visual-recognition network on a computing system. Regarding claim 25, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 23, wherein: “the neural network is a machine-learned convolutional network that includes at least a convolutional stage and an attention stage that comprises a relative attention mechanism that is configured to apply ” – Srinivas teaches this limitation in part. Srinivas teaches a visual-recognition backbone architecture built from a ResNet backbone and therefore is a machine-learned convolutional network architecture. Shrinivas states: “Our proposed architecture BoTNet is a hybrid model that uses both convolutions and self-attention.” (Shrinivas, pg. 2, Figure. 2 text) Shrinivas further teaches ResNet backbone in which earlier portions remain convolutional and later bottleneck blocks are replaced with attention blocks, stating: “Replace only the final three bottleneck blocks of a ResNet with BoT blocks without any other changes. Or in other words, take a ResNet and only replace the final three 3 x 3 convolutions with MHSA layers” (Shrinivas, pg. 2, § 1. Introduction) Shrinivas also explains that this design uses convolutions to efficiently learn feature maps and then uses self-attention to process and aggregate the information, i.e., a hybrid design using both convolutional processing and attention processing. “Our proposed architecture BoTNet is a hybrid model that uses both convolutions and self-attention.” (Shrinivas, pg. 2, Figure 2 Text) It would have been obvious to a POSITA to use a neural network at least a convolutional stage and an attention stage because Srinivas expressly teaches a hybrid architecture that uses both convolutions and self-attention, with convolutional potions of the backbone and later attention blocks, to improve visual-recognition performance while leveraging the known complementary benefits of convolution and attention. Regarding claim 26, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 23, wherein “the neural network includes at least an S0 stage, an S1 stage, an S2 stage, an S3 stage, and an S4 stage.” – Srinivas teaches this limitation. Srinivas teaches a ResNet-based backbone having an initial convolutional stem followed by four successive convolutional stage groups. In particular, Srinivas’s architecture tables shows an initial c1 convolutional stem, followed by c2, c3, c4, and c5 stages: PNG media_image1.png 450 474 media_image1.png Greyscale (Srinivas, pg. 4, § 3. Method) Srinivas further states that: “A ResNet typically has 4 stages (or blockgroups) commonly referred to as [c2, c3, c4, c5]” (Srinivas, pg. 4, § 3. Method) It would have been obvious to a person of ordinary skill in the art to implement the neural network with at least five successive stages because Shrinivas expressly teaches that staged convolutional backbones uses an initial stem and four subsequent stage groups for progressive visual-feature extraction and processing. A POSITA would have understood the claimed S0-S4 stages to correspond to such successive stages, regardless of the particular naming convention. Regarding claim 27, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 26, wherein “the S0 stage comprises a ” – Srinivas teaches this limitation in part. Srinivas teaches an initial convolutional stem stage c1 comprising a 7x7 convolution with 64 channels and stride 2, followed by the later c2-c5-feature-processing stages. (Srinivas, pg. 4, Table 1, § 3. Method). Srinivas does not expressly teach that the initial convolutional stem comprises two convolutional layers. However, Srinivas teaches that conventional visual backbones commonly use multiple stacked convolutional layers and that stacking convolutional layers improves feature-processing performance: “Most landmark backbone architectures … use multiple layers of 3 x 3 convolutions.” (Srinivas, pg. 1, § 1. Introduction) and: “Although stacking more layers indeed improves the performance of these backbones” (Srinivas, pg. 1, § 1. Introduction) Srinivas does not teach the remaining portion of the limitation: “” However, the recited “” convolutional stem would have been obvious to a person of ordinary skill in the art before the effective filing date to implement Srinivas’s initial convolutional stem using two successive convolutional layers instead of a single convolutional layer. Srinivas teaches both the use of an initial convolutional stage to extract and spatially downsample image features and the known performance benefit of stacking convolutional layers. A person of ordinary skill therefore would have selected two successive convolutional layers as a predictable implementation of the stem to perform staged local-feature extraction before the later feature-processing and attention stages, with the number of stem layers selected according to the desired balance of feature extraction, computational cost, and model depth. Accordingly, Srinivas’s initial convolutional stem, modified to include two successive convolutional layers in accordance with Srinivas’s teaching of stacked convolutional processing, renders the additional limitation of claim 27 obvious. Regarding claim 28, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 26, wherein “S1 stage comprises one or more convolutional blocks with squeeze excitation.” – Srinivas teaches this limitation. Srinivas teaches a staged ReNet-based backbone having successive convolutional block groups after the initial stem. In particular, Srinivas teaches that: “A ResNet typically has 4 stages (or blockgroups) commonly referred to as [c2, c3, c4, c5]” (Srinivas, pg. 4, § 3. Method) And Srinivas’s architecture table shows that those stages are made of ResNet blocks. See Srinivas, pg. 4, § 3. Method. Srinivas further teaches “squeeze excitation” in the context of convolutional blocks. In particular, Srinivas states: “Other aspects of improved training of backbone architectures for image classification has been the use of Squeeze-Excitation (SE) blocks” (Srinivas, pg. 9, § 4.8.4 Effect of SE blocks, SiLU and lower weight decay) Srinivas also teaches, in direct comparison to its own architecture: “it is possible to get visible gains on top of R50 when placing SE blocks in all bottleneck blocks throughout the ResNet and not just in c5.” (Srinivas, pg. 20, § A.6. Comparison to Squeeze-Excite) Srinivas teaches that the first post-stem stage of the ResNet backbone is a convolutional stage comprising convolutional blocks, and also teaches that SE blocks can be placed “in all bottleneck blocks throughout the ResNet and not just in c5”. A POSITA would have understood that applying SE to the early convolutional stage blocks, including the first post-steam stage corresponding to S1, was one of the straightforward placements encompassed by Srinivas’s teaching of using SE throughout the ResNet. Doing so would have been a predictable variation for improving feature recalibration in the convolutional blocks while retraining the hybrid Srinivas architecture. It would have been obvious to a person of ordinary skill in the art to configure the S1 stage to comprise one or more convolutional blocks with squeeze excitation because Shrinivas expressly teaches both ingredients: a first convolutional stage composed of convolutional blocks, and the use of SE blocks throughout the ResNet bottleneck blocks rather than only in c5. A POSITA would have recognized that applying SE to the first convolutional stage bocks was a routine and beneficial implementation choice within that teachings. Regarding claim 31, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 26, wherein “a number of channels is doubled for at least one of the S1 stage, the S2 stage, the S3 stage, or the S4 stage.” – Srinivas teaches this limitation. Srinivas teaches a staged backbone architecture in which the channel count increases across successive stages. In particular, Srinivas’s architecture table shows stage outputs progressing, i.e., c2 … 1x1, 256, … , c5 … 1x1, 2048: PNG media_image1.png 450 474 media_image1.png Greyscale (Srinivas, pg. 4, § 3. Method) Srinivas further teaches that these successive backbone stages / block groups: “A ResNet typically has 4 stages (or blockgroups) commonly referred to as [c2, c3, c4, c5]” (Srinivas, pg. 4, § 3. Method) It would have been obvious to a person of ordinary skill in the art to configure at least one of the claimed S1-S4 stages so that the number of channels is doubled because Srinivas expressly teaches staged visual-recognition backbones in which the output channel count increases by doubling across successive stages to support progressively richer feature representations. A POSITA would have understood this as a conventional and beneficial design choice in staged convolutional / hybrid backbones. Regarding claim 32, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 26, wherein “a width of the S0 stage is less than or equal to a width of the S1 stage.” – Srinivas teaches this limitation. Srinivas teaches an initial stem stage followed by a first subsequent stage in a staged backbone architecture. In particular, Srinivas’s architecture table shows an initial C1 stage having: “7x7, 64, stride 2” (Srinivas, pg. 4, § 3. Method, table text) Srinivas further shows the next stage, c2, as having blocks with width at least: “1x1, 64” (Shrinivas, pg. 4, § 3. Method, table text) Thus, Srinivas teaches that the width of the initial stage is 64, and the width of the next stage is 64 or greater, i.e., the width of the initial stage is less than or equal to the width to the width of the next stage. It would have been obvious to a person of ordinary skill in the art to configure the width of the initial stage to be less than or equal to the width of the following stage because Shrinivas expressly teaches a staged backbone in which the initial stem does not exceed the width of the next stage, consistent with conventional visual-network design favoring nondecreasing feature width across successive stages. Regarding claim 33, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 26, wherein “each of the S0 stage, the S1 stage, and the S4 stage comprises two blocks, and wherein the S2 stage and the S3 stage each comprise greater than two blocks.” – Srinivas does not teach this limitation. Liu, however, teaches this limitation. Liu teaches a hierarchical staged architecture having a defined number of successive processing blocks in each stage. Liu expressly teaches the Swin-T architecture having: “layer numbers = {2; 2; 6; 2}” (Liu, pg. 5, § 3.3. Architecture Variants) Liu’s Figure 3 likewise depicts four successive stages respectively comprising two blocks, two blocks, six blocks, and two blocks: PNG media_image2.png 251 693 media_image2.png Greyscale (Liu, pg. 4, Figure 3) Liu does not expressly label its architecture as five stages having the distribution recited in claim 33. However, Liu’s six-block third stage comprises six consecutive processing blocks. It would have been obvious to a person of ordinary skill in the art before the effective filing date, when applying Liu’s disclosed block-depth allocation to the five-stage architecture taught by Srinivas, to perform Liu’s six consecutive third-stage blocks into two consecutive stages containing there blocks each. The resulting five-stage distribution would be: Claimed stage Liu-based block allocation S0 2 blocks S1 2 blocks S2 First 3 blocks of Liu’s six-block third stage S3 Remaining 3 blocks of Liu’s six-block third stage S4 2 blocks Thus, the resulting distribution is: 2 ,   2 ,   3 ,   3 ,   2 wherein S0, S1, and S4 each comprise two blocks and S2 and S3 each comprise greater than two blocks. A person of ordinary skill would have been motivated to divide Liu’s six-block intermediate processing group into two successive three-block stages when implementing Liu’s block-depth allocation in Srinivas’s five-stage organization to provide a modular five-stage architecture while retaining Liu’s disclosed total number, order, and processing function of the blocks. Such partitioning would not require changing the operations performed by Liu’s six consecutive blocks and would predictably preserve the processing capacity of Liu’s intermediate stage. Further, claim 33 does not require a change in spatial resolution, channel dimension, or block type at every boundary between the recited stages. Therefore, organizing Liu’s six consecutive intermediate blocks as two successive groups of three blocks satisfies the claimed stage and block-count arrangement. Regarding claim 34, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 23, wherein “a special resolution gradually decreases over two or more network stages of the neural network.” – Srinivas teaches this limitation. Srinivas teaches a staged backbone architecture in which spatial resolution decreases across successive stages. In particular, Srinivas’s architecture table shows stage outputs progressing: PNG media_image1.png 450 474 media_image1.png Greyscale (Srinivas, pg. 4, § 3. Method) Srinivas further teaches that these successive backbone stages / block groups: “A ResNet typically has 4 stages (or blockgroups) commonly referred to as [c2, c3, c4, c5]” (Srinivas, pg. 4, § 3. Method) It would have been obvious to a person of ordinary skill in the art to configure the neural network so that spatial resolution gradually decreases over two or more network stages because Srinivas expressly teaches a staged visual-recognition backbone with progressively reduced spatial resolutions across successive stages, which was a conventional and beneficial design choice for extracting increasingly abstract features while controlling computation. Regarding claim 35, Srinivas in view of d’Ascoli and in further view of Liu, the computer-implemented method of claim 23, wherein “the sum of the static convolution kernel and the adaptive attention matrix is applied to the input data prior to performing the SoftMax normalization.” – Srinivas does not teach this limitation. d'Ascoli, however, teaches this limitation. d’Ascoli defines positional self-attention as: “ A i j h = s o f t m a x   ( Q i h K j h T + v p o s h T r i j ) ” (d’Ascoli, pg. 4, Eq. 4) d'Ascoli teaches that Q i h K j h T is the content-dependent attention term and that v p o s h T r i j is a position-dependent term based on relative spatial displacement between positions i and j . d'Ascoli further teaches that the positional term may express a convolutional layer according to convolutional-kernel offsets. Thus, d’Ascoli teaches summing the static convolutional term and adaptive attention term prior to SoftMax normalization. (d’Ascoli, pgs. 3-4, Eq. (4)-(5), and accompanying text) Thus: Q i h K j h T corresponds to the adaptive attention matrix; v p o s h T r i j corresponds to the static convolutional position term; and Equation (4) sums those terms inside the argument of SoftMax. Accordingly, d’Ascoli teaches applying the sum of the static convolution-kernel and the adaptive attention matrix prior to performing SoftMax normalization. For the reasons discussed with respect to claim 23, it would have been obvious to implement d’Ascoli’s convolutional position term in Srinivas’s global attention mechanism using d’Ascoli’s expressly disclosed pre-SoftMax formulation, predictably resulting in the arrangement recited in claim 35. Regarding claim 36, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 23, wherein “the sum of the static convolution kernel and the adaptive attention matrix is applied to the input data sequence to performing the SoftMax normalization by the relative attention mechanism.” – Srinivas does not teach this limitation. d’Ascoli, however, teaches this limitation. d’Ascoli teaches that gated positional self-attention (GPSA) separately applies SoftMax normalization to the adaptive content-attention term and the convolutional positional term, and thereafter sums the normalized terms: “GPSA layers sum the content and positional terms after the softmax, with their relative importance governed by a learnable gating parameter λ h (one for each attention head).” (d’Ascoli, pg. 5, § Positional gating) d'Ascoli provides the corresponding formulation: “ A i j h : = 1 - σ λ h   s o f t m a x Q i h K j h T + σ λ h s o f t m a x ( v p o s h T r i j ) ” (d’Ascoli, pg. 5, § Positional gating, Eq. 7) d'Ascoli further applies the resulting attention matrix to the input representation X : “ G P S A h X ≔ n o r m a l i z e A h X W v a l h ” (d’Ascoli, pg. 5, Eq. 6) Thus, d’Ascoli separately performs SoftMax normalization on the adaptive content-attention term Q i h K j h T and the convolutional positional term v p o s h T r i j , sums the normalized terms, and subsequently applies resulting attention matrix to the input data.For the reasons discussed with respect to claim 23, it would have been obvious to implement d’Ascoli’s convolutional position term in Srinivas’s global attention mechanism using d’Ascoli’s expressly disclosed post-SoftMax GPSA formulation, predictably resulting in the arrangement recited in claim 36. Regarding claim 37, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 23, wherein “the prediction comprises a computer vision output.” – Srinivas teaches this limitation. Srinivas teaches a visual-recognition backbone architecture for computer-vision. In particular, Srinivas states: “We present BoTNet, a conceptually simple yet powerful backbone architecture that incorporates self-attention for multiple computer vision tasks including image classification, object detection and instance segmentation.” (Srinivas, pg. 1, § Abstract) Thus, Srinivas teaches that the network output / prediction is a computer vision output, such as a classification result, detection result, or instance-segmentation result. It would have been obvious to a person of skill in the art to implement the claimed prediction as a computer vision output because Srinivas expressly teaches using the neural-network backbone for computer-vision tasks including image classification, object detection, and instance segmentation, each of which yields a computer vision output. Regarding claim 38, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 23, wherein “the prediction comprises a classification output.” – Srinivas teaches this limitation. Shrinivas teaches a visual-recognition backbone architecture for computer-vision tasks including image classification. In particular, Srinivas states: “We present BoTNet, a conceptually simple yet powerful backbone architecture that incorporates self-attention for multiple computer vision tasks including image classification, object detection and instance segmentation.” (Srinivas, pg. 1, § Abstract) Shrinivas further explains that It presents an adaptation of the BoTNeT design “for image classification” and reports models that achieve strong top-1 accuracy on the ImageNet benchmark which is a classification output. It would have been obvious to a person of ordinary skill in the art to implement the claimed prediction as a classification output because Shrinivas expressly teaches using the neural-network backbone for image classification, which inherently yields a classification output. Regarding claim 39, Srinivas in view of d’Ascoli and in further view of Liu, teach the computer-implemented method of claim 23, wherein “one or more convolutional stages of the neural network are sequentially prior to one or more attention stages of the neural network.” – Srinivas teaches this limitation. Srinivas teaches a hybrid neural-network backbone in which convolutional portions occur earlier in the network and attention portions occur later. In particular, Srinivas states: “Use convolutions to efficiently learn abstract and low resolution featuremaps from large images; (2) Use global (all2all) self-attention to process and aggregate the information contained in the featuremaps captured by convolutions.” (Srinivas, pg. 2, § Introduction) Shrinivas further teaches: “Replace only the final three bottleneck blocks of a ResNet with BoT blocks without any other changes. Or in other words, take a ResNet and only replace the final three 3 x 3 convolutions with MHSA layers” (Srinivas, pg. 2, § Introduction) It would have been obvious to a person of ordinary skill in the art to arrange one or more convolutional stages sequentially prior to one or more attention stages because Srinivas expressly teaches using convolutions earlier to efficiently extract feature maps and using self-attention later to process and aggregate those features, reflecting the known complementary advantages of early convolutional feature extraction and later attention-based relational modeling. Regarding claim 40, Srinivas in view of d’Ascoli, and further in view of Liu, teach a computing system comprising: “one or more processors; and one or more non-transitory computer-readable media that store instructions that when executed by the one or more processors, cause the computer system to perform operations comprising:” – Srinivas teaches a computer-implemented neural-network architecture executed using computing hardware, including TPU-v3 hardware, for performing image classification, object detection, and instance segmentation. (Srinivas, pg. 1 and 4). It would have been obvious to store the instructions implementing Srinivas’s neural-network operations on non-transitory computer-readable media for execution by one or more processors, as a conventional implementation of the disclosed computer-executed neural network. Claim 40 otherwise recites operations materially corresponding to those of claim 23 in computing-system form. Accordingly, the findings, mappings, and rationales set forth above with respect to claim 23 are incorporated herein. In particular: Srinivas teaches obtaining input data, providing the input data to a neural network, processing the input data using global all-to-all self-attention, and receiving a prediction based on that processing. Srinivas further teaches summing an adaptive content-attention term and a relative-position term within the attention operation. (Srinivas, pgs. 1-4, Fig. 4) d'Ascoli teaches that the relative-position term may constitute a static convolutional position term and may be summed with adaptive content attention either before or after SoftMax normalization. (d’Ascoli, pgs. 3-5, Eq. (4), (5), and (7)). Liu teaches implementing the relative-position term using a plurality of learned parameters selected according to relative spatial position. (Liu, pg. 5, § 3.2, Eq. (4)). Thus, for the reasons discussed with respect to claim 23, the combined teachings render obvious instructions that cause the computing system to: “obtaining input data [and] providing the input data to a neural network that includes a statice convolution kernel and an adaptive attention matrix, wherein the neural network processes the input data by applying a sum of a static convolution kernel and the adaptive attention matrix, either prior to or subsequent to performing a Softmax normalization, to the input data, wherein the static convolution kernel includes a plurality of parameters indexed by a relative spatial position and the static convolutional kernel employs a global receptive field;” and to receive a prediction based on the neural network processing the input data. Accordingly, Srinivas in view of d’Ascoli, and further in view of Liu, renders claim 40 obvious under 35 U.S.C. § 103. Regarding claim 41 Claim 41 is rejected under 35 U.S.C. § 103 as being unpatentable over Srinivas in view of d’Ascoli and further in view of Liu. The rejection of claim 41 relies on the same findings, mappings, and citations set forth above with respect to claim 25, which is the corresponding method claim, except that claim 41 is directed to the computing system of claim 40. Srinivas teaches a computing system / neural network backbone for visual-recognition processing using a hybrid model that uses both convolutions and self-attention and that retains convolutional portions of a ResNet backbone while replacing later portions with MHSA / attention blocks, thereby teaching a machine-learned convolutional network including at least a convolutional stage and an attention stage; Srinivas further teaches a relative attention mechanism through its use of relative distance encodings and a summed attention formulation q K T + q R T ; d’Ascoli teaches that the static term within the attention mechanism can be convolutional in nature by introducing a soft convolutional inductive bias and a self-attention layer that can be initialized as a convolutional layer; and Liu teaches applying a static term together with an adaptive attention term prior to SoftMax via SoftMax( Q K T d + B ) V . Accordingly, for the reasons set forth above with respect to claim 25, the combined teachings render obvious the system of claim 41 in which the neural network is a machine-learned convolutional network that includes at least a convolutional stage and an attention stage that comprises a relative attention mechanism configured to apply the sum of the static convolution kernel and the adaptive attention matrix, either prior to or subsequent to performing SoftMax normalization, to the input data. Regarding claim 42 Claim 42 is rejected under 35 U.S.C. § 103 as being unpatentable over Srinivas in view of d’Ascoli and further in view of Liu. The rejection of claim 42 relies on the same findings, mappings, and citations set forth above with respect to claim 26, which is the corresponding method claim, except that claim 42 is directed to the computing system of claim 40. Srinivas teaches a staged backbone architecture having an initial convolutional stem followed by four successive stages / block groups, namely an initial c1 stem and subsequent c2, c3, c4, c5 stages, and further teaches that “ResNet typically has 4 stages (or blockgroups) commonly referred to as [c2,c3,c4,c5]”, thereby teaching a neural network having at least five successive stages corresponding to the claimed S0, S1, S2, S3, and S4 stages; accordingly, in view of the reasons set forth with respect to claim 40, the combined teachings render obvious the system of claim 42 in which the neural network includes at least an S0 stage, an S1 stage, an S2 stage, an S3 stage, and an S4 stage. Claims 29 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Srinivas in view of d'Ascoli in further view of Liu in further view of Mingxing Tan (EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks) and further in view of Mark Sandler (MobileNetV2: Inverted Residuals and Linear Bottlenecks). Regarding claim 29, Srinivas in view of d’Ascoli in further view of Liu in further view of Tan, and further in view of Sandler, teach the computer-implemented method of claim 28, wherein “the one or more convolutional blocks of the S1 stage comprise ” – Srinivas teaches this limitation in part. Srinivas teaches a staged convolutional backbone with a post-stem convolutional stage comprising convolutional blocks, thus teaching the general S1-stage convolution-block context. Srinivas architecture table shows an initial c1 stem followed by c2, and the c2 entry is a stack of ResNet bottleneck blocks, i.e. a post-stem convolutional stage comprising convolutional blocks: PNG media_image1.png 450 474 media_image1.png Greyscale (Srinivas, pg. 4, § 3. Method) Srinivas also states that: “A ResNet typically has 4 stages (or blockgroups) commonly referred to as [c2, c3, c4, c5]” (Srinivas, pg. 4, § 3. Method) This disclosure supports c2 as the first post-stem stage. Srinivas does not teach these portions of the limitation: “” Tan, however, teaches these portions of the limitation: “” – Tan teaches that its main building block is mobile inverted bottleneck MBConv and that EfficientNet-B0 is organized by stages, including an early post-stem MBConv stage. Specifically, Tan states that its: “main building block is mobile inverted bottleneck MBConv … to which we also add squeeze-and-excitation optimization” (Tan, § 4. EfficientNet Architecture) And Tan’s stage table shows an initial stem followed by “Stage 2 MBConv1, k3x3” and later MBConv stages: “ PNG media_image3.png 358 730 media_image3.png Greyscale ” (Tan, § 4. EfficientNet Architecture, Table 1) Tan does not teach the remaining portion of the limitation: “” Sandler, however, teaches this remaining portion of the limitation: “” – Sandler teaches: “a novel layer module: the inverted residual with linear bottleneck.” (Sandler, § 1. Introduction) And Sandler further teaches that this module: “takes as an input a low-dimensional compressed representation which is first expanded to high dimension” (Sandler, § 1. Introduction) And that: “Features are subsequently projected back to a low-dimensional representation with a linear convolution.” (Sandler, § 1. Introduction) Sandler also teaches that the architecture: “is based on an inverted residual structure where the shortcut connections are between the thin bottleneck layers.” (Sandler, § Abstract) “The intermediate expansion layer uses lightweight depthwise convolutions to filter features as a source of non-linearity” (Sandler, § Abstract) It would have been obvious to a person of ordinary skill in the art to use MBConv blocks in the first post-stem convolutional stage of the Srinivas-based staged network because Tan expressly teaches MBConv as its main building block, including in an early post-stem stage, and Sandler teaches the defining inverted-residual / linear-bottleneck behavior of expansion to a higher-dimensional intermediate representation followed by projection back to a bottleneck representation. A POSITA would have recognized that substituting the known MBConv block structure into the known first post-stem convolutional stage of a staged vision backbone was predictable variation for improving efficiency while preserving representational capacity. Regarding claim 30, Srinivas in view of d’Ascoli in further view of Liu in further view of Tan, and further in view of Sandler, teach the computer-implemented method of claim 26, wherein “each of the S2 stage, S3 stage, or S4 stage comprising a convolutional stage comprise a mobile inverted bottleneck convolution (MBConv) block.” – Srinivas teaches this limitation in part. Srinivas teaches the staged convolutional-network context, including later stages corresponding to S2, S3, and S4. In particular, Srinivas teaches that: “A ResNet typically has 4 stages (or blockgroups) commonly referred to” (Srinivas, pg. 4, § 3. Method) And Srinivas teaches that these stacks: “consist of multiple bottleneck blocks with residual connections” (Srinivas, pg. 4, § 3. Method) Srinivas does not teach this portion of the limitation: “” Tan, however, teaches this remaining portion of the limitation: “” – Tan teaches that its: “main building block is mobile inverted bottleneck MBConv” (Tan, § 4. EfficientNet Architecture) And Tans stage table shows multiple successive later stages using MBConv blocks, including: “ PNG media_image4.png 161 356 media_image4.png Greyscale ” (Tan, § 4. EfficientNet Architecture, Table 1) It would have been obvious to a person of ordinary skill in the art to use MBConv blocks in later convolutional stages of the BoTNet-based staged network because Tan expressly teaches MBConv as the main building block across multiple successive later stages, and a POSITA would have recognized that using those known MBConv blocks in the later convolutional stages of a staged hybrid backbone was a predictable way to implement efficient staged convolutional processing. Accordingly, the combined teachings of Srinivas and Tan render obvious the claimed limitation. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Paul Coleman whose telephone number is (571)272-4687. The examiner can normally be reached Mon-Fri. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PAUL COLEMAN/ Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Jul 19, 2023
Application Filed
Aug 16, 2023
Response after Non-Final Action
Apr 20, 2026
Non-Final Rejection mailed — §103
Jul 20, 2026
Response Filed
Aug 06, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748961
NEUROMORPHIC COMPUTING DEVICE WITH THREE-DIMENSIONAL MEMORY
4y 4m to grant Granted Sep 29, 2026
Patent 12731010
MULTIVARIABLE TIME-SERIES FEATURE EXTRACTION
3y 7m to grant Granted Sep 08, 2026
Patent 12711416
MODEL MODIFICATION AND DEPLOYMENT
5y 3m to grant Granted Aug 18, 2026
Patent 12688400
COMPUTATIONAL NEURAL NETWORK APPARATUS, CARD, METHOD, AND READABLE STORAGE MEDIUM
3y 6m to grant Granted Jul 21, 2026
Patent 12665745
MACHINE LEARNING/ARTIFICIAL INTELLIGENCE (ML/AI) SYSTEM WITH PROTECTED NEURAL NETWORKS
3y 5m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+47.4%)
3y 8m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 26 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month