Prosecution Insights
Last updated: August 18, 2026
Application No. 18/305,676

NON-LINEAR MULTI-DIMENSIONAL COST FUNCTION FOR ARTIFICIAL INTELLIGENCE INFERENCE

Final Rejection §102§103
Filed
Apr 24, 2023
Examiner
BOSTWICK, SIDNEY VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Samsung Electronics Co., Ltd.
OA Round
2 (Final)
52%
Grant Probability
Moderate
3-4
OA Rounds
1y 1m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 52% of resolved cases
52%
Career Allowance Rate
76 granted / 147 resolved
-3.3% vs TC avg
Strong +37% interview lift
Without
With
+36.9%
Interview Lift
resolved cases with interview
Typical timeline
4y 5m
Avg Prosecution
41 currently pending
Career history
214
Total Applications
across all art units

Statute-Specific Performance

§101
25.3%
-14.7% vs TC avg
§103
45.2%
+5.2% vs TC avg
§102
4.9%
-35.1% vs TC avg
§112
24.3%
-15.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 147 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Remarks This Office Action is responsive to Applicants' Amendment filed on April 27, 2026, in which claims 1, 2, 13, 14, and 20 are currently amended. Claim 10 is canceled. Claims 1-9 and 11-20 are currently pending. Response to Arguments The rejections to claims 1-9 and 11-20 under 35 U.S.C. § 112(b) are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claims 1-9 and 11-20 under 35 U.S.C. 101 based on amendment have been considered and are persuasive. The rejections to claims 1-9 and 11-20 under 35 U.S.C. § 101 are hereby withdrawn, as necessitated by applicant's amendments and remarks made to the rejections. Applicant’s arguments with respect to rejection of claims 1-9 and 11-20 under 35 U.S.C. 102/103 based on amendment have been considered, however, are not persuasive. With respect to Applicant's arguments on p. 10 of the Remarks submitted 4/27/2026 that "Gondimalla does not disclose "linearization"", Examiner respectfully disagrees. Gondimalla's transfer metric is linearized because it reduces the multidimensional, piecewise CNN memory-traffic problem to a scalar additive transfer count X. The provided exemplary embodiment from the instant disclosure is explicitly non-limiting such that neither the instant specification nor the instant claim require the linearized metric to be approximate, stochastic, or produced by a particular linearization technique. Gondimalla explicitly defines X as "the optimal number of transfers," initializes X as |Li|+|Lj|, and updates it by the additive recurrence OP[i,j].X = OP[i,popt]X+OP[popt,j].X (Eqn. 5). That is a linearized scalar representation of the performance parameter, total off-chip traffic. It linearizes the otherwise multidimensional CNN memory-footprint problem into additive counts of feature-map transfers across partition boundaries. With respect to Applicant's arguments on p. 11 of the Remarks submitted 4/27/2026 that "Gondimalla does not disclose a "Non-Linear Constraint" on a performance parameter", Examiner respectfully disagrees. With respect to Applicant's argument "Gondimalla does not disclose any non-linear cost function. Instead, Gondimalla merely discusses minimizing a single metric (off-chip transfer count X), applying a cache-capacity feasibility condition, and computes transfer counts accordingly. See, e.g., page 11-12 of Gondimalla. However, a cache-capacity check (i.e., whether a span fits in cache) is a binary feasibility constraint - not a non-linear cost function", this argument improperly isolates X from the DP feasibility logic that determines how X is computed. Gondimalla is not merely "minimizing a single metric" in the abstract. It discloses a constrained dynamic-programming cost formulation. The Gondimalla paper says Occam uses a "dynamic programming algorithm" to "optimally partition a given CNN" and "guarantee the least off-chip traffic" for "a given on-chip capacity" ([p. 21]). That is the claimed performance parameter, X, optimized under a capacity constraint. Gondimalla explicitly discloses a dynamic-programming cost formulation in which the off-chip transfer metric X is initialized as |Li|+|Lj| only when the span footprint satisfies the cache capacity inequality |DC(i,j)|+E|Wk|<C. Otherwise, the metric is updated by partitioning the span and minimizing the sum of subspan transfer metric. Because the inequality controls the branch of the DP cost calculation, the cost formulation is piecewise/discontinuous and therefore non-linear. The claimed "non-linear constraint" is derived from the same cost formulation as the thresholded condition on the span-footprint term that determined whether the base cost or partitioned cost applies. In other words the piecewise formulation of Gondimalla's cost function is OP[i,j].X={|Li|+|Lj|, if |DC(i,j|+E|Wk|<C, otherwise min_p(OP[i,p].X+OP[p,j].X}. The disclosed DP formulation is objectively a thresholded, piecewise/discontinuous optimization over multidimensional CNN tensor footprints. It is non-linear because the constraint/inequality produces a discontinuous feasibility boundary. A binary feasibility condition is still a constraint and more importantly, here it controls the branch of the dynamic programming recurrence. that makes the overall cost formulation piecewise/discontinuous/nonlinear. This reasoning also addresses Applicant's arguments on p. 11 of the Remarks submitted 4/27/2026 that "Gondimalla Does Not Disclose Computing an Updated Linearized Metric Based on a Non-Linear Constraint". With respect to Applicant’s arguments on p. 11 of the Remarks submitted 4/27/2026 that “Gondimalla is silent regarding “obtaining an algorithm for a computational graph including a series of interconnected nodes each representing a mathematical operation; grouping multiple nodes, among the series of interconnected nodes, into a plurality of sequences; tiling at least one node, among the multiple nodes, such that associated tensor data fits within an internal memory of a hardware device,” Examiner respectfully disagrees. Gondimalla discloses ([p. 6] "we propose a dynamic programming algorithm that optimally partitions a given CNN" [p. 9] "we partition a given CNN into sets of contiguous layers in which each partition’s dependence closure fits in the cache" [p. 6] "To minimize capacity misses, the largest tile that fits in the cache, called a maximal tile, is used" [p. 3] "we partition the CNN into sets of contiguous layers so that each partition’s dependence closure fits on chip (each partition reads its input map from and writes its output map to off-chip)"). For at least these reasons and those further detailed below, Examiner asserts that it is reasonable and appropriate to maintain the rejections under 35 USC 102/103. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-6, 8, 9, 11-18, and 20 are rejected under U.S.C. §102(a)(1) as being anticipated by Gondimalla (“Occam: Optimal Data Reuse for Convolutional Neural Networks”, 2022). Regarding claim 1, Gondimalla teaches A method comprising: obtaining an algorithm for a computational graph;([p. 2] "we describe CNNs in terms of access patterns" [p. 14 §4] "We implement Occam in CUDA. which is used prevalently to implement CNNs." CNN interpreted as an algorithm for a computational graph in light of the instant specification at least at ([¶0059] "a neural network computational graph (e.g., neural network graph 500-a, neural network graph 500-b, etc.) is a graphical representation of operations and data flow in a neural network (e.g., such as in an ANN, CNN, etc.).").) including a series of interconnected nodes each representing a mathematical operation;([p. 4] "As mentioned in Section 1, a CNN comprises many layers, each of which employs many filters to extract higher-order features. A layer’s output map is the next layer’s input map" [p. 4] "Each output map cell is computed by a dot product of a filter and a subset of cells in the input maps of the same “cuboidal” dimensions as the filter." A CNN is a type of ANN which has a series of interconnected nodes representing mathematical operations by definition, which is explicitly supported by the instant specification ([¶0055] "An artificial neural network (ANN) is a hardware or a software component that includes a number of connected nodes 505 (i.e., artificial neurons), which may loosely correspond to the neurons in a human brain") and reinforced by Gondimalla) grouping multiple nodes, among the series of interconnected nodes, into a plurality of sequences; ([p. 6] "we propose a dynamic programming algorithm that optimally partitions a given CNN" [p. 9] "we partition a given CNN into sets of contiguous layers in which each partition’s dependence closure fits in the cache" Nodes could be interpreted as layers or neurons, either of which is satisfied by Gondimalla) tiling at least one node, among the multiple nodes, such that associated tensor data fits within an internal memory of a hardware device; ([p. 6] "To minimize capacity misses, the largest tile that fits in the cache, called a maximal tile, is used" [p. 3] "we partition the CNN into sets of contiguous layers so that each partition’s dependence closure fits on chip (each partition reads its input map from and writes its output map to off-chip)") computing an initial linearized metric for a performance parameter for performing the algorithm using a hardware device;([p. 2] "CNNs are amenable to prefetching and multithreading, the problem is memory bandwidth" [p. 11] "If single-layer spans (SPAN(i,i + 1)) fit in the cache, we initialize OP[i,i + 1], 0 ≤ i < l" [p. 14] "we carefully scale the simulated GPU to match a slice of the above system that performs a single inference out of the 32-image mini-batch. This scaling reduces the compute bandwidth" [p. 11] "X, the optimal number of transfers" base span/initial number of off-chip feature map transfers OP[i,j].X interpreted as initial linearized metric for a performance parameter (total off-chip traffic) for performing the algorithm using a hardware device) wherein the performance parameter is associated with a non-linear multi-dimensional cost function; ([p. 2] "CNNs are amenable to prefetching and multithreading, the problem is memory bandwidth" [p. 11] "If single-layer spans (SPAN(i,i + 1)) fit in the cache, we initialize OP[i,i + 1], 0 ≤ i < l" [p. 10] "the layers must be partitioned into spans such that (1) each SPAN(i,j) satisfies the capacity constraint that the dependence closure (i.e., |DC(i, j)|) and weights (i.e., Σj−1 k=i |Wk|) of the span must fit in the cache, and (2) the total amount of data transferred off-chip at the partition boundaries is minimized [...] In a valid PBS, each span’s footprint fits in the cache: ∀0 ≤ i <l,(|DC(pi,pi+1)| + Σpi+1−1 k=pi |Wk|) < C" Gondimalla explicitly discloses a dynamic-programming cost formulation in which the off-chip transfer metric X is initialized as |Li|+|Lj| only when the span footprint satisfies the cache capacity inequality |DC(i,j)|+E|Wk|<C. Otherwise, the metric is updated by partitioning the span and minimizing the sum of subspan transfer metric. Because the inequality controls the branch of the DP cost calculation, the cost formulation is piecewise/discontinuous and therefore non-linear. The claimed "non-linear constraint" is derived from the same cost formulation as the thresholded condition on the span-footprint term that determined whether the base cost or partitioned cost applies. In other words the piecewise formulation of Gondimalla's cost function is OP[i,j].X={|Li|+|Lj|, if |DC(i,j|+E|Wk|<C, otherwise min_p(OP[i,p].X+OP[p,j].X}. The disclosed DP formulation is objectively a thresholded, piecewise optimization over multidimensional CNN tensor footprints.) computing an updated linearized metric based on the initial linearized metric and a non-linear constraint on the performance parameter ([p. 11] "The above table update step accurately tracks the number of off-chip transfers (i.e., X) of each of the two resulting spans" [p. 11] "If single-layer spans (SPAN(i,i + 1)) fit in the cache, we initialize OP[i,i + 1], 0 ≤ i < l" See also Eqn. 5 which shows the linearized metric update. If span fits in cache interpreted as non-linear constraint on the performance parameter which determines the linearized metric update.) wherein the non-linear constraint is derived from the non-linear multi-dimensional cost function;([p. 2] "CNNs are amenable to prefetching and multithreading, the problem is memory bandwidth" [p. 11] "If single-layer spans (SPAN(i,i + 1)) fit in the cache, we initialize OP[i,i + 1], 0 ≤ i < l" [p. 10] "the layers must be partitioned into spans such that (1) each SPAN(i,j) satisfies the capacity constraint that the dependence closure (i.e., |DC(i, j)|) and weights (i.e., Σj−1 k=i |Wk|) of the span must fit in the cache, and (2) the total amount of data transferred off-chip at the partition boundaries is minimized [...] In a valid PBS, each span’s footprint fits in the cache: ∀0 ≤ i <l,(|DC(pi,pi+1)| + Σpi+1−1 k=pi |Wk|) < C" Gondimalla explicitly discloses a dynamic-programming cost formulation in which the off-chip transfer metric X is initialized as |Li|+|Lj| only when the span footprint satisfies the cache capacity inequality |DC(i,j)|+E|Wk|<C. Otherwise, the metric is updated by partitioning the span and minimizing the sum of subspan transfer metric. Because the inequality controls the branch of the DP cost calculation, the cost formulation is piecewise/discontinuous and therefore non-linear. The claimed "non-linear constraint" is derived from the same cost formulation as the thresholded condition on the span-footprint term that determined whether the base cost or partitioned cost applies. In other words the piecewise formulation of Gondimalla's cost function is OP[i,j].X={|Li|+|Lj|, if |DC(i,j|+E|Wk|<C, otherwise min_p(OP[i,p].X+OP[p,j].X}. The disclosed DP formulation is objectively a thresholded, piecewise optimization over multidimensional CNN tensor footprints.) compiling instructions for performing the algorithm based on the updated linearized metric ([p. 12] "The DP algorithm is used to optimize partitions offline (like compiler optimizations)" [p. 13] "We implement Occam in CUDA. which is used prevalently to implement CNN […] Occam requires changes to the CUDA kernels of CNN implementations") programming the hardware device based on the compiled instructions to implement the computational graph, wherein the computational graph is implemented according to the updated linearized metric. ([p. 12] "The DP algorithm is used to optimize partitions offline (like compiler optimizations) […] Occam captures the enormous inter-image filter reuse by placing each partition – its filters and dependence closure – in a separate chip (e.g., a GPU, TPU, or FPGA)"). Regarding claim 2, Gondimalla teaches The method of claim 1, further comprising: identifying a plurality of layers of the computational graph;(Gondimalla [p. 2] "the large number of layers (e.g., 34 in ResNet), and filters per layers […] result in heavy compute and large intermediate data" See also FIG. 1) grouping the plurality of layers into a plurality of sequences; and(Gondimalla [p. 9] "We define a SPAN (i, j) as the convolution computations starting with Li as the input and ending with Lj as the output" [p. 10] "the layers must be partitioned into spans such that (1) each SPAN (i, j) satisfies the capacity constraint that the dependence closure" SPAN interpreted as sequence of layers) performing a dynamic programming process to identify a subset of sequences from the plurality of sequences, wherein the initial linearized metric is based on the subset of sequences, and wherein the dynamic programming process is based on a linearity constraint on the performance parameter.(Gondimalla [p. 12] "The DP algorithm is used to optimize partitions offline (like compiler optimizations) […] Occam captures the enormous inter-image filter reuse by placing each partition – its filters and dependence closure – in a separate chip (e.g., a GPU, TPU, or FPGA)" DP interpreted as dynamic programming.). Regarding claim 3, Gondimalla teaches The method of claim 2, further comprising: identifying a tiling of the plurality of layers, wherein the hardware device is programmed based on the tiling.(Gondimalla [p. 14] "We implement our DP algorithm as a stand-alone JavaScript application that takes as input the network parameters and produces the optimal partitions and tile dimensions" [p. 15] "we use Occam’s optimal partition algorithm for Layer Fusion to compute the partitions with the largest square tile whose dependence closure for a given partition would fit in the cache (a different tile size for each partition). Because the tiles are suboptimal even though the partitions are optimal, Layer Fusion does not capture full reuse. We verified the functional correctness of these schemes by comparing against an unmodified ConvNet" See also FIG. 6). Regarding claim 4, Gondimalla teaches The method of claim 2, further comprising: performing an additional dynamic programming process to identify an updated subset of sequences from the plurality of sequences, wherein the updated linearized metric is based on the updated subset of sequences, wherein the hardware device is programmed based on the updated subset of sequences.(Gondimalla [p. 16 §5.1] "In Table 2, we present the optimal partitions and tile dimensions for our networks for 3-MB on chip memory. While the tile dimensions for Occam and Layer Fusion are different, we use Occam’s partition algorithm to derive partitions for Layer Fusion (see Section 4). The partitions are shown using the start layer for each" See Table 2 which shows a plurality of additional DP processes, each corresponding to a different model). Regarding claim 5, Gondimalla teaches The method of claim 4, further comprising: determining an initial weight for the performance parameter; and(Gondimalla [p. 10] "weights (i.e., Σj−1 k=i |Wk |) of the span must fit in the cache" Wk interpreted as initial weight for the performance parameter) computing an initial score for the subset of sequences based on the initial linearized metric and the initial weight.(Gondimalla [p. 10] "(|DC(pi ,pi+1) | + Σpi+1−1 k=pi |Wk |)" Interpreted as initial score for the subset of sequences (defined by pi, pi+1) based on the initial linearized metric and the initial weight). Regarding claim 6, Gondimalla teaches The method of claim 5, further comprising: computing a weighted sum of a plurality of linearized metrics including the initial linearized metric based on a plurality of initial weights including the initial weight, wherein the initial score is based on the weighted sum.(Gondimalla [p. 11] "OP[i,j].X = OP[i,popt].X +OP[popt,j].X" See Eqn. 5 and Eqn. 7). Regarding claim 8, Gondimalla teaches The method of claim 5, further comprising: computing scores for a plurality of subsets of sequences, respectively, wherein the subset of sequences is selected from the plurality of subsets of sequences based on the scores.(Gondimalla [p. 2] "CNNs are amenable to prefetching and multithreading, the problem is memory bandwidth" [p. 11] "If single-layer spans (SPAN(i,i + 1)) fit in the cache, we initialize OP[i,i + 1], 0 ≤ i < l" [p. 14] "we carefully scale the simulated GPU to match a slice of the above system that performs a single inference out of the 32-image mini-batch. This scaling reduces the compute bandwidth" [p. 11] "X, the optimal number of transfers" [p. 11] "we consider every possible partition of SPAN (i, j) and pick the partition point p that yields the fewest transfers in the two resulting sub-spans, SPAN (i,p) and SPAN (p, j)." X of OP[i,j].X interpreted as score for subset SPAN(i,j) which is selected based on X.). Regarding claim 9, Gondimalla teaches The method of claim 5, further comprising: computing a non-linear term based on the initial weight and the non-linear constraint, wherein the initial score is based on the non-linear term.(Gondimalla [p. 9] "DC(i,j) defines the dependence closure of one row-plane of the output feature map in Lj extending back to the feature map of layer Li"). Regarding claim 11, Gondimalla teaches The method of claim 1, further comprising: modifying a design for the hardware device based on the updated linearized metric, wherein the hardware device is programmed based on the modified design.(Gondimalla [p. 14] "We implement our DP algorithm as a stand-alone JavaScript application that takes as input the network parameters and produces the optimal partitions and tile dimensions. We then feed these outputs to ConvNet." See FIG. 5). Regarding claim 12, Gondimalla teaches The method of claim 1, further comprising: modifying an algorithm for the computational graph based on the updated linearized metric, wherein the hardware device is programmed based on the modified algorithm.(Gondimalla [p. 1] "we propose a dynamic programming algorithm that optimally partitions a given CNN to guarantee the least off-chip traffic at the partition boundaries for a given on-chip capacity" Partitioning the CNN using dynamic programming interpreted as synonymous with modifying an algorithm (the CNN) for the computational graph based on the updated linearized metric, wherein the hardware device is programmed based on the modified algorithm.). Regarding claim 13, claim 13 is directed towards an apparatus for performing the method of claim 1. Therefore, the rejection applied to claim 1 also applies to claim 13. Claim 13 also recites additional elements a processor and a memory storing instructions and in electronic communication with the processor, the processor being configured to execute the instructions to (Gondimalla [p. 3] "Finally, the input-stationary approach [11] holds on-chip input/output maps and fetches filters from off-chip, ignoring filter reuse across images (e.g., TPU). Occam’s partitions reside on different chips (e.g., GPU, TPU, or FPGA) available in a multi-accelerator environment such as data centers, forming a pipeline so that a partition’s filters and dependence closure remain on-chip as different images pass through (i.e., each partition incurs off-chip traffic only for its inputs and outputs)"). Similarly, regarding claims 14-18, claims 14-18 are directed towards an apparatus for performing the method of claims 2-6, therefore, the rejections applied to claims 2-6 also apply to claims 14-18. Regarding claim 20, claim 20 is directed towards a non-transitory computer readable medium storing code, the code comprising instructions executable by a processor for performing the method of claim 1. Therefore, the rejection applied to claim 1 also applies to claim 20. Claim 20 also recites additional elements A non-transitory computer readable medium storing code, the code comprising instructions executable by a processor to (Gondimalla [p. 3] "Finally, the input-stationary approach [11] holds on-chip input/output maps and fetches filters from off-chip, ignoring filter reuse across images (e.g., TPU). Occam’s partitions reside on different chips (e.g., GPU, TPU, or FPGA) available in a multi-accelerator environment such as data centers, forming a pipeline so that a partition’s filters and dependence closure remain on-chip as different images pass through (i.e., each partition incurs off-chip traffic only for its inputs and outputs)"). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 7 and 19 are rejected under U.S.C. §103 as being unpatentable over the combination of Gondimalla and Hu (“Learning Anytime Predictions in Neural Networks via Adaptive Loss Balancing”, 2019). Regarding claim 7, Gondimalla teaches The method of claim 5. However, Gondimalla doesn't explicitly teach determining an updated weight for the performance parameter based on the initial score; and computing an updated score for the subset of sequences based on the updated weight and the updated linearized metric, wherein the subset of sequences is selected based on the updated score. Hu, in the same field of endeavor, teaches determining an updated weight for the performance parameter based on the initial score; and computing an updated score for the subset of sequences based on the updated weight and the updated linearized metric, wherein the subset of sequences is selected based on the updated score.([p. 4] "we form the following joint optimization over 0 and Bi for general losses without probability models:" See Eqn. 3 where li is a performance parameter, the initial L for i=0, Bi is the updated weight for i+1, and the updated score is L for i+1). Gondimalla as well as Hu are directed towards neural network optimization. Therefore, Gondimalla as well as Hu are reasonably pertinent analogous art. It would have been obvious before the effective filing date of the claimed invention to use the AdaLoss function in Hu with the OCCAM system in Gondimalla by equating B with W and l with X. Hu provides as additional motivation for combination ([p. 2] “we find the joint optimization equivalent to optimizing the geometric mean of the expected training losses, an objective that treats the relative improvement of each loss equally. Empirically, we show on multiple models and visual recognition data-sets that the proposed adaptive weights outperform natural, non-adaptive weighting schemes”). Regarding claim 19, claim 19 is directed towards an apparatus for performing the method of claim 7. Therefore, the rejection applied to claim 7 also applies to claim 19. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Wu (US12190404B2) is directed towards resource constrained computational graph compilation based on nonlinear cost functions. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SIDNEY VINCENT BOSTWICK whose telephone number is (571)272-4720. The examiner can normally be reached M-F 7:30am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached on (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SIDNEY VINCENT BOSTWICK/Examiner, Art Unit 2124
Read full office action

Prosecution Timeline

Apr 24, 2023
Application Filed
Dec 20, 2025
Non-Final Rejection (signed) — §102, §103
Jan 26, 2026
Non-Final Rejection mailed — §102, §103
Feb 25, 2026
Applicant Interview (Telephonic)
Feb 25, 2026
Examiner Interview Summary
Apr 27, 2026
Response Filed
Jul 22, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699874
Leveraging Redundancy in Attention with Reuse Transformers
3y 10m to grant Granted Aug 04, 2026
Patent 12675673
NEURAL NETWORK PROCESSING DEVICE, METHOD, AND COMPUTER-READABLE RECORDING MEDIUM
3y 7m to grant Granted Jul 07, 2026
Patent 12645914
INSTRUCTION PRUNING FOR NEURAL NETWORKS
3y 6m to grant Granted Jun 02, 2026
Patent 12626139
SECRET SOFTMAX FUNCTION CALCULATION SYSTEM, SECRET SOFTMAX FUNCTION CALCULATION APPARATUS, SECRET SOFTMAX FUNCTION CALCULATION METHOD, SECRET NEURAL NETWORK CALCULATION SYSTEM, SECRET NEURAL NETWORK LEARNING SYSTEM, AND PROGRAM
4y 3m to grant Granted May 12, 2026
Patent 12619815
Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation
1y 6m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
52%
Grant Probability
89%
With Interview (+36.9%)
4y 5m (~1y 1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 147 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month