Prosecution Insights
Last updated: October 02, 2026
Application No. 18/575,139

METHOD AND APPARATUS FOR FUSING LAYERS OF DIFFERENT MODELS

Non-Final OA §101§103
Filed
Dec 28, 2023
Priority
Nov 30, 2021 — nonprovisional of PCTCN2021134286
Examiner
YI, DAVID
Art Unit
Tech Center
Assignee
Intel Corporation
OA Round
1 (Non-Final)
69%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
178 granted / 258 resolved
+9.0% vs TC avg
Strong +64% interview lift
Without
With
+64.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 11m
Avg Prosecution
8 currently pending
Career history
263
Total Applications
across all art units

Statute-Specific Performance

§101
18.6%
-21.4% vs TC avg
§103
46.3%
+6.3% vs TC avg
§102
12.9%
-27.1% vs TC avg
§112
19.0%
-21.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 258 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1 – 20 are cancelled by the applicant. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 21 - 40 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception without significantly more. Regarding Claim 21: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites the abstract ideas of: search layers from different models and determine whether to perform layer fusing; This limitation is directed to the abstract idea of a mental process as searching layers and determine whether to perform fusing is analogous to evaluation and judgment, which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). combine input data for the instructions in the layers from different models into a combined input data; This limitation is directed to the abstract idea of a mental process as combining input data is analogous to evaluation and judgment, which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application. A layer fusing apparatus, comprising interface circuitry; processor circuitry coupled to the interface circuitry and configured to. This limitation recites generic computer components such as processors and circuits, which invokes a computer merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to integrate the exception into a practical application. fuse instructions in the layers from different models into a fused instruction in response to determining to perform layer fusing; The claim further recites fusing instructions in the layers. However, fusing instructions merely provides instructions to apply the mental concept and therefore does not integrate the judicial exception into a practical application. The claim does not improve the functioning of a computer or another technology (MPEP 2106.05(f)). load the combined input data for the fused instruction from the continuous storage area in the memory to perform the fused instruction; This limitation is directed to mere data gathering, which is an insignificant extra-solution activity [see MPEP 2106.05(g)(3)] and therefore fails to integrate the judicial exception into a practical application. store output data obtained after performing the fused instruction into a continuous storage area in the memory. This limitation recites as storing data in a memory, which is considered an insignificant extra solution activity, as it is merely storing data which is a conventional computer function, therefore, it does not impose meaningful limits on the claim such that it is not nominally or tangentially related to the invention per 2106.05(g). Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception. A layer fusing apparatus, comprising interface circuitry; processor circuitry coupled to the interface circuitry and configured to. This limitation invokes a computer merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to amount to significantly more than the judicial exception. fuse instructions in the layers from different models into a fused instruction in response to determining to perform layer fusing; The additional elements, fusing instructions in the layers, do not amount to significantly more than the abstract idea. Fusing instructions merely provides instructions to apply the mental concept. The claim does not improve the functioning of a computer or another technology (MPEP 2016.04(d)(1)). load the combined input data for the fused instruction from the continuous storage area in the memory to perform the fused instruction; This limitation is directed to receiving or transmitting data over a network, which the courts have recognized as well-understood, routine, conventional activity when they are claimed at a high level of generality or as insignificant extra-solution activity [see MPEP 2106.05(d) II. i] and therefore fails to amount to significantly more than the judicial exception. store output data obtained after performing the fused instruction into a continuous storage area in the memory. This limitation recites storing data which amounts to storing and retrieving information in memory, further considered well-understood, routine and conventional under MPEP 2106.05(d) II (iv). Step 2A Prong Two and Step 2B: Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible. Regarding Claim 22: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites the abstract ideas of: the processor circuitry is further configured to determine whether to perform layer fusing based on a first fusing metric. This limitation is directed to the abstract idea of a mental process as determining whether to perform layer fusing is analogous to an analogous to evaluation and judgment, which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). Regarding Claim 23: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Claim 23 does not recite abstract ideas other than the ones recited at claim 21. Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application. the first fusing metric is a degree of saturation indicating instruction utilization in a layer that characterizes a ratio of the portion of a register actually utilized by the instruction to the whole register allocated for the instruction. This limitation recites description of the first metric is a degree of saturation indicating instruction utilization in a layer. Therefore, this limitation amounts to merely indicating a field of use or technological environment [see MPEP 2106.05(h)] and fails to integrate the judicial exception into a practical application. Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception. the first fusing metric is a degree of saturation indicating instruction utilization in a layer that characterizes a ratio of the portion of a register actually utilized by the instruction to the whole register allocated for the instruction. This limitation recites description of the of the first metric is a degree of saturation indicating instruction utilization in a layer. Therefore, this limitation amounts to merely indicating a field of use or technological environment [see MPEP 2106.05(h)] and therefore fails to amount to significantly more than the judicial exception. Step 2A Prong Two and Step 2B: Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible. Regarding Claim 24: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites the abstract ideas of: the processor circuitry is further configured to determine whether to perform layer fusing based on a second fusing metric. This limitation is directed to the abstract idea of a mental process as determining whether to perform layer fusing is analogous to an analogous to evaluation and judgment, which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). Regarding Claim 25: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites the abstract ideas of: the second fusing metric is an impact factor calculated for layers from different models, the impact factor indicating whether performing layer fusing will add sync points overhead that would cause performance hits. This limitation is directed to the abstract idea of mathematical concepts, as calculating a factor is analogous to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.) Regarding Claim 26: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites the abstract ideas of: the processor circuitry is further configured to determine whether to perform layer fusing based on a third fusing metric. This limitation is directed to the abstract idea of a mental process as determining whether to perform layer fusing is analogous to an analogous to evaluation and judgment, which can be performed by the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) III. C.). Regarding Claim 27: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Yes, the claim recites the abstract ideas of: the third metric is a score calculated for layers from different models, the score indicating the benefit it may get after performing layer fusing. This limitation is directed to the abstract idea of mathematical concepts, as calculating a score is analogous to a mathematical calculation (see MPEP 2106.04(a)(2) I. C.) Regarding Claim 28: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Claim 28 does not recite abstract ideas other than the ones recited at claim 21. Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application. the processor circuitry is further configured to allocate a continuous buffer pool as a shared storage area in the memory for the fused instructions. This limitation recites generic computer, which invokes a computer merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to integrate the exception into a practical application. Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception. the processor circuitry is further configured to allocate a continuous buffer pool as a shared storage area in the memory for the fused instructions. This limitation invokes a computer merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to amount to significantly more than the judicial exception. Step 2A Prong Two and Step 2B: Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible. Regarding Claim 29: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Claim 29 does not recite abstract ideas other than the ones recited at claim 21. Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application. the processor circuitry is further configured to provide an interface for loading the combined input data for the fused instruction and storing the output data obtained after performing operation of the fused instruction. This limitation recites as an insignificant extra solution activity, as receiving data and obtaining an output under BRI, is mere data gathering per MPEP 2106.05(g)(3). Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception. the processor circuitry is further configured to provide an interface for loading the combined input data for the fused instruction and storing the output data obtained after performing operation of the fused instruction. This limitation is directed to receiving data, which the courts have recognized as well-understood, routine, conventional activity when they are claimed at a high level of generality or as insignificant extra-solution activity [see MPEP 2106.05(d) II. i] and therefore fails to amount to significantly more than the judicial exception. Step 2A Prong Two and Step 2B: Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible. Regarding Claim 30: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Claim 30 does not recite abstract ideas other than the ones recited at claim 21. Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application. the instructions in the layers from different models comprise Single Instruction Multiple Data (SIMD) instructions that includes Vector Neural Network Instruction (VNNI), Tile matrix multiply unit (TMUL) and Advanced Matrix Extension (AMX). Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception. the instructions in the layers from different models comprise Single Instruction Multiple Data (SIMD) instructions that includes Vector Neural Network Instruction (VNNI), Tile matrix multiply unit (TMUL) and Advanced Matrix Extension (AMX). Step 2A Prong Two and Step 2B: Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible. Regarding claims 31- 39: Claims 31 - 39 recites analogous limitations to claims 21 - 29 (respectively) and therefore they are rejected on the same grounds as claims 21 - 29. Regarding Claim 40: Step 1 – Is the claim to a process, machine, manufacture, or composition of matter? Yes Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon? Claim 40 does not recite abstract ideas other than the ones recited at claim 31. Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application? – No, there are no additional elements that integrate the judicial exception into a practical application. A computer-readable storage medium with program instructions stored thereon which, when executed by a processor, cause the processor to implement the method of claim 31. This limitation recites generic computer, which invokes a computer merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to integrate the exception into a practical application. Step 2B – Does the claim recite any additional elements that amount to significantly more than the judicial exception? – No, there are no additional elements that amount to significantly more than the judicial exception. A computer-readable storage medium with program instructions stored thereon which, when executed by a processor, cause the processor to implement the method of claim 31. This limitation invokes a computer merely as a tool for performing an existing process [see MPEP 2106.05(f)(2)] and therefore fails to amount to significantly more than the judicial exception. Step 2A Prong Two and Step 2B: Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. The claim is ineligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or non-obviousness. Claim(s) 21 - 22, 24 - 27, 29, 31-32, 34-37, 39-40 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayanan (NPL, “Accelerating Deep Learning Workloads through Efficient Multi-Model Execution”, dated on 12/07/2018, by Narayanan et al – hereinafter Narayanan) in view of Chen (USPB, US20220129289A1, PCT file data: 08/25/2020, by Chen et al, hereinafter Chen). Referring to Claim 21, Narayanan teaches: A layer fusing apparatus, comprising interface circuitry; processor circuitry coupled to the interface circuitry and configured to. See Narayanan at [Page 3, bottom]:” The compiler composes models in a model batch into a single computation DAG, fuses preprocessing pipelines, and performs layer fusion across models where possible. The runtime then executes this optimized DAG on the GPU, trying to run non-dependent computations concurrently." Examiner interprets the compiler performing layer fusion across models as equivalent as the function of the layer fusing as claimed. Also See Narayanan at [Page 5, bottom]:” Experiments were run on two machines: 1) a machine in a private cluster with 28 CPU cores and an NVIDIA P100 GPU (henceforth called P100), and 2) a p3.2xlarge instance on Amazon EC2 with 8 hyperthreads and an NVIDIA V100 GPU (henceforth called V100). Experiments were run using CUDA 9.0 and CuDNN 7.0.” Examiner interprets the machines with EC2, CPUs and GPUs as equivalent as the interface circuitry and processor circuitry as claimed. Thus, Narayanan teaches the limitation. search layers from different models and determine whether to perform layer fusing; See Narayanan at [Page 4, mid]: "Each node in the computation DAG represents a layer (convolution, fully-connected layer, activation) of one of the models in the model batch; each edge represents a data dependency” and Narayanan at [Page 4, bottom]:” HiveMind can then search the graph for all instances of these relationships to determine the sets of nodes to be fused." Examiner interprets HiveMind searching the graph for all instances and determining the sets of nodes to be fused as equivalent as searching layers from different models and determining to perform layer fusing since each node in the computation DAG (Directed Acyclic Graph) represents a layer of one model. Also see Narayanan at [Page 4, bottom]:” Layer fusion is applicable in three settings: 1) when stateful operators in different models share the same underlying weights (and where gradients for the shared weights across different models can be safely aggregated), 2) when stateful operators share the same inputs and have same output shapes, and 3) when non-stateful operators operate on inputs of the same shape” Narayanan discloses three settings that whether the layer fusion is applicable, which is as equivalent as the metrics that determining whether to perform layer fusing in claim 22, claim 24, and claim 26. combine input data for the instructions in the layers from different models into a combined input data; See Narayanan at [Page 2, mid]:” Second, model search workloads like Efficient Neural Architecture Search (ENAS) [14] as well as model ensembles used for inference often feature models with shared weights – we can exploit this to concatenate inputs, increasing the computational intensity of operations like convolutions and reducing the kernel launch overhead. Finally, ensembles of fine-tuned models can share the first k layers, allowing for inputs to be shared while concatenating weights.” Examiner interprets concatenating inputs and allowing inputs to be shared while concatenating weights as equivalent as combining input data into a combined input data. allocate a continuous storage area in a memory for the combined input data; See Narayanan at [Page 5, top]: "The first step is ensuring that the input and output buffers of the new fused layer are contiguous, as cuDNN's GPU kernel implementations require inputs and outputs to be allocated contiguous memory." Narayanan discloses allocating contiguous memory for the input buffers and also for the output. Thus, Narayanan teaches the limitation. load the combined input data for the fused instruction from the continuous storage area in the memory to perform the fused instruction; See Narayanan at [Page 4, Algorithm 1]: PNG media_image1.png 205 656 media_image1.png Greyscale As discussed above, Narayanan discloses the contiguous memory space is allocated for the input buffer. In Algorithm 1, line 4, “NewInputBuffer = REASSIGNINPUTBUFFERS (prevNodes)” the new input buffer is created (or reassigned) by function REASSIGNINPUTBUFFERS(…) with the parameter prevNodes, and stored in the memory as a combined input data. Also, in line 6, newInputBuffer is loaded by the function CREATENODE(…) as the parameter to create newNode, which is as equivalent as loading the combined input data form the memory to perform the fused instruction. store output data obtained after performing the fused instruction into a continuous storage area in the memory. See Narayanan at [Page 4, Algorithm 1]: in line 5, newOutputBuffer is created (or reassigned) by function REASSIGNOUTPUTBUFFERS(…). Thus, newOutputBuffer is stored in the memory. As discussed above, Narayanan discloses the contiguous memory space is allocated to the output buffer. Thus, Narayanan teaches the limitation. However, it fails to teach: fuse instructions in the layers from different models into a fused instruction in response to determining to perform layer fusing; Chen teaches: fuse instructions in the layers from different models into a fused instruction in response to determining to perform layer fusing; See Chen at [0074]: "an operation instruction in NCLAPI is an abstract representation of transformation, which can be used to represent a deep learning processing layer, or can be used to represent general computations." And see Chen at [0245]: "The fusion operation instruction includes an operation instruction obtained by combining a plurality of operation instructions according to a calling sequence." As discussed above, Narayanan discloses fusing nodes (a node represents a layer) from different models. Chen discloses that a deep-learning layer may be represented by an operation instruction and that a plurality of such operation instructions may be combined into one fusion operation instruction. Accordingly, applying Chen’s known instruction-level representation to Narayanan’s fused node would have resulted in fusing the instructions associated with the selected layers into the claimed fused instruction. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Narayanan with the above teachings of Chen by fusing the layers from different models, as taught by Narayanan, fusing the instructions into a fused instruction, as taught by Chen. The modification would have been obvious because one of ordinary skill in the art would be motivated to effectively improving the performance optimization effect of deep learning algorithms in corresponding hardware platforms, See Chen at [0010]: “With the use of the deep learning algorithm compiling method , the device , and the related product provided by various aspects of the embodiments of the present disclosure , the compilation process can be adaptively changed according to different types of operation instructions , thereby greatly improving compilation flexibility and efficiency , effectively improving the performance optimization effect of deep learning algorithms in corresponding hardware platforms , and then further improving the processing performance of deep learning processor.” Referring to Claim 22, Narayanan-Chen teaches the layer fusing apparatus of claim 21, Narayanan also teaches: the processor circuitry is further configured to determine whether to perform layer fusing based on a first fusing metric. As discussed in claim 21, Narayanan discloses three settings that whether the layer fusion is applicable, which is as equivalent as the metric that determines whether to perform layer fusing. The same motivation that was utilized for combining Narayanan with Chen as set forth in claim 21 is equally applicable to claim 22. Referring to Claim 24, Narayanan-Chen teaches the layer fusing apparatus of claim 21, Narayanan also teaches: the processor circuitry is further configured to determine whether to perform layer fusing based on a second fusing metric. As discussed in claim 21, Narayanan discloses three settings that whether the layer fusion is applicable, which is as equivalent as the metric that determines whether to perform layer fusing. The same motivation that was utilized for combining Narayanan with Chen as set forth in claim 21 is equally applicable to claim 24. Referring to Claim 25, Narayanan-Chen teaches the layer fusing apparatus of claim 21, Narayanan fails to teach: the second fusing metric is an impact factor calculated for layers from different models, the impact factor indicating whether performing layer fusing will add sync points overhead that would cause performance hits. Chen teaches: the second fusing metric is an impact factor calculated for layers from different models, the impact factor indicating whether performing layer fusing will add sync points overhead that would cause performance hits. See Narayanan at [Page 5, mid]:” Data Dependencies. DAGs in HiveMind can have nodes with multiple in-edges or out-edges, on account of layers with skip connections, or layer fusion. These layers could be executed on different streams, and thus HiveMind needs to manage these data dependencies. HiveMind uses the CUDA Event API to ensure that kernels are executed only once their inputs are available.” And see Chen at [0123]: “In a possible implementation manner, a programming model can be designed based on three factors, which are: data flow, execution flow, and control flow….In terms of the execution flow, an embodiment of the present disclosure divides the execution mode of a computational model into three categories, which are: calling layer by layer: calling all operations in a computational model one by one; fusion and calling: performing operation fusion on an entire computational model, and then calling a fusion operation; and fusing and calling by segment: performing operation fusion on a computational model by segment, and then calling a fusion operation by segment. An NCLAPI programming model is distinguished according to the following three types of execution methods: the execution method of layer-by-layer calling corresponds to an imperative programming model; the execution method of fusion calling corresponds to a declarative programming model; the execution method of fusing and calling by segment corresponds to a mixed programming model.” Also see Chen at [Page 212]:” The cost model estimates the total time (including data format conversion and other overhead) a deep learning processor takes to perform an operation. Main factors that the model considers are: memory access time of the operation, operating time, and the overlap rate of the two. The three factors directly determine the performance of the program. Specifically, the deep learning processor includes an independent operation unit and a memory access unit. A computation instruction and a memory access instruction can be executed in a pipeline or overlapped in execution.” Chen discloses designing the execution mode including layer-by-layer, full fusion, and segmented fusion execution-based on control-flow factors that include device synchronization and further discloses calculating total execution time including overhead using a cost model. Applying that calculated synchronization-related factor to Narayanan’s candidate layers from different models indicates whether the proposed fusion would add CUDA-event synchronization overhead. The same motivation that was utilized for combining Narayanan with Chen as set forth in claim 21 is equally applicable to claim 25. Referring to Claim 26, Narayanan-Chen teaches the layer fusing apparatus of claim 21, Narayanan also teaches: the processor circuitry is further configured to determine whether to perform layer fusing based on a third fusing metric. As discussed in claim 21, Narayanan discloses three settings that whether the layer fusion is applicable, which is as equivalent as the metric that determines whether to perform layer fusing. The same motivation that was utilized for combining Narayanan with Chen as set forth in claim 21 is equally applicable to claim 26. Referring to Claim 27, Narayanan-Chen teaches the layer fusing apparatus of claim 21, Narayanan also teaches: the third metric is a score calculated for layers from different models, the score indicating the benefit it may get after performing layer fusing. See Narayanan at [Page 6, mid]: “Layer Fusion. We show the impact of layer fusion on the AlexNet inference workload across 4 models in Figure 5b. The number of layers fused is varied from 0 to 16; we see that in the best case; layer fusion provides a 1.6X speedup on the P100 and a 1.3X speedup on the V100.” Narayanan discloses layer fusion provides a 1.6X and a 1.3X speedup on P100 and V100, which is as equivalent as the benefit after performing layer fusing, and the time in Fig 5(b) is as equivalent as the score calculated for layers since the time is less the benefit is higher. PNG media_image2.png 300 1007 media_image2.png Greyscale The same motivation that was utilized for combining Narayanan with Chen as set forth in claim 21 is equally applicable to claim 27. Referring to Claim 29, Narayanan-Chen teaches the layer fusing apparatus of claim 21, Narayanan also teaches: the processor circuitry is further configured to provide an interface for loading the combined input data for the fused instruction and storing the output data obtained after performing operation of the fused instruction. As discussed in claim 1, Narayanan teaches computer components such as Amazon EC2, CPU, GPU as the interface circuitry, loading the input buffer as the combined input data, allocating a memory space for output buffer, and storing the output buffer. Thus, Narayanan teaches the limitation. The same motivation that was utilized for combining Narayanan with Chen as set forth in claim 21 is equally applicable to claim 29. Referring to Claims 31 – 32, these claims are rejected on the same basis as claims 21 - 22, mutatis mutandis, since they are analogous claims. Referring to Claims 34 – 37, these claims are rejected on the same basis as claims 24 - 27, mutatis mutandis, since they are analogous claims. Referring to Claim 39, these claims are rejected on the same basis as claim 29, mutatis mutandis, since they are analogous claims. Referring to Claim 40, Narayanan-Chen teaches the layer fusing apparatus of claim 21 (and claim 31), Narayanan also teaches: A computer-readable storage medium with program instructions stored thereon which, when executed by a processor, cause the processor to implement the method of claim 31. See Narayanan at [Page 2, mid]: “First, small-model and small-minibatch workloads contain GPU kernels with low computational intensity (number of floating point operations performed per byte loaded from memory) – this means that the GPU’s compute units are often stalled on memory reads.” Examiner interprets the memory as equivalent as the computer-readable storage medium. The same motivation that was utilized for combining Narayanan with Chen as set forth in claim 21 is equally applicable to claim 40. Claim(s) 23 and 33 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayanan - Chen in view of Lang (NPL, “Make the most out of your SIMD investments: counter control flow divergence in compiled query pipelines”, Pub. Date: 07/16/2019 by Lang et al – hereinafter Lang). Referring to Claim 23, Narayanan-Chen teaches the layer fusing apparatus of claim 21. However, it fails to teach: the first fusing metric is a degree of saturation indicating instruction utilization in a layer that characterizes a ratio of the portion of a register actually utilized by the instruction to the whole register allocated for the instruction. Lang teaches: the first fusing metric is a degree of saturation indicating instruction utilization in a layer that characterizes a ratio of the portion of a register actually utilized by the instruction to the whole register allocated for the instruction. See Lang at [Page 762, right-bottom]:” Figure 6b visualizes the same workload with in-register buffering, following consume everything semantics. …In this example, we require the utilization to be at least 75% (six out of eight lanes need to be active).” Lang discloses the in-register buffering and a utilization metric that six out of eight lanes (75%), which is as equivalent as the ratio of the portion of a register actually utilized by the instruction to the whole register allocated for the instruction as claimed. Thus, Lang teaches the limitation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Narayanan - Chen with the above teachings of Lang by fusing the layers from different models, as taught by Narayanan - Chen, a degree of saturation indicating instruction utilization in a layer that characterizes a ratio of the portion of a register actually utilized by the instruction to the whole register allocated for the instruction, as taught by Lang. The modification would have been obvious because one of ordinary skill in the art would be motivated to significantly increasing the lane utilization and reducing the overall execution time, See Lang at [762, right-bottom]: “The purple and black vertical lines indicate that tuples are written to the buffers, or read from the buffer, respectively. Compared to the divergent implementation, the lane utilization has significantly increased, and the overall execution time has reduced.” Referring to Claim 33, these claims are rejected on the same basis as claim 23, mutatis mutandis, since they are analogous claims. Claim(s) 28 and 38 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayanan - Chen in view of Fang (USPB: US20180307973A1, Pub. Dated on 10/25/2018 by Fang et al – hereinafter Fang). Referring to Claim 28, Narayanan-Chen teaches the layer fusing apparatus of claim 21. However, it fails to teach: the processor circuitry is further configured to allocate a continuous buffer pool as a shared storage area in the memory for the fused instructions. Fang teaches: the processor circuitry is further configured to allocate a continuous buffer pool as a shared storage area in the memory for the fused instructions. See Fang at [0108]: "The buffer pool 16 is configured for storing the neural network data and the computation result of the computation unit, including intermediate computation result and final computation result." See Fang at [0121]: "In the present disclosure, the DPU core no longer adopts the traditional input buffer->computation complex->output buffer structure. Instead, it proposes a novel structure, including: a buffer pool, a multi-channeled data writing scheduling unit and a multi-channeled data reading scheduling unit. With this kind of configuration, the DPU may support concurrent request for any number of writing/reading channels, and each writing/reading channel may have access to the whole buffer pool. Thus, the DPU may request, distribute and utilize the buffer space according to actual computation demand, especially considering the characteristics of the CNN algorithm. Therefore, the on-chip buffer utilization may be increased significantly." As discussed in claim 21, Narayanan-Chen discloses allocating the continuous memory space for input buffer and output buffer for of the fused layer. Fang discloses the structure including a buffer pool, multi-channeled data writing and reading scheduling unit that supports concurrent I/O operations, which is as equivalent as the buffer pool as a shared storage area in the memory. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Narayanan - Chen with the above teachings of Fang by allocating the continuous memory space for input buffer and output buffer for of the fused layer, as taught by Narayanan - Chen, allocate a continuous buffer pool as a shared storage area in the memory, as taught by Fang. The modification would have been obvious because one of ordinary skill in the art would be motivated to make the computation unit in operating status as much as possible and maximizing its efficiency, See Fang at [0103]: “However, in the embodiment of the present disclosure, by arranging a plurality instruction units (in this case, two instruction units) in the DPU core, the processor may obtain two independent instruction steams concurrently, eliminating the dependency between different instructions. Moreover, when the execution of one instruction steam pauses, the processor may switch to execute the other instruction stream. In this way, a plurality of instruction steams may share the same computation unit, so that the computation unit will be in operating status (rather than idle or waiting mode) as much as possible, maximizing its efficiency.” Referring to Claim 38, these claims are rejected on the same basis as claim 28, mutatis mutandis, since they are analogous claims. Claim(s) 30 is/are rejected under 35 U.S.C. 103 as being unpatentable over Narayanan - Chen in view of Intel (NPL, “Intel® Architecture Instruction Set Extensions and Future Features Programming Reference”, Pub. Date: June 2020 by Intel Corporation – hereinafter Intel). Referring to Claim 30, Narayanan-Chen teaches the layer fusing apparatus of claim 21. However, it fails to teach: the instructions in the layers from different models comprise Single Instruction Multiple Data (SIMD) instructions that includes Vector Neural Network Instruction (VNNI), Tile matrix multiply unit (TMUL) and Advanced Matrix Extension (AMX). Intel teaches: the instructions in the layers from different models comprise Single Instruction Multiple Data (SIMD) instructions that includes Vector Neural Network Instruction (VNNI), Tile matrix multiply unit (TMUL) and Advanced Matrix Extension (AMX). See Intel at [Page 1-1, mid]:” The base of the 512-bit SIMD instruction extensions are referred to as Intel AVX-512 Foundation instructions. They include extensions of the Intel AVX and Intel AVX2 family of SIMD instructions but are encoded using a new encoding scheme with support for 512-bit vector registers, up to 32 vector registers in 64-bit mode, and conditional processing using opmask registers.” See Intel at [Page V, top]:” Updated Table 1-2 “Recent Instruction Set Extensions /Features Introduction in Intel® 64 and IA-32 Processors” to list the AVX512_VNNI instruction set architecture on a separate line due to presence on future processors available sooner than previously listed.” See Intel at [Page 3-1, top]: "Intel AMX instructions are synchronous in the Intel architecture instruction stream and the memory loaded and stored by the tile instructions is coherent with respect to the host's memory accesses. There are no restrictions on interleaving of Intel architecture and Intel AMX code or restrictions on the resources the host can use in parallel with Intel AMX (e.g., Intel AVX-512)." And see Intel at [Page 3-2, top]:” The TMUL instruction set (defined to be CPUID bits AMX-BF16 and AMX-INT8) only supports reg-reg operations.” Intel discloses SIMD, AMX, TMUL AND VNNI instructions. Some of instructions do not stand alone, that distinction does not affect the rejection because the limitation is not limited to special register functions, or any register width, but requires only the utilized-to-total ratio for the register allocated to the instruction. Thus, Intel teaches the limitation. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Narayanan - Chen with the above teachings of Intel by the instructions in the layers from different models, as taught by Narayanan - Chen, Single Instruction Multiple Data (SIMD) instructions that includes Vector Neural Network Instruction (VNNI), Tile matrix multiply unit (TMUL) and Advanced Matrix Extension (AMX) , as taught by Intel. The modification would have been obvious because one of ordinary skill in the art would be motivated to deliver comprehensive set of functionality and higher performance of instruction set extensions to cover a diverse range of application domains and programming usages to users. See Intel at [1-1, top]: “The instruction set extensions cover a diverse range of application domains and programming usages. The 512-bit SIMD vector SIMD extensions, referred to as Intel® Advanced Vector Extensions 512 (Intel® AVX-512) instructions, deliver comprehensive set of functionality and higher performance than Intel® Advanced Vector Extensions (Intel® AVX) and Intel® Advanced Vector Extensions 2 (Intel® AVX2) instructions. Intel AVX, Intel AVX2 and many Intel AVX-512 instructions are covered in Intel® 64 and IA-32 Architectures Software Developer’s Manual sets. The reader can refer to them for basic and more background information related to various features referenced in this document.” Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIAYUE MA whose telephone number is (571)272-9658. The examiner can normally be reached between 9 am to 5 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Jiayue Ma/ Examiner, Art Unit 2126 /DAVID YI/Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Dec 28, 2023
Application Filed
Aug 24, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748965
WILDFIRE IGNITION PREDICTION WITH SWARM NEURAL NETWORK ENSEMBLE
4y 10m to grant Granted Sep 29, 2026
Patent 12737035
Gyst Technology
5y 7m to grant Granted Sep 15, 2026
Patent 12730556
Input Interface Display Method and Terminal
3y 5m to grant Granted Sep 08, 2026
Patent 12718120
RECOMMENDING MODEL CONTRIBUTIONS BASED ON FEDERATED LEARNING LINEAGE
4y 10m to grant Granted Aug 25, 2026
Patent 12694657
RECOGNITION, REIDENTIFICATION AND SECURITY ENHANCEMENTS USING AUTONOMOUS MACHINES
3y 9m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
69%
Grant Probability
99%
With Interview (+64.5%)
3y 11m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 258 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month