Prosecution Insights
Last updated: August 17, 2026
Application No. 17/122,598

TECHNIQUES FOR OPTIMIZING NEURAL NETWORKS

Non-Final OA §103
Filed
Dec 15, 2020
Examiner
SITIRICHE, LUIS A
Art Unit
2126
Tech Center
2100 — Computer Architecture & Software
Assignee
NVIDIA Corporation
OA Round
5 (Non-Final)
78%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
368 granted / 474 resolved
+22.6% vs TC avg
Strong +21% interview lift
Without
With
+21.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
12 currently pending
Career history
496
Total Applications
across all art units

Statute-Specific Performance

§101
23.2%
-16.8% vs TC avg
§103
40.9%
+0.9% vs TC avg
§102
13.6%
-26.4% vs TC avg
§112
12.9%
-27.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 474 resolved cases

Office Action

§103
DETAILED ACTION Claims 1, 9, 17, 25, 33 are amended. Claims 1-41 are pending (33-39 withdrawn). Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/08/2026 has been entered. Information Disclosure Statement The information disclosure statement filed 06/08/2026 fails to comply with 37 CFR 1.98(a)(2), which requires a legible copy of each cited foreign patent document; each non-patent literature publication or that portion which caused it to be listed; and all other information or that portion which caused it to be listed. Specifically, the 2nd NPL reference cited in the IDS by Wu Linyang is titled as “A Deep Learning Compilation Framework with Co-optimization of Operations and Data”, however, the filed copy by this author is titled “A computation and data unified compile architecture for deep learning”, therefore, it seems this is not the correct copy filed for the cited NPL. It has been placed in the application file, but the information referred to therein has not been considered. Response to Arguments The Applicant’s arguments regarding the rejection of above claims have been fully considered. In reference to Applicant’s arguments about: Rejections 35 USC 112 (a). Examiner’s response: Rejections are withdrawn in view of amendment and applicant’s arguments. In reference to Applicant’s arguments about: Rejections 35 USC 103. Examiner’s response: Applicant asserts that Khan and Zlateski fail to teach the limitations of amended independent claim 1 (and analogous claims 9, 17 and 25), however, examiner respectfully disagrees. Regarding Applicant’s arguments: “Khan describes an LSTM node that generates state data and can then store that data for as long as needed to be used by the LSTM node. The "previous value" for the state stored in a cell of the LSTM node in Khan would have been generated by that same LSTM node, and any subsequent use of that stored "previous value" would be by the same LSTM node. Khan does not teach caching data generated by an intermediate activation of one or more layers that is distinct from a portion of one or more layers that then uses the data at a first time step, and the data is then used again by yet another distinct portion of the one or more layers at a second time step”, Examiner would like to provide the broadest reasonable interpretation (BRI) of the claim limitations, specifically, the term “different/distinct portion”. Applicant asserts that Khan reuses the previous value by the same LSTM node, however, Examiner understands that Khan shows previous values (generated data) being stored as long as needed, and then decides how much of that information/values needs to be passed along to the future, which is interpreted to be re-used (used again) in future time steps. Based on the BRI of Khan’s teachings, Khan reusing the cached data in a future time step is broadly considered to be different to a previous time step. The claim limitation, as currently written, does not describe specifically what are the metes and bounds of “different portions” to determine how two portions are considered different from each other, therefore, Examiner still interprets Khan’s teachings to be analogous. In regards to the “cache”, even though Khan implicitly teaches “caching” as it stores data for future reuse, Zlateski teaches this cache buffer reuse for the purpose of memory bandwidth reduction, and it would be obvious to be combined with Khan’s approach of having past information from a previous state and deciding to remember it, or store it, in this cache from Zlateski, for later reuse or passing it along to the future, or subsequent time step. In addition, even though Examiner understands that Khan broadly teaches the claimed “different portions” as explained above, a new art is brought (Vinyals et al, previously cited as pertinent art in the last Office Action) to clarify and explicitly teach the reuse of cached data for different portions (see updated rejection below in view of the combination of Khan, Zlatestki, and Vinyal). Rejections are still maintained. In reference to Applicant’s arguments about: Rejections under 35 USC 101. Examiner’s response: Rejections are withdrawn in view of amendment and applicant’s arguments. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1, 4 – 7, 9, 12 – 14, 16 – 18, 22, 24 – 25, 29 – 30, 32 and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Khan et al (RNN-LSTM-GRU based language transformation) (hereafter referred to as Khan) in view of Zlateski et al (US Pub. No. 2020/0160181- hereinafter Zlateski) and further in view of Vinyals et al (US 2017/0200076- hereinafter Vinyals). Regarding claim 1, Khan discloses: A processor, comprising: one or more circuits to cause: generation of data for input to a portion of one or more layers or a neural network for use during a first time step, the generation performed by one or more intermediate activations of the one or more layers distinct from the portion of the one or more layers (Khan: “The deep neural network has shown a remarkable performance in many areas of computer technology. Neural networks are a core component of deep learning. The approach to build a neural network for machine translation is termed as neural machine translation.” [*Examiner note: i.e., the neural network is implemented by a computer to perform machine translation, which includes a processor] [Page 13010, 2.3 Deep Neural Network]. In addition, Khan teaches “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., storing previous values and deciding when to pass them along to the next hidden state (interpreted as the intermediate outputs of one or more layers distinct from the portion) is being interpreted as data previously generated to the current time step, interpreted as the first time step] [Page 13016, 3.5 LSTM and GRU cell].; the generated data to be cached and used by the portion of the one or more layer at the first time step (Khan teaches “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., storing previous values and deciding when to pass them along to the next hidden state (interpreted as the intermediate outputs of one or more layers) is being interpreted as caching and reusing data generated by one or more layers of a neural network] [Page 13016, 3.5 LSTM and GRU cell]), and the cached data to be reused by a different portion of the one or more layers at a second time step subsequent to the first time step, the different portion distinct from the one or more intermediate activations (Khan teaches “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., storing previous values and deciding when to pass them along to the next hidden state (interpreted as the intermediate outputs of one or more layers) is being interpreted as caching and reusing data generated by one or more layers of a neural network] [Page 13016, 3.5 LSTM and GRU cell]. In addition, Khan teaches “The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future” [*Examiner note: i.e., having past information from the previous state and deciding to remember it to determine how useful it is for a future time step is interpreted as cached for reuse by a different portion of the one or more layers at a subsequent time step, interpreted as the second time step, as future steps are interpreted to be different] [Page 13016, 3.5 LSTM and GRU cell]). Even though Khan implicitly teaches the “caching” for reuse, as explained above, Zlateski explicitly teaches it (Zlateski at [0100]: “Moreover, unlike prior approaches, the computation of some embodiments, within each pyramid task, can allow maximizing cache buffer reuse and/or reduction of memory bandwidth traffic, which can allow great savings in the amount of overall memory that needs to be used at any given point in the computation (e.g. a process may not need to store a whole layer's data in memory at the same time). This property can be a critical component enabling efficient execution of sparse CNNs”. [Examiner note: i.e., this cache buffer reuse for the purpose of memory bandwidth reduction would be obvious to be combined with Khan’s approach of having past information from a previous state and deciding to remember it, or store it, in this cache from Zlateski, for later reuse or passing it along to the future, or subsequent time step]). It would have obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan and Zlateski. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention in order to reduce memory bandwidth traffic, thereby allowing great saving savings in the amount of overall memory that needs to be used at any given point in the computation (Zlateski: “which can allow great savings in the amount of overall memory that needs to be used at any given point in the computation (e.g. a process may not need to store a whole layer's data in memory at the same time)” [0100]). Even though Khan implicitly teaches the cached data to be reused by a different portion of the one or more layers at a second time step, Vinyals teaches, in an analogous system, the cached data to be reused by a different portion of the one or more layers at a second time step (see Vinyals at [0029]: “For instance, each LSTM memory block can include one or more cells that each include an input gate, a forget gate, and an output gate that allow the cell to store previous states for the cell, e.g., for use in generating a current activation or to be provided to other components of the LSTM neural network”. Therefore, the reuse of the data in previous states to be provided to other components of the LSTM neural network is interpreted to the claimed ‘different portion of the one or more layers’). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski and Vinyals. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. One of ordinary skill would have motivation to combine Khan, Zlateski and Vinyals to optimize the training of the neural network by determining an optimal order for the target outputs that maximally simplifies a prescribed task, even where the optimal order is not known a priori (see Vinyals at [0044]: “However, because experimental results have demonstrated that the order in which the target outputs in a target output set can impact how effectively the recurrent neural network 302 is trained, the training system 300 is configured to determine an optimal order for the target outputs that maximally simplifies a prescribed task, even where the optimal order is not known a priori. The operations for training a recurrent neural network, including determining optimal orders for target output sets when training the recurrent neural network”). Referring to independent Claims 9, 17, and 25, they are rejected on the same basis as independent claim 1, mutatis mutandis, since they are analogous claims. Regarding claim 4, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 1 as shown in the rejection above. Khan also discloses: wherein the generated data is cached to store a first two or more results at the first time step, concurrently with a second two or more results from a previous time step (Khan: “RNN encoder encodes the input into a hidden sequence h = {h1, h2, …, hi}, where ht is a hidden state of the encoder at time step t” [*Examiner note: i.e., the hidden state ht is one result of the current time step t] [Page 13014, 3.3 Transliteration system] … “Hidden state at time step t is ht which represents input at same time step xt, modified by a weight vector U added to a previous hidden state ht-1 multiplied by its own hidden state matric V. The sum of weight input and hidden state is squashed by a function which is a logistic sigmoid activation function.” [*Examiner note: i.e., the weight vector U added to the previous hidden state is being interpreted as a second result of the previous time step] [Page 13014, 3.3 Transliteration system] … “Both LSTM and GRU are designed to keep the state from the previous activation rather than replacing the entire activation … the model is not washing out the new input every single time but keeps the relevant information and passes it down to the next time steps of the network.” [*Examiner note: i.e., the single time is the current time step, which is being interpreted as the previous time step. The next time step is being interpreted as a first time step] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 5, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 1 as shown in the rejection above. Khan also discloses: wherein the data includes two or more results generated by a gated recurrent unit (GRU) layer (Khan: “There are many variations in the structure of LSTM; one of them is called gated recurrent unit or GRU. GRU uses two gates: rest gate and update gate. The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future.” [*Examiner note: i.e., the result of the rest gate and the result of the update gate are being interpreted as two or more results generated by the GRU layer] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 6, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 1 as shown in the rejection above. Khan also discloses: wherein the data includes two or more results generated by a long short-term memory (LSTM) layer (Khan: “LSTM is not much different from RNNs except that they use a different function to compute hidden state in addition to the gating mechanism. It consists of three gates: an input gate, forget gate and an output gate to control the information to pass through.” [*Examiner note: i.e., the computed hidden state and the result of the output gate are being interpreted as two or more results of the current time step] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 7, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 1 as shown in the rejection above. Khan also discloses: wherein the one or more circuits are to cause the neural network to generate an inferencing result based, at least in part, on the reuse of the cached data (Khan: “The first network is used to encode the input sequence of the source language into a sequence vector which used by another network to decode that vector into an output sequence of the target language.” [*Examiner note: i.e., the another network to decode is being interpreted as one or more circuits causing the neural network to generate a sequence of the target language (i.e., an inferencing result) based on the output of the encoder (i.e., the neural network)] [Page 13014, 3.3 Transliteration System]. In addition, Khan teaches “The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future” [*Examiner note: i.e., having past information from the previous state and deciding to remember it to determine how useful it is for a future time step is interpreted as cached for reuse by a different portion of the one or more layers at a subsequent time step] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 12, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 9 as shown in the rejection above. Khan also discloses: wherein the one or more layers are convolutional gated recurrent unit (GRU) layers (Khan: “There are many variations in the structure of LSTM; one of them is called gated recurrent unit or GRU. GRU uses two gates: rest gate and update gate. The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future.” [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 13, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 9 as shown in the rejection above. Khan also discloses: wherein the one or more layers are long short-term memory (LSTM) layers (Khan: “LSTM is not much different from RNNs except that they use a different function to compute hidden state in addition to the gating mechanism. It consists of three gates: an input gate, forget gate and an output gate to control the information to pass through.” [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 14, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 9 as shown in the rejection above. Khan also discloses: wherein the cached data is based, at least in part, on an update gate weight and a reset gate weight (Khan: “GRU uses two gates: rest gate and update gate. The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future.” [*Examiner note: i.e., the rest gate is being interpreted as a reset gate, the update gate’s value is being interpreted as an update gate weight, and the rest gate’s value is being interpreted as a reset gate weight] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 16, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 9 as shown in the rejection above. Khan also discloses: wherein the one or more processors cause the neural network to generate an inferencing result based, at least in part, on the reuse of the cached data (Khan: “The first network is used to encode the input sequence of the source language into a sequence vector which used by another network to decode that vector into an output sequence of the target language.” [*Examiner note: i.e., the another network to decode is being interpreted as one or more circuits causing the neural network to generate a sequence of the target language (i.e., an inferencing result) based on the output of the encoder (i.e., the neural network)] [Page 13014, 3.3 Transliteration System]. In addition, Khan teaches “The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future” [*Examiner note: i.e., having past information from the previous state and deciding to remember it to determine how useful it is for a future time step is interpreted as cached for reuse by a different portion of the one or more layers at a subsequent time step] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 18, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 17 as shown in the rejection above. Khan also discloses: wherein the one or more layers are recurrent layers (Khan: “Our proposed model for machine transliteration is based on Seq2Seq deep learning model which uses bidirectional attention-based encoder and decoder architecture as illustrated in Fig. 4. Encoder–decoders are two separate recurrent neural networks.” [*Examiner note: i.e., the Encoder is a recurrent neural network, and the layers of the RNN encoder are being interpreted as recurrent layers] [Page 13014, 3.3 Transliteration System] … “The encoder encodes the input sequence into fixed dimension vector representation which helps to compute the new internal state at time step t using history of sequence vector ht-1” [*Examiner note: i.e., the fixed dimension vector representation is used to generate the new internal state and therefore does not dependent on the new internal state. Rather, the internal state depends on the fixed dimension vector. Thus, the fixed dimension vector representation being interpreted as state independent data] [Page 13014, 3.3 Transliteration System]). Regarding claim 22, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 17 as shown in the rejection above. Khan also discloses: wherein the one or more layers are gated recurrent unit (GRU) layers (Khan: “There are many variations in the structure of LSTM; one of them is called gated recurrent unit or GRU. GRU uses two gates: rest gate and update gate. The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future.” [*Examiner note: i.e., the result of the rest gate and the result of the update gate are being interpreted as two or more results generated by the GRU layer] [Page 13016, 3.5 LSTM and GRU cell]), and the cached data concurrently includes a first two or more results at the first time step and a second two or more results from a previous time step (Khan: “RNN encoder encodes the input into a hidden sequence h = {h1, h2, …, hi}, where ht is a hidden state of the encoder at time step t” [*Examiner note: i.e., the hidden state ht is one result of the current time step t] [Page 13014, 3.3 Transliteration system] … “Hidden state at time step t is ht which represents input at same time step xt, modified by a weight vector U added to a previous hidden state ht-1 multiplied by its own hidden state matric V. The sum of weight input and hidden state is squashed by a function which is a logistic sigmoid activation function.” [*Examiner note: i.e., the weight vector U added to the previous hidden state is being interpreted as a second result of the previous time step] [Page 13014, 3.3 Transliteration system] … “Both LSTM and GRU are designed to keep the state from the previous activation rather than replacing the entire activation … the model is not washing out the new input every single time but keeps the relevant information and passes it down to the next time steps of the network.” [*Examiner note: i.e., the single time is the current time step, which is being interpreted as the previous time step. The next time step is being interpreted as a first time step] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 24, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 17 as shown in the rejection above. Khan also discloses: wherein the instructions, which if performed by the one or more processors, further cause the one or more processors to cause the neural network to generate an inferencing result based, at least in part, on the reuse of the cached data (Khan: “The first network is used to encode the input sequence of the source language into a sequence vector which used by another network to decode that vector into an output sequence of the target language.” [*Examiner note: i.e., the decoder is being interpreted as one or more circuits that generate a sequence of the target language (i.e., an inferencing result) based on the output of the encoder (i.e., the neural network)] [Page 13014, 3.3 Transliteration System]. In addition, Khan teaches “The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future” [*Examiner note: i.e., having past information from the previous state and deciding to remember it to determine how useful it is for a future time step is interpreted as cached for reuse by a different portion of the one or more layers at a subsequent time step] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 29, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 25 as shown in the rejection above. Khan also discloses: wherein the data includes two or more results generated by a gated recurrent unit (GRU) layer (Khan: “There are many variations in the structure of LSTM; one of them is called gated recurrent unit or GRU. GRU uses two gates: rest gate and update gate. The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future.” [*Examiner note: i.e., the result of the rest gate and the result of the update gate are being interpreted as two or more results generated by the GRU layer] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 30, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 25 as shown in the rejection above. Khan also discloses: wherein the data includes two or more results generated by a long short-term memory (LSTM) layer (Khan: “LSTM is not much different from RNNs except that they use a different function to compute hidden state in addition to the gating mechanism. It consists of three gates: an input gate, forget gate and an output gate to control the information to pass through.” [*Examiner note: i.e., the computed hidden state and the result of the output gate are being interpreted as two or more results of the current time step] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 32, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 25 as shown in the rejection above. Khan also discloses: further comprising generating, by the neural network, an inferencing result based, at least in part, on the reuse of the cached data (Khan: “The first network is used to encode the input sequence of the source language into a sequence vector which used by another network to decode that vector into an output sequence of the target language.” [*Examiner note: i.e., the decoder is being interpreted as one or more circuits that generate a sequence of the target language (i.e., an inferencing result) based on the output of the encoder (i.e., the neural network)] [Page 13014, 3.3 Transliteration System]. In addition, Khan teaches “The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future” [*Examiner note: i.e., having past information from the previous state and deciding to remember it to determine how useful it is for a future time step is interpreted as cached for reuse by a different portion of the one or more layers at a subsequent time step] [Page 13016, 3.5 LSTM and GRU cell]). Regarding claim 40, the combination of Khan, Zlateski and Vinyals discloses the processor of claim 1, wherein the one or more intermediate activations are from one or more intermediate portions of the one or more layers (Khan teaches “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., storing previous values and deciding when to pass them along to the next hidden state (interpreted as the intermediate outputs from one or more intermediate portions) is being interpreted as caching and reusing data generated by one or more layers of a neural network] [Page 13016, 3.5 LSTM and GRU cell]). Claims 2, 10 – 11, 19 – 20, 26, and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Khan et al. (RNN-LSTM-GRU based language transformation) (hereafter referred to as Khan) in view of Zlateski et al (US Pub. No. 2020/0160181- hereinafter Zlateski), in view of Vinyals et al (US 2017/0200076- hereinafter Vinyals), and further in view of Goyal et al. (Recurrent Independent Mechanisms) (hereafter referred to as Goyal). Regarding claim 2, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 1 as shown in the rejection above. Khan also discloses: wherein the one or more layers are recurrent layers (Khan: “Our proposed model for machine transliteration is based on Seq2Seq deep learning model which uses bidirectional attention-based encoder and decoder architecture as illustrated in Fig. 4. Encoder–decoders are two separate recurrent neural networks.” [*Examiner note: i.e., the Encoder is a recurrent neural network, and the layers of the RNN encoder are being interpreted as recurrent layers] [Page 13014, 3.3 Transliteration System] … “The encoder encodes the input sequence into fixed dimension vector representation which helps to compute the new internal state at time step t using history of sequence vector ht-1” [*Examiner note: i.e., the fixed dimension vector representation is used to generate the new internal state and therefore does not dependent on the new internal state. Rather, the internal state depends on the fixed dimension vector. Thus, the fixed dimension vector representation being interpreted as state independent data] [Page 13014, 3.3 Transliteration System]). Khan fails to disclose: the data is state independent However, Goyal discloses: and the data is state independent (Goyal: “Our approach to modelling a dynamical system of interest divides the overall model into k small subsystems (or modules), each of which is recurrent in order to be able to capture the dynamics in the observed sequences. We refer to these subsystems as Recurrent Independent Mechanisms (RIMs), where each RIM has distinct functions that are learned automatically from data. We refer to RIM k at time step t as having vector-valued state ht,k” [*Examiner note: i.e., the data generating by the recurrent independent mechanisms is being interpreted as data that is state independent] [Page 2, 2 RIMs with Sparse Interactions] … “RIMs can be used as a drop-in replacement for an LSTM/GRU layer.” [*Examiner note: i.e., one could replace one or more layers of a LSTM/GRU neural network with a layer using recurrent independent mechanisms (i.e., data that is state independent] [Page 5, 2.4 Communication between RIMs]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Goyal. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Goyal teaches a recurrent layer with state independent data that can replace any LSTM or GRU layer. One of ordinary skill would be motivated to combine Khan, Zlateski, Vinyals and Goyal to replace one or more LSTM or GRU layers of Khan with one or more of the independent recurrent layers taught by Goyal. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because not all data is improved from reliance on other states, as well as being computationally intensive to always treat data as dependent (Goyal: “Many systems of interest comprise multiple dynamical processes that operate relatively independently and only occasionally have meaningful interactions. Despite this, most machine learning models employ the opposite inductive bias, i.e., that all processes interact. This can lead to poor generalization and lack of robustness to changing task distributions. We have proposed a new architecture, Recurrent Independent Mechanisms (RIMs), in which we learn multiple recurrent modules that are independent by default, but interact sparingly.” [Page 9, 5 Conclusion]). Regarding claim 10, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 9 as shown in the rejection above. Khan also discloses: wherein the one or more layers are recurrent layers (Khan: “Our proposed model for machine transliteration is based on Seq2Seq deep learning model which uses bidirectional attention-based encoder and decoder architecture as illustrated in Fig. 4. Encoder–decoders are two separate recurrent neural networks.” [*Examiner note: i.e., the Encoder is a recurrent neural network, and the layers of the RNN encoder are being interpreted as recurrent layers] [Page 13014, 3.3 Transliteration System] … “The encoder encodes the input sequence into fixed dimension vector representation which helps to compute the new internal state at time step t using history of sequence vector ht-1” [*Examiner note: i.e., the fixed dimension vector representation is used to generate the new internal state and therefore does not dependent on the new internal state. Rather, the internal state depends on the fixed dimension vector. Thus, the fixed dimension vector representation being interpreted as state independent data] [Page 13014, 3.3 Transliteration System]). Khan fails to disclose that the data is state independent However, Goyal discloses: and the data is state independent (Goyal: “Our approach to modelling a dynamical system of interest divides the overall model into k small subsystems (or modules), each of which is recurrent in order to be able to capture the dynamics in the observed sequences. We refer to these subsystems as Recurrent Independent Mechanisms (RIMs), where each RIM has distinct functions that are learned automatically from data. We refer to RIM k at time step t as having vector-valued state ht,k” [*Examiner note: i.e., the data generating by the recurrent independent mechanisms is being interpreted as data that is state independent] [Page 2, 2 RIMs with Sparse Interactions] … “RIMs can be used as a drop-in replacement for an LSTM/GRU layer.” [*Examiner note: i.e., one could replace one or more layers of a LSTM/GRU neural network with a layer using recurrent independent mechanisms (i.e., data that is state independent] [Page 5, 2.4 Communication between RIMs]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Goyal. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Goyal teaches a recurrent layer with state independent data that can replace any LSTM or GRU layer. One of ordinary skill would be motivated to combine Khan, Zlateski, Vinyals and Goyal to replace one or more LSTM or GRU layers of Khan with one or more of the independent recurrent layers taught by Goyal. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because not all data is improved from reliance on other states, as well as being computationally intensive to always treat data as dependent (Goyal: “Many systems of interest comprise multiple dynamical processes that operate relatively independently and only occasionally have meaningful interactions. Despite this, most machine learning models employ the opposite inductive bias, i.e., that all processes interact. This can lead to poor generalization and lack of robustness to changing task distributions. We have proposed a new architecture, Recurrent Independent Mechanisms (RIMs), in which we learn multiple recurrent modules that are independent by default, but interact sparingly.” [Page 9, 5 Conclusion]). Regarding claim 11, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 9 as shown in the rejection above. Khan also discloses: wherein the one or more layers generate state dependent data [and state independent data] (Khan: “Hidden state at time step t is ht which represents input at same time step xt, modified by a weight vector U added to a previous hidden state ht-1 multiplied by its own hidden state matric V. The sum of weight input and hidden state is squashed by a function which is a logistic sigmoid activation function.” [*Examiner note: i.e., the weight input depends on the previous hidden state ht-1 and therefore is being interpreted as state dependent data] [Page 13014, 3.3 Transliteration system]), Khan fails to disclose: wherein the one or more layers generate [state dependent data and] state independent data and the cached data is state independent data generated by the one or more layers. However, Goyal discloses: wherein the one or more layers generate [state dependent data and] state independent data (Goyal: “Our approach to modelling a dynamical system of interest divides the overall model into k small subsystems (or modules), each of which is recurrent to be able to capture the dynamics in the observed sequences. We refer to these subsystems as Recurrent Independent Mechanisms (RIMs), where each RIM has distinct functions that are learned automatically from data. We refer to RIM k at time step t as having vector-valued state ht,k” [*Examiner note: i.e., the data generating by the recurrent independent mechanisms is being interpreted as data that is state independent] [Page 2, 2 RIMs with Sparse Interactions] … “RIMs can be used as a drop-in replacement for an LSTM/GRU layer.” [*Examiner note: i.e., one could replace one or more layers of a LSTM/GRU neural network with a layer using recurrent independent mechanisms (i.e., data that is state independent] [Page 5, 2.4 Communication between RIMs]). and the cached data is state independent data generated by the one or more layers (Goyal: “Although the RIMs operate independently by default, the attention mechanism allows sharing of information among the RIMs. Specifically, we allow the activated RIMs to read from all other RIMs (activated or not). The intuition behind this is that non-activated RIMs are not related to the current input, so their value needs not change. However they may still store contextual information relevant for activated RIMs later on.” [*Examiner note: stored contextual information relevant for activated RIMs is being interpreted as cached state independent data] [Page 5, 2.4 Communication between RIMs]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Goyal. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Goyal teaches a recurrent layer with state independent data that can replace any LSTM or GRU layer. One of ordinary skill would be motivated to combine Khan, Zlateski, Vinyals and Goyal to replace one or more LSTM or GRU layers of Khan with one or more of the independent recurrent layers taught by Goyal. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because not all data is improved from reliance on other states, as well as being computationally intensive to always treat data as dependent (Goyal: “Many systems of interest comprise multiple dynamical processes that operate relatively independently and only occasionally have meaningful interactions. Despite this, most machine learning models employ the opposite inductive bias, i.e., that all processes interact. This can lead to poor generalization and lack of robustness to changing task distributions. We have proposed a new architecture, Recurrent Independent Mechanisms (RIMs), in which we learn multiple recurrent modules that are independent by default, but interact sparingly.” [Page 9, 5 Conclusion]). Regarding claim 19, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 17 as shown in the rejection above. Khan also discloses: the one or more layers are convolutional gated recurrent unit (GRU) layers (Khan: “There are many variations in the structure of LSTM; one of them is called gated recurrent unit or GRU. GRU uses two gates: rest gate and update gate. The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future.” [*Examiner note: i.e., the result of the rest gate and the result of the update gate are being interpreted as two or more results generated by the GRU layer] [Page 13016, 3.5 LSTM and GRU cell]). Khan fails to disclose: wherein the cached data is state-independent data However, Goyal discloses: wherein the cached data is state-independent data (Goyal: “Although the RIMs operate independently by default, the attention mechanism allows sharing of information among the RIMs. Specifically, we allow the activated RIMs to read from all other RIMs (activated or not). The intuition behind this is that non-activated RIMs are not related to the current input, so their value needs not change. However they may still store contextual information relevant for activated RIMs later on.” [*Examiner note: stored contextual information relevant for activated RIMs is being interpreted as cached state independent data] [Page 5, 2.4 Communication between RIMs]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Goyal. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Goyal teaches a recurrent layer with state independent data that can replace any LSTM or GRU layer. One of ordinary skill would be motivated to combine Khan, Zlateski, Vinyals and Goyal to replace one or more LSTM or GRU layers of Khan with one or more of the independent recurrent layers taught by Goyal. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because not all data is improved from reliance on other states, as well as being computationally intensive to always treat data as dependent (Goyal: “Many systems of interest comprise multiple dynamical processes that operate relatively independently and only occasionally have meaningful interactions. Despite this, most machine learning models employ the opposite inductive bias, i.e., that all processes interact. This can lead to poor generalization and lack of robustness to changing task distributions. We have proposed a new architecture, Recurrent Independent Mechanisms (RIMs), in which we learn multiple recurrent modules that are independent by default, but interact sparingly.” [Page 9, 5 Conclusion]). Regarding claim 20, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 17 as shown in the rejection above. Khan also discloses: the one or more layers are long short-term memory (LSTM) layers (Khan: “LSTM is not much different from RNNs except that they use a different function to compute hidden state in addition to the gating mechanism. It consists of three gates: an input gate, forget gate and an output gate to control the information to pass through.” [*Examiner note: i.e., the computed hidden state and the result of the output gate are being interpreted as two or more results of the current time step] [Page 13016, 3.5 LSTM and GRU cell]). Khan fails to disclose: wherein the cached data is state-independent data. However, Goyal discloses: wherein the cached data is state-independent data (Goyal: “Although the RIMs operate independently by default, the attention mechanism allows sharing of information among the RIMs. Specifically, we allow the activated RIMs to read from all other RIMs (activated or not). The intuition behind this is that non-activated RIMs are not related to the current input, so their value needs not change. However they may still store contextual information relevant for activated RIMs later on.” [*Examiner note: stored contextual information relevant for activated RIMs is being interpreted as cached state independent data] [Page 5, 2.4 Communication between RIMs]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Goyal. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Goyal teaches a recurrent layer with state independent data that can replace any LSTM or GRU layer. One of ordinary skill would have motivation to combine Khan, Zlateski, Vinyals and Goyal to replace one or more LSTM or GRU layers of Khan with one or more of the independent recurrent layers taught by Goyal. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because not all data is improved from reliance on other states, as well as being computationally intensive to always treat data as dependent (Goyal: “Many systems of interest comprise multiple dynamical processes that operate relatively independently and only occasionally have meaningful interactions. Despite this, most machine learning models employ the opposite inductive bias, i.e., that all processes interact. This can lead to poor generalization and lack of robustness to changing task distributions. We have proposed a new architecture, Recurrent Independent Mechanisms (RIMs), in which we learn multiple recurrent modules that are independent by default, but interact sparingly.” [Page 9, 5 Conclusion]). Regarding claim 26, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 25 as shown in the rejection above. Khan also discloses: wherein the one or layers are recurrent layers (Khan: “Our proposed model for machine transliteration is based on Seq2Seq deep learning model which uses bidirectional attention-based encoder and decoder architecture as illustrated in Fig. 4. Encoder–decoders are two separate recurrent neural networks.” [*Examiner note: i.e., the Encoder is a recurrent neural network, and the layers of the RNN encoder are being interpreted as recurrent layers] [Page 13014, 3.3 Transliteration System] … “The encoder encodes the input sequence into fixed dimension vector representation which helps to compute the new internal state at time step t using history of sequence vector ht-1” [*Examiner note: i.e., the fixed dimension vector representation is used to generate the new internal state and therefore does not dependent on the new internal state. Rather, the internal state depends on the fixed dimension vector. Thus, the fixed dimension vector representation being interpreted as state independent data] [Page 13014, 3.3 Transliteration System]) Khan fails to disclose: the data is state independent. However, Goyal discloses: and the data is state independent (Goyal: “Our approach to modelling a dynamical system of interest divides the overall model into k small subsystems (or modules), each of which is recurrent in order to be able to capture the dynamics in the observed sequences. We refer to these subsystems as Recurrent Independent Mechanisms (RIMs), where each RIM has distinct functions that are learned automatically from data. We refer to RIM k at time step t as having vector-valued state ht,k” [*Examiner note: i.e., the data generating by the recurrent independent mechanisms is being interpreted as data that is state independent] [Page 2, 2 RIMs with Sparse Interactions] … “RIMs can be used as a drop-in replacement for an LSTM/GRU layer.” [*Examiner note: i.e., one could replace one or more layers of a LSTM/GRU neural network with a layer using recurrent independent mechanisms (i.e., data that is state independent] [Page 5, 2.4 Communication between RIMs]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Goyal. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Goyal teaches a recurrent layer with state independent data that can replace any LSTM or GRU layer. One of ordinary skill would be motivated to combine Khan, Zlateski, Vinyals and Goyal to replace one or more LSTM or GRU layers of Khan with one or more of the independent recurrent layers taught by Goyal. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because not all data is improved from reliance on other states, as well as being computationally intensive to always treat data as dependent (Goyal: “Many systems of interest comprise multiple dynamical processes that operate relatively independently and only occasionally have meaningful interactions. Despite this, most machine learning models employ the opposite inductive bias, i.e., that all processes interact. This can lead to poor generalization and lack of robustness to changing task distributions. We have proposed a new architecture, Recurrent Independent Mechanisms (RIMs), in which we learn multiple recurrent modules that are independent by default, but interact sparingly.” [Page 9, 5 Conclusion]). Regarding claim 28, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 25 as shown in the rejection above. The combination of Khan and Zlateski fails to disclose: wherein the cached data includes a first two or more state independent results for the first time step and a second two or more state independent results from a previous time step. However, Goyal discloses: wherein the cached data includes a first two or more state independent results for the first time step and a second two or more state independent results from a previous time step (Goyal: “We refer to these subsystems as Recurrent Independent Mechanisms (RIMs), where each RIM has distinct functions that are learned automatically from data. We refer to RIM k at time step t as having vector-valued state ht,k , where t = 1, …, T” [*Examiner note: i.e., timestep t = 1 is being interpreted as a first time step, time step t = 2 is being interpreted as a second time step. The vector-valued state is being interpreted as one state independent result] [Page 2, 2 RIMs with Sparse Interactions] … we want each RIM to have its own independent dynamics operating by default, and occasionally to interact with other relevant RIMs and selected elements of the encoded input… and operate on few key/value variables at a time selected using an attention mechanism” [*Examiner note: at each time step, the layer with RIMs generates output by operating on selected elements (i.e., a second state independent result)] [Pages 2 – 3, 2 RIMs with Sparse Interactions]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Goyal. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Goyal teaches a recurrent layer with state independent data that can replace any LSTM or GRU layer. One of ordinary skill would have motivation to combine Khan, Zlateski and Goyal to replace one or more LSTM or GRU layers of Khan with one or more of the independent recurrent layers taught by Goyal. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because not all data is improved from reliance on other states, as well as being computationally intensive to always treat data as dependent (Goyal: “Many systems of interest comprise multiple dynamical processes that operate relatively independently and only occasionally have meaningful interactions. Despite this, most machine learning models employ the opposite inductive bias, i.e., that all processes interact. This can lead to poor generalization and lack of robustness to changing task distributions. We have proposed a new architecture, Recurrent Independent Mechanisms (RIMs), in which we learn multiple recurrent modules that are independent by default, but interact sparingly.” [Page 9, 5 Conclusion]). Claims 3, 8, 15, 21, 23, 27, and 31 are rejected under 35 U.S.C. 103 as being unpatentable over Khan et al. (RNN-LSTM-GRU based language transformation) (hereafter referred to as Khan) in view of Zlateski et al (US Pub. No. 2020/0160181- hereinafter Zlateski), in view of Vinyals et al (US 2017/0200076- hereinafter Vinyals) and further in view of Rishi et al. (WO 2019002880 A1) (hereafter referred to as Rishi). Regarding claim 3, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 1 as shown in the rejection above. Khan also discloses: wherein the generated data is cached in one or more [circular] buffers (Khan: “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., the cell storing previous values and deciding when to pass them along to the next hidden state is being interpreted as a buffer] [Page 13016, 3.5 LSTM and GRU cell]. Further, Zlateski teaches at [0100]: “Moreover, unlike prior approaches, the computation of some embodiments, within each pyramid task, can allow maximizing cache buffer reuse and/or reduction of memory bandwidth traffic”). The combination of Khan, Zlateski, Vinyals fails to disclose: one or more circular buffers. However, Rishi discloses: one or more circular buffers (Rishi: “the feedback loop comprises a circular buffer and/or a state buffer” [*Examiner note: emphasis added] [Page 6, Paragraph 2] … “the one or more neural networks comprises any of Long: Short Term Memory (LSTM) neural networks; Recurrent neural networks (RNN), Gated Recurrent Unit (GRU); internal state machines; and/or circular buffers” [Page 7, Paragraph 1]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Rishi. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Rishi teaches using a circular buffer with a LSTM or GRU neural network. One of ordinary skill would have motivation to combine Khan, Zlateski and Rishi to replace the buffer taught by Khan with one or more circular buffers taught by Rishi. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because circular buffers improve the speed of the neural network (Rishi: “Circular buffers are capable of increasing the speed of within the neural network by storing states over a buffer and implements a feedback loop.” [Page 6, Paragraph 3]). Regarding claim 8, the combination of Khan, Zlateski and Vinyals discloses all the limitations of claim 1 as shown in the rejection above. Khan also discloses: wherein the one or more layers are recurrent layers, the cached data concurrently includes a first two or more results for the first time step and a second two or more results from a previous time step (Khan: “RNN encoder encodes the input into a hidden sequence h = {h1, h2, …, hi}, where ht is a hidden state of the encoder at time step t” [*Examiner note: i.e., the hidden state ht is one result of the current time step t] [Page 13014, 3.3 Transliteration system] … “Hidden state at time step t is ht which represents input at same time step xt, modified by a weight vector U added to a previous hidden state ht-1 multiplied by its own hidden state matric V. The sum of weight input and hidden state is squashed by a function which is a logistic sigmoid activation function.” [*Examiner note: i.e., the weight input is being interpreted as a second result of the current time step] [Page 13014, 3.3 Transliteration system] … “Both LSTM and GRU are designed to keep the state from the previous activation rather than replacing the entire activation … the model is not washing out the new input every single time but keeps the relevant information and passes it down to the next time steps of the network.” [*Examiner note: i.e., the single time is the current time step, which is being interpreted as a previous time step. The next time step is being interpreted as a first time step] [Page 13016, 3.5 LSTM and GRU cell]), and wherein the cached data, including the first two or more results and the second two or more results, is stored in a single [circular] buffer (Khan: “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values” [*Examiner note: i.e., the cell is being interpreted as a single buffer that stores the previous values (i.e., the two or more results) for each time step] [Page 13016, 3.5 LSTM and GRU cell]). Khan fails to disclose: one or more circular buffers. However, Rishi discloses: one or more circular buffers (Rishi: “the feedback loop comprises a circular buffer and/or a state buffer” [*Examiner note: emphasis added] [Page 6, Paragraph 2] … “the one or more neural networks comprises any of Long: Short Term Memory (LSTM) neural networks; Recurrent neural networks (RNN), Gated Recurrent Unit (GRU); internal state machines; and/or circular buffers” [Page 7, Paragraph 1]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Rishi. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Rishi teaches using a circular buffer with a LSTM or GRU neural network. One of ordinary skill would be motivated to combine Khan, Zlateski and Rishi to replace the buffer taught by Khan with one or more circular buffers taught by Rishi. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because circular buffers improve the speed of the neural network (Rishi: “Circular buffers are capable of increasing the speed of within the neural network by storing states over a buffer and implements a feedback loop.” [Page 6, Paragraph 3]). Regarding claim 15, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 9 as shown in the rejection above. Khan also discloses: wherein the cached data concurrently includes a first two or more results for the first time step and a second two or more results from a previous time step (Khan: “RNN encoder encodes the input into a hidden sequence h = {h1, h2, …, hi}, where ht is a hidden state of the encoder at time step t” [*Examiner note: i.e., the hidden state ht is one result of the current time step t] [Page 13014, 3.3 Transliteration system] … “Hidden state at time step t is ht which represents input at same time step xt, modified by a weight vector U added to a previous hidden state ht-1 multiplied by its own hidden state matric V. The sum of weight input and hidden state is squashed by a function which is a logistic sigmoid activation function.” [*Examiner note: i.e., the weight input is being interpreted as a second result of the previous time step] [Page 13014, 3.3 Transliteration system]), and wherein the first two or more results are concatenated to form a first concatenated tensor, the second two or more results are concatenated to form a second concatenated tensor (Khan: “The sum of weight input and hidden state is squashed by a function which is a logistic sigmoid activation function. The encoder encodes the input sequence into fixed dimension vector representation which helps to compute the new internal state at time step t” [*Examiner note: i.e., for each time step (i.e., for both the first and the second timestep), the two or more results are concatenated to form a vector representation. A vector is a tensor.] [Page 13014, 3.3 Transliteration system] … “Both LSTM and GRU are designed to keep the state from the previous activation rather than replacing the entire activation … the model is not washing out the new input every single time but keeps the relevant information and passes it down to the next time steps of the network.” [*Examiner note: i.e., the single time is the current time step, which is being interpreted as a first time step. The next time step is being interpreted as a second time step] [Page 13016, 3.5 LSTM and GRU cell]), and the first concatenated tensor and the second concatenated tensor are stored in a single [circular] buffer (Khan: “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., the cell storing previous values and deciding when to pass them along to the next hidden state is being interpreted as a buffer] [Page 13016, 3.5 LSTM and GRU cell]). Khan fails to disclose: a single circular buffer. However, Rishi discloses: a single circular buffer (Rishi: “the feedback loop comprises a circular buffer and/or a state buffer” [*Examiner note: emphasis added] [Page 6, Paragraph 2] … “the one or more neural networks comprises any of Long: Short Term Memory (LSTM) neural networks; Recurrent neural networks (RNN), Gated Recurrent Unit (GRU); internal state machines; and/or circular buffers” [Page 7, Paragraph 1]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Rishi. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Rishi teaches using a circular buffer with a LSTM or GRU neural network. One of ordinary skill would be motivated to combine Khan, Zlateski and Rishi to replace the buffer taught by Khan with one or more circular buffers taught by Rishi. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because circular buffers improve the speed of the neural network (Rishi: “Circular buffers are capable of increasing the speed of within the neural network by storing states over a buffer and implements a feedback loop.” [Page 6, Paragraph 3]). Regarding claim 21, the combination of Khan and Zlateski discloses all the limitations of claim 17 as shown in the rejection above. Khan also discloses: wherein the instructions, which if performed by the one or more processors, further cause the one or more processors to cache the generated data in one or more [circular] buffers (Khan: “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., the cell storing previous values and deciding when to pass them along to the next hidden state is being interpreted as a buffer] [Page 13016, 3.5 LSTM and GRU cell]). Khan fails to disclose: one or more circular buffers. However, Rishi discloses: one or more circular buffers (Rishi: “the feedback loop comprises a circular buffer and/or a state buffer” [*Examiner note: emphasis added] [Page 6, Paragraph 2] … “the one or more neural networks comprises any of Long: Short Term Memory (LSTM) neural networks; Recurrent neural networks (RNN), Gated Recurrent Unit (GRU); internal state machines; and/or circular buffers” [Page 7, Paragraph 1]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Rishi. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Rishi teaches using a circular buffer with a LSTM or GRU neural network. One of ordinary skill would be motivated to combine Khan, Zlateski and Rishi to replace the buffer taught by Khan with one or more circular buffers taught by Rishi. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because circular buffers improve the speed of the neural network (Rishi: “Circular buffers are capable of increasing the speed of within the neural network by storing states over a buffer and implements a feedback loop.” [Page 6, Paragraph 3]). Regarding claim 23, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 17 as shown in the rejection above. Khan also discloses: wherein the cached data concurrently includes a first two or more results at the first time step and a second two or more results from a previous time step (Khan: “RNN encoder encodes the input into a hidden sequence h = {h1, h2, …, hi}, where ht is a hidden state of the encoder at time step t” [*Examiner note: i.e., the hidden state ht is one result of the current time step t] [Page 13014, 3.3 Transliteration system] … “Hidden state at time step t is ht which represents input at same time step xt, modified by a weight vector U added to a previous hidden state ht-1 multiplied by its own hidden state matric V. The sum of weight input and hidden state is squashed by a function which is a logistic sigmoid activation function.” [*Examiner note: i.e., the weight input is being interpreted as a second result of the current time step] [Page 13014, 3.3 Transliteration system] … “Both LSTM and GRU are designed to keep the state from the previous activation rather than replacing the entire activation … the model is not washing out the new input every single time but keeps the relevant information and passes it down to the next time steps of the network.” [*Examiner note: i.e., the single time is the current time step, which is being interpreted as a previous time step. The next time step is being interpreted as a first time step] [Page 13016, 3.5 LSTM and GRU cell]), and wherein the instructions, which if performed by the one or more processors, further cause the one or more processors to cache the data in one or more [circular] buffers (Khan: “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., the cell storing previous values and deciding when to pass them along to the next hidden state is being interpreted as a buffer] [Page 13016, 3.5 LSTM and GRU cell]). Khan fails to disclose: one or more circular buffers. However, Rishi discloses: one or more circular buffers (Rishi: “the feedback loop comprises a circular buffer and/or a state buffer” [*Examiner note: emphasis added] [Page 6, Paragraph 2] … “the one or more neural networks comprises any of Long: Short Term Memory (LSTM) neural networks; Recurrent neural networks (RNN), Gated Recurrent Unit (GRU); internal state machines; and/or circular buffers” [Page 7, Paragraph 1]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Rishi. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Rishi teaches using a circular buffer with a LSTM or GRU neural network. One of ordinary skill would have motivation to combine Khan, Zlateski and Rishi to replace the buffer taught by Khan with one or more circular buffers taught by Rishi. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because circular buffers improve the speed of the neural network (Rishi: “Circular buffers are capable of increasing the speed of within the neural network by storing states over a buffer and implements a feedback loop.” [Page 6, Paragraph 3]). Regarding claim 27, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 25 as shown in the rejection above. Khan also discloses: wherein the one or more layers are recurrent layers (Khan: “Our proposed model for machine transliteration is based on Seq2Seq deep learning model which uses bidirectional attention-based encoder and decoder architecture as illustrated in Fig. 4. Encoder–decoders are two separate recurrent neural networks.” [*Examiner note: i.e., the Encoder is a recurrent neural network, and the layers of the RNN encoder are being interpreted as recurrent layers] [Page 13014, 3.3 Transliteration System] … “The encoder encodes the input sequence into fixed dimension vector representation which helps to compute the new internal state at time step t using history of sequence vector ht-1” [*Examiner note: i.e., the fixed dimension vector representation is used to generate the new internal state and therefore does not dependent on the new internal state. Rather, the internal state depends on the fixed dimension vector. Thus, the fixed dimension vector representation being interpreted as state independent data] [Page 13014, 3.3 Transliteration System]), and caching data includes caching the data in one or more [circular] buffers (Khan: “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., the cell storing previous values and deciding when to pass them along to the next hidden state is being interpreted as a buffer] [Page 13016, 3.5 LSTM and GRU cell]). Khan fails to disclose: one or more circular buffers. However, Rishi discloses: one or more circular buffers (Rishi: “the feedback loop comprises a circular buffer and/or a state buffer” [*Examiner note: emphasis added] [Page 6, Paragraph 2] … “the one or more neural networks comprises any of Long: Short Term Memory (LSTM) neural networks; Recurrent neural networks (RNN), Gated Recurrent Unit (GRU); internal state machines; and/or circular buffers” [Page 7, Paragraph 1]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Rishi. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Rishi teaches using a circular buffer with a LSTM or GRU neural network. One of ordinary skill would be motivated to combine Khan, Zlateski and Rishi to replace the buffer taught by Khan with one or more circular buffers taught by Rishi. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because circular buffers improve the speed of the neural network (Rishi: “Circular buffers are capable of increasing the speed of within the neural network by storing states over a buffer and implements a feedback loop.” [Page 6, Paragraph 3]). Regarding claim 31, the combination of Khan, Zlateski, Vinyals discloses all the limitations of claim 25 as shown in the rejection above. Khan also discloses: wherein the one or more layers are recurrent layers, the generated data is cached to store a first two or more results at the first time step concurrently with a second two or more results from a previous time step, and wherein caching data includes caching the data in a single [circular] buffer (Khan: “LSTMs have a cell that stores the previous values and hold onto it unless a forget gate tells the cell to forget those values. In another way, it keeps the previous iterations as long as needed. Input gate adds a new input to the cell, and an output gate decides when to pass along the vectors from the cell to the next hidden state.” [*Examiner note: i.e., the cell storing previous values and deciding when to pass them along to the next hidden state is being interpreted as a single buffer] [Page 13016, 3.5 LSTM and GRU cell]). Khan fails to disclose: a single circular buffer. However, Rishi discloses: a single circular buffer (Rishi: “the feedback loop comprises a circular buffer and/or a state buffer” [*Examiner note: emphasis added] [Page 6, Paragraph 2] … “the one or more neural networks comprises any of Long: Short Term Memory (LSTM) neural networks; Recurrent neural networks (RNN), Gated Recurrent Unit (GRU); internal state machines; and/or circular buffers” [Page 7, Paragraph 1]) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Rishi. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Rishi teaches using a circular buffer with a LSTM or GRU neural network. One of ordinary skill would be motivated to combine Khan, Zlateski and Rishi to replace the buffer taught by Khan with one or more circular buffers taught by Rishi. One having ordinary skill in the art would have been motivated to make this change before the effective filing date of the claimed invention because circular buffers improve the speed of the neural network (Rishi: “Circular buffers are capable of increasing the speed of within the neural network by storing states over a buffer and implements a feedback loop.” [Page 6, Paragraph 3]). Claim 41 is rejected under 35 U.S.C. 103 as being unpatentable over Khan et al. (RNN-LSTM-GRU based language transformation) (hereafter referred to as Khan) in view of Zlateski et al (US Pub. No. 2020/0160181- hereinafter Zlateski) in view of Vinyals et al (US 2017/0200076- hereinafter Vinyals) and further in view of Bacchiani et al. (US 9,620,145) (hereafter referred to as Bacchiani). Regarding Claim 41, the combination of Khan, Zlateski and Vinyals discloses the processor of claim 1, however, fails to teach wherein the one or more intermediate activations include state independent data for use across multiple time steps. Bacchiani teaches, in an analogous system, wherein the one or more intermediate activations include state independent data for use across multiple time steps (see Bacchiani at Claim 11: “a first neural network that has been trained using speech examples that are each assigned a context-dependent state by clustering context independent states based on activations at a hidden layer of a second neural network that was trained to provide outputs corresponding to context-independent states, the first neural network being configured to provide outputs corresponding to one or more context-dependent states” [*Examiner note: the activation of the second neural network’s outputs corresponding to context-independent states is analogous to the claimed “intermediate outputs include state independent data”]. Further, Khan teaches, as explained in the rejection of claim 1, the data across multiple time steps: “The rest gate is used in a model to decide how much of the past information needs to forget or remember. It uses the information in previous state and the next input candidate to decide the information to forget. The update gate helps the model to determine how much of the past information from previous time steps needs to be passed along to the future” [*Examiner note: i.e., having past information from the previous state and deciding to remember it to determine how useful it is for a future time step is interpreted as data across multiple time steps] [Page 13016, 3.5 LSTM and GRU cell])). It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Khan, Zlateski, Vinyals and Bacchiani. Khan teaches a neural network capable of deciding how much past data to remember generated by one or more LSTM or GRU layers. Zlateski teaches the use of efficient cache buffer. Vinyals teaches reusing the stored data from previous state for use in other components of the LSTM. Bacchiani teaches wherein the one or more intermediate outputs include state independent data generated by one or more intermediate activations. One of ordinary skill would have motivation to combine Khan, Zlateski and Bacchiani to provide an enhanced speech recognition engine and provide accurate transcriptions of an utterance (see Bacchiani at Column 1: Background “Automatic speech recognition is an important technology that can be used in mobile devices and other devices. In general, automatic speech recognition attempts to provide accurate transcriptions of what a person has said”). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to LUIS A SITIRICHE whose telephone number is (571)270-1316. The examiner can normally be reached M-F 9am-6pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LUIS A SITIRICHE/Primary Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

Show 6 earlier events
May 27, 2025
Request for Continued Examination
May 29, 2025
Response after Non-Final Action
Jun 10, 2025
Non-Final Rejection mailed — §103
Sep 10, 2025
Response Filed
Nov 26, 2025
Final Rejection mailed — §103
Feb 26, 2026
Request for Continued Examination
Mar 09, 2026
Response after Non-Final Action
Aug 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705301
ANALYSIS OF CLUSTERED DATA
4y 2m to grant Granted Aug 11, 2026
Patent 12682276
METHODS AND SYSTEMS FOR PROCESSING UNSTRUCTURED AND UNLABELLED DATA
4y 9m to grant Granted Jul 14, 2026
Patent 12670232
MODEL PREDICTION CONFIDENCE UTILIZING DRIFT
4y 8m to grant Granted Jun 30, 2026
Patent 12664481
Systems and Methods for Predictive Coding Utilizing Confidence Levels
5y 2m to grant Granted Jun 23, 2026
Patent 12651040
SAMPLING FROM A SET SPINS WITH CLAMPING
4y 6m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+21.4%)
3y 7m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 474 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month