Prosecution Insights
Last updated: October 01, 2026
Application No. 18/697,211

USING MEMORY TO AUGMENT SELF-ATTENTION IN NEURAL NETWORKS

Non-Final OA §103
Filed
Mar 29, 2024
Priority
Oct 06, 2021 — provisional 63/252,616 +1 more
Examiner
BRAHMACHARI, MANDRITA
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
77%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
324 granted / 422 resolved
+16.8% vs TC avg
Strong +29% interview lift
Without
With
+28.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
26 currently pending
Career history
444
Total Applications
across all art units

Statute-Specific Performance

§101
12.1%
-27.9% vs TC avg
§103
57.4%
+17.4% vs TC avg
§102
6.1%
-33.9% vs TC avg
§112
17.4%
-22.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 422 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION The action is in response to claims dated 3/29/2024 Claims pending in the case: 1-20, 22-29 Cancelled claims: 21 Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 7, 11-14, 17, 22-25, 27 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gupta (Memory-efficient Transformers via Top-k Attention). Regarding Claim 1, Gupta teaches, a system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to implement a neural network (Gupta: Pg. 1 section 1 [1-2]: transformer neural network; Pg. 5 section 3: benchmarking – computer implementation) configured to process a network input to generate an output sequence comprising a respective output element at each of multiple output positions in an output order (Gupta: Pg. 2 [1]: computing output one chunk at a time), the output sequence being partitioned into a plurality of subsequences each comprising a different subset of the output elements at respective output positions (Gupta: Pg. 2 [1]: computing output one chunk at a time), the generating comprising, for each subsequence and for each particular output position in the subsequence, processing an input sequence comprising any previous output elements at respective previous output positions preceding the particular output position in the subsequence (Gupta: Pg. 4, section Query chunking: "we partition the queries into chunks and process them sequentially, one chunk at a time"; Pg. 5, section Improving efficiency through top-k attention: "a highly compressed but accurate representation of QKAT and requires only …. storage") , the neural network comprising a sequence of one or more network blocks, the sequence comprising at least one self-attention network block (Gupta: Abstract, Pg. 1 section 1 [1-2]: transformer neural network; Pg. 5, section 3 [1]: "a single self-attention layer" ) configured to perform, for each output position in each subsequence after a first subsequence of the output sequence (Gupta: Pg. 2, [1]: "chunking the query vectors when computing the output one chunk at a time"), operations comprising: obtaining a block input sequence generated from the input sequence for the output position and comprising a respective block input at each of the previous output positions in the subsequence (Gupta: Pg. 2, [1]: "chunking the query vectors when computing the output one chunk at a time"); for each particular previous output position in the subsequence, applying a self-attention mechanism over the block inputs at the previous output positions to generate a respective attention output for the particular previous output position (Gupta: Pg. 2, [1]: "chunking the query vectors when computing the output one chunk at a time"), wherein applying the self-attention mechanism comprises: determining a query from the block input at the particular previous output position (Gupta: Pg. 2, [1]: "chunking the query vectors when computing the output one chunk at a time"), obtaining, from a memory configured to store previous keys and corresponding previous values generated by the neural network when generating previous subsequences preceding the subsequence in the output sequence, one or more particular previous keys and corresponding particular previous values according to a similarity between the determined query and the particular previous values (Gupta: Pg. 5, section -Improving efficiency through top-k attention: "avoid recomputing activations ... a highly compressed but accurate representation of QKT and ... we can cache it in addition to Q, K, v~), and using the query, previous keys, and previous values to generate the attention output for the particular previous output position (Gupta: Pg. 4, sections -Query chunking and Input checkpointing, Pg. 5 section-Improving efficiency through top-k attention: the current query is processed according to the cached Q, K, V); and generating, from the attention outputs corresponding to respective previous output positions, a block output sequence comprising a respective block output at each of the previous output positions in the subsequence (Gupta: Pg. 2 [1], Pg. 9 section 5 [1]: chunking the query vectors when computing the output one chunk at a time sequentially). The examiner finds that sequentially processing query by chunking the query vectors read on the subsequences as claimed in the last limitation. Hence the limitations as claimed are obvious over the teachings in Gupta. Regarding Claim 2, Gupta teaches the limitations as claimed in claim 1 and, wherein: the self-attention mechanism is a first self-attention mechanism, and the at least one self-attention network block is configured to perform, for each output position in each subsequence, operations further comprising: for each particular previous output position in the subsequence, applying a second self-attention mechanism over the block inputs at the previous output positions to generate a respective second attention output for the particular previous output position, wherein applying the second self-attention mechanism comprises: determining a second query from the block input at the particular previous output position, determining second keys from the block inputs at the plurality of previous output positions, determining second values from the block inputs at the plurality of previous output positions, and using the second query, second keys, and second values to generate the second attention output for the particular previous output position (Gupta: Pg. 2 [1], Pg. 4, sections -Query chunking and Input checkpointing, Pg. 9 section 5 [1]: chunking the query vectors and processing sequentially; Abstract, Pg. 7 section 4: multi-head attention layers- each head performing its own self-attention computation). Regarding Claim 3, Gupta teaches the limitations as claimed in claim 2 and, wherein, for each particular previous output position, the corresponding query and second query are the same (Gupta: Pg. 2 [1], Pg. 9 section 5 [1]: chunking the query vectors when computing the output one chunk at a time sequentially). The Examiner further notes that the query being same is not functionally involved in the steps recited. Thus, this information appears arbitrary and will not distinguish the claimed invention from the prior art in terms of patentability. Regarding Claim 4, Gupta teaches the limitations as claimed in claim 2 and, wherein generating the block output sequence comprises, for each particular previous output position, combining (i) the attention output for the particular previous output position and (ii) the second attention output for the particular previous output position (Gupta: Abstract, Pg. 2 [1-4], Pg. 7 section 4: multi-head attention layers- each head performing its own self-attention computation). Regarding Claim 7, Gupta teaches the limitations as claimed in claim 1 and, wherein obtaining one or more particular previous keys and corresponding particular previous values according to a similarity between the determined query and the particular previous values comprises: determining, from the previous keys stored by the memory, the one or more particular previous keys that are closest to the determined query according to a distance metric; and obtaining (i) the one or more determined particular previous keys and (ii) for each determined particular previous key, the corresponding particular previous value stored by the memory (Gupta: Abstract, Pg. 2 [1], Pg 3-5 section 2.2: keep k largest similarity scores with respect of the L keys in Top-k Attention). Regarding Claim(s) 11-14, 17, this/these claim(s) is/are similar in scope as claim(s) 1-4, 7 respectively. Therefore, this/these claim(s) is/are rejected under the same rationale. Regarding Claim(s) 22-25, 27, this/these claim(s) is/are similar in scope as claim(s) 1-4, 7 respectively. Therefore, this/these claim(s) is/are rejected under the same rationale. Claim(s) 5-6, 8-10, 15-16, 18-20, 26, 28-29 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gupta (Memory-efficient Transformers via Top-k Attention) in view of Byun (US 20220230623). Regarding Claim 5, Gupta teaches the limitations as claimed in claim 4 and, wherein, for each particular previous output position, combining (i) the attention output for the particular previous output position and (ii) the second attention output for the particular previous output position comprises summing the attention output and the second attention output position (Gupta: Abstract, Pg. 2 [1-4], Pg. 3 section 2.1 [2], Pg. 7 section 4: multi-head attention layers- each head performing its own self-attention computation); The limitation appear to claim the fundamentals of muti-head attention which combines all of the attention outputs; Nonetheless Byun teaches, combining the attention outputs (Byun: Figs 4A-4B, [90-91, 100-101, 103, 110]: combine multi-head attention outputs); It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Gupta and Byun because the combination would enable using multi-head attention outputs to generate combined output. One of ordinary skill in the art would have been motivated to combine the teachings because the combination utilizes the inherent benefits of using Multi-head attention to improve deep learning models by running multiple attention operations in parallel, allowing the network to capture diverse relationships and feature maps. Regarding Claim 6, Gupta teaches the limitations as claimed in claim 2 and, Byun further teaches, the operations further comprising: providing, for each previous output position in the subsequence, the corresponding second key and second value to the memory for future stages (Byun: Figs 4A-4B, [87, 110]: Key values generated for future stages). The same motivation to combine as stated above applies. Regarding Claim 8, Gupta teaches the limitations as claimed in claim 1 and, wherein using the query, previous keys, and previous values to generate the attention output for the particular previous output position comprises: for each of the one or more previous keys, generating a respective weight value by determining a product of the query and the previous key; and computing a weighted sum of the previous values, wherein each previous value is weighted according to the weight value generated from the corresponding previous key (Gupta: Pg. 3 section 2.1, Pg 4 equation 4: Multi head Attention); The limitation appear to claim the fundamentals of muti-head attention operation; Nonetheless Byun teaches, generating weights in multi-head operation (Byun: [94, 106]: updating weights to train network); The same motivation to combine as stated above applies. Regarding Claim 9, Gupta and Byan teach the limitations as claimed in claim 8 and, wherein, for each of the one or more previous keys, generating a respective weight value comprises normalizing the determined product of the query and previous key using one or more of: a dimensionality of the query, a sum of the products of the query and respective previous keys, or a sum of a set of second products computed between the query and respective second keys determined from the block inputs at the previous output positions in the subsequence (Gupta: Abstract, Pg, 4 [4], Pg. 7 section 4: BERT-base transformer architecture; multi-head attention layers- each head performing its own self-attention computation) (Byan [88-91]: generating outputs using key values). Regarding Claim 10, Gupta teaches the limitations as claimed in claim 1 and, Gupta and Byan further teach, wherein the sequence of network blocks further comprises one or more local self-attention network blocks that are each configured to perform, at each stage, operations comprising, for each output position in each subsequence after a first subsequence of the output sequence: obtaining a local self-attention block input sequence generated from the input sequence for the output position and comprising a respective local self-attention block input at each of the previous output positions in the subsequence; for each particular previous output position in the subsequence, applying a third self-attention mechanism over the local self-attention block inputs at the previous output positions to generate a respective third attention output for the particular previous output position, wherein applying the third self-attention mechanism comprises: determining a third query from the local self-attention block input at the particular previous output position ,determining third keys from the local self-attention block inputs at the previous output positions in the subsequence, determining third values from the local self-attention block inputs at the previous output positions in the subsequence, and using the third query, third keys, and third values to generate the third attention output for the particular previous output position (Gupta: Abstract, Pg, 4 [4], Pg. 7 section 4: BERT-base transformer architecture; multi-head attention layers- each head performing its own self-attention computation) (Byan: Figs. 4A-B, [87-91, 94, 108]: generating outputs using key values). The same motivation to combine as stated above applies. Regarding Claim(s) 15-16, 18-20, this/these claim(s) is/are similar in scope as claim(s) 5-6, 8-10respectively. Therefore, this/these claim(s) is/are rejected under the same rationale. Regarding Claim(s) 26, 28-29 this/these claim(s) is/are similar in scope as claim(s) 6, 8-9 respectively. Therefore, this/these claim(s) is/are rejected under the same rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure in the attached 892. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MANDRITA BRAHMACHARI whose telephone number is (571)272-9735. The examiner can normally be reached Monday to Friday, 11 am to 8 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 571 272 4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Mandrita Brahmachari/Primary Examiner, Art Unit 2144
Read full office action

Prosecution Timeline

Mar 29, 2024
Application Filed
Sep 16, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737047
Finger-Mounted Device With Sensors and Haptics
2y 7m to grant Granted Sep 15, 2026
Patent 12731049
METHODS AND SYSTEMS FOR ANOMALY AND PATTERN DETECTION OF UNSTRUCTURED BIG DATA
4y 9m to grant Granted Sep 08, 2026
Patent 12705487
METHOD FOR SIMPLIFYING AN ARTIFICIAL NEURAL NETWORK
4y 1m to grant Granted Aug 11, 2026
Patent 12705521
METHOD FOR DETERMINING AN ISOLATED OPERATING POINT ASSOCIATED WITH AN ISOLATED REGIME, METHOD FOR DETERMINING AN OPTIMAL SET OF PARAMETERS OF A MEASUREMENT MEANS AND SYSTEM THEREFOR
3y 7m to grant Granted Aug 11, 2026
Patent 12694313
QUANTUM CIRCUIT SIMULATION
4y 3m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+28.9%)
2y 11m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 422 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month