Prosecution Insights
Last updated: October 02, 2026
Application No. 18/980,852

NATURAL LANGUAGE PROCESSING TECHNIQUES USING MULTI-CONTEXT SELF-ATTENTION MACHINE LEARNING FRAMEWORKS

Non-Final OA §101§103
Filed
Dec 13, 2024
Priority
Feb 25, 2022 — provisional 63/314,073 +1 more
Examiner
MASTERS, KRISTEN MICHELLE
Art Unit
Tech Center
Assignee
Optum Inc.
OA Round
1 (Non-Final)
65%
Grant Probability
Moderate
1-2
OA Rounds
1y 2m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 65% of resolved cases
65%
Career Allowance Rate
33 granted / 51 resolved
+4.7% vs TC avg
Strong +24% interview lift
Without
With
+24.1%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
24 currently pending
Career history
87
Total Applications
across all art units

Statute-Specific Performance

§101
36.1%
-3.9% vs TC avg
§103
51.1%
+11.1% vs TC avg
§102
7.2%
-32.8% vs TC avg
§112
3.2%
-36.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 51 resolved cases

Office Action

§101 §103
Detailed Action This communication is in response to the Application filed on 12/13/2024. Claims 1-20 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 3/21/2025 and 3/21/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Independent claim 1 recites, “1. A computer-implemented method comprising: generating, by one or more processors and using a multi-context convolutional self-attention machine learning framework, a cross-context token representation based at least in part on an input text token of an input text sequence, wherein the multi-context convolutional self-attention machine learning framework comprises: (a) a shared token embedding machine learning model, and [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] (b) a context-specific self-attention machine learning model associated with a distinct context window size of a plurality of distinct context window sizes, and providing, by the one or more processors, a natural language processing output for the input text sequence based at least in part on the cross-context token representation. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] Regarding independent Claim 10, Claim 10 is a system claim with limitations similar to that of claim 1 and is rejected under the same rationale. Regarding independent Claim 18, Claim 18 is a non-transitory computer-readable media claim with limitations similar to that of claim 1 and is rejected under the same rationale. The Dependent Claims do not include additional limitations that could incorporate the abstract idea into a practical application or cause the Claim as a whole to amount to significantly more than the underlying abstract idea. This judicial exception is not integrated into a practical application. In particular, claim 1, 10 and 18 recite the additional element of “processor”. For example, in [0046] of the as filed specification, there is the description For example, the processing element 205 may be embodied as one or more complex programmable logic devices (CPLDs), microprocessors, multi-core processors, coprocessing entities, application-specific instruction-set processors (ASIPs), microcontrollers, and/or controllers. Further, the processing element 205 may be embodied as one or more other processing devices or circuitry. The term circuitry may refer to an entirely hardware embodiment or a combination of hardware and computer program products. Thus, the processing element 205 may be embodied as integrated circuits, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), hardware accelerators, other circuitry, and/or the like. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using a processor is noted as a general computer. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Further, the additional limitation in the claims noted above are directed towards insignificant solution activity. The claims are not patent eligible. Dependent claim 2 recites, “2. The computer-implemented method of claim 1, wherein the shared token embedding machine learning model is configured to generate an initial token embedding for the input text token and the context-specific self-attention machine learning model is configured to generate a context-specific token representation for the input text token based at least in part on the initial token embedding. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] no additional limitations. Dependent claim 3 recites, “3. The computer-implemented method of claim 2, wherein the multi-context convolutional self-attention machine learning framework further comprises a cross-context representation inference machine learning model that is configured to generate the cross-context token representation based at least in part on the context-specific token representation. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] no additional limitations. Dependent claim 4 recites, “4. The computer-implemented method of claim 3, wherein the cross-context representation inference machine learning model further comprises an attention-based machine learning model and is configured to: generate a sequence representation for the input text token using the attention-based machine learning model and based at least in part on the context-specific token representation for the input text token; and generate the cross-context token representation for the input text token based at least in part on the sequence representation. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] no additional limitations. Dependent claim 5 recites, “5. The computer-implemented method of claim 1, wherein the context-specific self-attention machine learning model is one of a plurality of context-specific self-attention machine learning models that are respectively associated with the plurality of distinct context window sizes. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] no additional limitations. Dependent claim 7 recites, “6. The computer-implemented method of claim 5, wherein the cross-context token representation is based at least in part at least one of a plurality of context-specific token representations respectively generated by the plurality of context-specific self-attention machine learning models. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] no additional limitations. Dependent claim 7 recites, “7. The computer-implemented method of claim 1, wherein the distinct context window size of the context-specific self-attention machine learning model is determined based at least in part on a sliding window size for the context-specific self-attention machine learning model. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] no additional limitations. Dependent claim 8 recites, “8. The computer-implemented method of claim 1, wherein the context-specific self-attention machine learning model comprises a convolutional neural network, is associated with a window size, and the distinct context window size is based at least in part on the window size. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] no additional limitations. Dependent claim 9 recites, “9. The computer-implemented method of claim 1, wherein the context-specific self-attention machine learning model comprises a convolutional neural network, is associated with a filter width, and the distinct context window size is based at least in part on the filter width. [This relates to a series of information processing/mental like steps and mathematical operations that a human can perform.] no additional limitations. As to dependent Claim 11, Claim 11 is a system claim with limitations similar to that of claim 2 and is rejected under the same rationale. As to dependent Claim 12, Claim 12 is a system claim with limitations similar to that of claim 3 and is rejected under the same rationale. As to dependent Claim 13, Claim 13 is a system claim with limitations similar to that of claim 4 and is rejected under the same rationale. As to dependent Claim 14, Claim 14 is a system claim with limitations similar to that of claim 5 and is rejected under the same rationale. As to dependent Claim 15, Claim 15 is a system claim with limitations similar to that of claim 6 and is rejected under the same rationale. As to dependent Claim 16, Claim 16 is a system claim with limitations similar to that of claim 7 and is rejected under the same rationale. As to dependent Claim 17, Claim 17 is a system claim with limitations similar to that of claim 8 and is rejected under the same rationale. As to dependent Claim 19, Claim 19 is a non-transitory computer-readable media claim with limitations similar to that of claim 9 and is rejected under the same rationale. As to dependent Claim 20, Claim 20 is a non-transitory computer-readable media claim with limitations similar to that of claim 8 and is rejected under the same rationale. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 5-7, 10- 12, and 14-18 are rejected under 35 U.S.C. 103 as being unpatentable over Gupta (U.S. Patent Number US-11580968-B1), in view of Pang (U.S. Patent Number US-11941346-B2). Regarding Claim 1, Gupta teaches 1. A computer-implemented method comprising: generating, by one or more processors and using a multi-context convolutional self-attention machine learning framework, a cross-context token representation based at least in part on an input text token of an input text sequence, wherein the multi-context convolutional self-attention machine learning framework comprises: (see Gupta (5:49-6:12)) “(31) In an embodiment, the encoder component performs various encoding techniques for the different signals used in modeling. One such signal includes the user utterances (for example, past utterances 302A and 302B, in the example of FIG. 2, and the current utterance 302N). In some embodiments, to encode the utterances, the encoder component uses directional self-attention architecture-based encoder (DiSAN) units 304A-304N. For example, the encoder component takes as input the current user utterance 302N, denoted here by custom character.sub.i∈custom character.sup.d.sup.e.sup.×k.sup.i where k, represents a number of tokens in i.sup.th user utterance. The encoder component then passes this vector through an embedding layer to generate a hidden state, h.sub.i∈custom character.sup.d.sup.h.sup.×k.sup.i, where d.sub.h is the hidden layer dimension. The encoder component then adds a learned absolute position embedding, P.sub.e∈custom character.sup.d.sup.h.sup.×k.sup.i, resulting in a new vector, h.sub.i=h.sub.i+P.sub.e. Among other benefits, the embedding enables the system to capture syntactic information related to the utterance. 32) In some embodiments, the input above is passed through a DiSAN unit that broadly comprises token-to-token (t2t) attention (for example, t2t 308A-308N, respectively corresponding to each utterance 302A-302N), masks, and source-to-token (s2t) attention (for example, t2t 310A-310N, respectively corresponding to each utterance 302A-302N). On input h.sub.i, for example, the encoder component first applies a multi-dimensional t2t attention layer that encodes the dependency between a token and all other tokens in the sentence. In addition, forward and backward masks are added to the attention computation to incorporate directional information.”) (see Gupta (9:40-59) “(76) FIG. 5 is a flow diagram illustrating operations 500 of a method for a contextual natural language understanding (cNLU) framework that is able to incorporate contextual signals of variable history length to perform joint intent classification (IC) and slot labeling (SL) tasks according to some embodiments. Some or all of the operations 500 (or other processes described herein, or variations, and/or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware or combinations thereof. The code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising instructions executable by one or more processors. The computer-readable storage medium is non-transitory. In some embodiments, one or more (or all) of the operations 500 are performed by contextual natural language understanding (cNLU) model of the other figures.”) (a) a shared token embedding machine learning model, and (see Gupta (8:31-42) “(62) In an embodiment, the memory is composed to two trainable word embeddings, A and C (shown in FIG. 4 as embedding 406A and 406C, respectively), referring to input memory and output memory, respectively, converting tokens into vectors {s.sub.k} and {g.sub.k}. An embedding 406B is built for the current user utterance as well. It maps the current utterance 410 {u.sub.t} into feature vectors {q.sub.i}. For each utterance vector the attention weights are computed against dialog history by calculating the inner-product between the current utterance and input memory {s.sub.k}…”) and providing, by the one or more processors, a natural language processing output for the input text sequence based at least in part on the cross-context token representation. (see Gupta 2:53-3:6) “(14) FIG. 1 is a diagram illustrating an environment for an example dialog system including a cNLU model according to some embodiments. As shown in FIG. 1, the dialog system 100 includes a cNLU model 102 and other optional components including a dialog manager 104, dialog policy component 106, and NLG model 108, among other possible components. At a high level, the dialog system 100 receives user utterances (e.g., a user utterance 110) from users (e.g., a user 112), uses a cNLU model to generate IC/SL predictions 114 (based on both the user utterance 110 and other contextual signals 116, as described in more detail herein), passes the IC/SL predictions 114 to a dialog manager 104 that maintains state information related to the dialog, a dialog policy manager that decides on a next action to be performed by the system based on data from the cNLU model 102 and dialog manager 104, and finally a NLG component which forms a reply that is converted into an agent utterance 118 (e.g., a text-based reply, voice-based reply, or the like) that can be understood by the user 112. Examples of a cNLU model 102 are described in more detail in sections hereinafter.”) (see Gupta (2:23-33) “(12) Natural language understanding (NLU) is an important component of such dialog systems and, in particular, for capturing semantic information from a user's utterances at each turn of a dialogue with a smart conversational agent. At a high level, NLU in this context involves two tasks: intent classification (IC) and slot labeling (SL). An intent in the context of IC identifies the goal underlying an expressed utterance (that is, what the user is trying to achieve with the utterance), and slots identify optional parameters of these intents (that is, information provided by the user in the utterance that is relevant to the user's intent).”) Gupta does not specifically teach (b) a context-specific self-attention machine learning model associated with a distinct context window size of a plurality of distinct context window sizes, However, Pang does teach this limitation (see Pang [0012-0015] In view of the existing issues in text summarization, embodiments described herein are directed to a document summarization model which takes into account all of the information of the source document while preserving efficient processing complexity. This is achieved by jointly encoding the input of textual tokens into two different levels of representations. A bottom-up representation of the tokens is generated via a transformer with only local self-attention within a predefined window of each input token. A top-level representation of the tokens is generated by pooling the bottom-up inferred tokens, and passing the pooled tokens through a transformer with full self-attention. The top-level representation is used to update the bottom-up representation by another transformer using cross-attention between the two representations. The output tokens are then sent to a decoder which produces the output summary of the document. [0013] In one embodiment, in the bottom-up inference, contextual embeddings of the tokens are computed with a specified number of layers of local self-attention. In particular, each token only attends to nearby tokens within a window of a specified size. The computation complexity is thereby limited in contrast to full self-attention models. [0014] In the top-down inference, documents are encoded into representations at a coarser granularity level or at a more abstract temporal scale. This is referred to as a top-level or segment representation. [0015] At the top-down inference, full self-attention may be efficiently used due to the coarser granularity, which allows these top-level units to capture global document context. The bottom-up inferred token representations may then be updated with the top-level representations. This may be achieved with cross-attention between the top-level (segment) and bottom-level (token) units. This injects global contextual information to token representations, which completes the combination of the bottom-up and top-down inference for token representation.”) Gupta and Pang are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Gupta to incorporate (b) a context-specific self-attention machine learning model associated with a distinct context window size of a plurality of distinct context window sizes, of Pang. This allows for greater efficiency and greater capture of context as recognized by Pang [0015]. Regarding independent Claim 10, Claim 10 is a system claim with limitations similar to that of claim 1 and is rejected under the same rationale. Furthermore, Gupta teaches A system comprising: one or more processors; and one or more memories storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: (see Gupta (14:5-30) “(102) In some embodiments, a computer system 800 includes one or more offload cards 870 (including one or more processors 875, and possibly including the one or more network interfaces 840) that are connected using an I/O interface 830 (e.g., a bus implementing a version of the Peripheral Component Interconnect-Express (PCI-E) standard, or another interconnect such as a QuickPath interconnect (QPI) or UltraPath interconnect (UPI)). For example, in some embodiments the computer system 800 may act as a host electronic device (e.g., operating as part of a hardware virtualization service) that hosts compute instances, and the one or more offload cards 870 execute a virtualization manager that can manage compute instances that execute on the host electronic device. As an example, in some embodiments the offload card(s) 870 can perform compute instance management operations such as pausing and/or un-pausing compute instances, launching and/or terminating compute instances, performing memory transfer/copying operations, etc. These management operations may, in some embodiments, be performed by the offload card(s) 870 in coordination with a hypervisor (e.g., upon a request from a hypervisor) that is executed by the other processors 810A-810N of the computer system 800.”) Regarding independent Claim 18, Claim 18 is a non-transitory computer-readable media claim with limitations similar to that of claim 1 and is rejected under the same rationale. Furthermore, Gupra teaches 18. One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: (see Gupta (14:35-55) “(103) In some embodiments, system memory 820 may be one embodiment of a computer-accessible medium configured to store program instructions and data as described above. However, in other embodiments, program instructions and/or data may be received, sent or stored upon different types of computer-accessible media. Generally speaking, a computer-accessible medium may include non-transitory storage media or memory media such as magnetic or optical media, e.g., disk or DVD/CD coupled to computer system 800 via I/O interface 830. A non-transitory computer-accessible storage medium may also include any volatile or non-volatile media such as RAM (e.g., SDRAM, double data rate (DDR) SDRAM, SRAM, etc.), read only memory (ROM), etc., that may be included in some embodiments of computer system 800 as system memory 820 or another type of memory. Further, a computer-accessible medium may include transmission media or signals such as electrical, electromagnetic, or digital signals, conveyed via a communication medium such as a network and/or a wireless link, such as may be implemented via network interface 840.”) As to Claim 2, Gupta in view of Pang teaches: 2. The computer-implemented method of claim 1, Furthermore, Gupta teaches wherein the shared token embedding machine learning model is configured to generate an initial token embedding for the input text token and the context-specific self-attention machine learning model is configured to generate a context-specific token representation for the input text token based at least in part on the initial token embedding. (see Gupta (5:49-6:12)) “(31) In an embodiment, the encoder component performs various encoding techniques for the different signals used in modeling. One such signal includes the user utterances (for example, past utterances 302A and 302B, in the example of FIG. 2, and the current utterance 302N). In some embodiments, to encode the utterances, the encoder component uses directional self-attention architecture-based encoder (DiSAN) units 304A-304N. For example, the encoder component takes as input the current user utterance 302N, denoted here by custom character.sub.i∈custom character.sup.d.sup.e.sup.×k.sup.i where k, represents a number of tokens in i.sup.th user utterance. The encoder component then passes this vector through an embedding layer to generate a hidden state, h.sub.i∈custom character.sup.d.sup.h.sup.×k.sup.i, where d.sub.h is the hidden layer dimension. The encoder component then adds a learned absolute position embedding, P.sub.e∈custom character.sup.d.sup.h.sup.×k.sup.i, resulting in a new vector, h.sub.i=h.sub.i+P.sub.e. Among other benefits, the embedding enables the system to capture syntactic information related to the utterance. 32) In some embodiments, the input above is passed through a DiSAN unit that broadly comprises token-to-token (t2t) attention (for example, t2t 308A-308N, respectively corresponding to each utterance 302A-302N), masks, and source-to-token (s2t) attention (for example, t2t 310A-310N, respectively corresponding to each utterance 302A-302N). On input h.sub.i, for example, the encoder component first applies a multi-dimensional t2t attention layer that encodes the dependency between a token and all other tokens in the sentence. In addition, forward and backward masks are added to the attention computation to incorporate directional information.”) (see Gupta (9:40-59) “(76) FIG. 5 is a flow diagram illustrating operations 500 of a method for a contextual natural language understanding (cNLU) framework that is able to incorporate contextual signals of variable history length to perform joint intent classification (IC) and slot labeling (SL) tasks according to some embodiments. Some or all of the operations 500 (or other processes described herein, or variations, and/or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware or combinations thereof. The code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising instructions executable by one or more processors. The computer-readable storage medium is non-transitory. In some embodiments, one or more (or all) of the operations 500 are performed by contextual natural language understanding (cNLU) model of the other figures.”) (see Gupta (8:31-42) “(62) In an embodiment, the memory is composed to two trainable word embeddings, A and C (shown in FIG. 4 as embedding 406A and 406C, respectively), referring to input memory and output memory, respectively, converting tokens into vectors {s.sub.k} and {g.sub.k}. An embedding 406B is built for the current user utterance as well. It maps the current utterance 410 {u.sub.t} into feature vectors {q.sub.i}. For each utterance vector the attention weights are computed against dialog history by calculating the inner-product between the current utterance and input memory {s.sub.k}…”) a natural language processing output for the input text sequence; (see Gupta (2:23-33) “(12) Natural language understanding (NLU) is an important component of such dialog systems and, in particular, for capturing semantic information from a user's utterances at each turn of a dialogue with a smart conversational agent. At a high level, NLU in this context involves two tasks: intent classification (IC) and slot labeling (SL). An intent in the context of IC identifies the goal underlying an expressed utterance (that is, what the user is trying to achieve with the utterance), and slots identify optional parameters of these intents (that is, information provided by the user in the utterance that is relevant to the user's intent).”) As to Claim 3, Gupta in view of Pang teaches: 3. The computer-implemented method of claim 2, Furthermore, Gupta teaches, wherein the multi-context convolutional self-attention machine learning framework further comprises a cross-context representation inference machine learning model that is configured to generate the cross-context token representation based at least in part on the context-specific token representation. (See Gupta (15:65-166) “(80) In one embodiment, the ML model includes: an encoder component used to encode the user utterance and the contextual information for each turn into a plurality of turn vectors, a context fusion component including a multi-dimensional attention layer to be applied to the plurality of turn vectors, and an intent classification and slot label prediction component used to obtain the intent classification and the one or more slots labels.”) As to claim 5 Gupta in view of Pang teaches 5. The computer-implemented method of claim 1, Furthermore Pang teaches wherein the context-specific self-attention machine learning model is one of a plurality of context-specific self-attention machine learning models that are respectively associated with the plurality of distinct context window sizes. (see Pang [0041] “In some embodiments, the Summarization module 330 may further include the bottom-up inference module 331, top-down inference module 332, and a cross-attention module 333. The bottom-up inference module 331 (examiner interprets machine learning models as “331, 332, and 333”) is configured to produce bottom-up representation tokens of an input document using local self-attention. For example, as discussed with reference to local self-attention 104 of FIG. 1, and step 210 of FIG. 2.”)(see Pang [0046] The same encoder-decoder architecture was tested for all datasets. The tested encoder has 8 bottom-up inference layers and 4 top-down inference layers for tokens, and 2 self-attention layers for segments. The decoder has 12 layers. The encoder layers for tokens (12 layers) and the decoder layers are all initialized from BART described in Lewis et al., BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension, in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871-7880, 2020, except the parameters for token-segment cross-attention in the top-down inference layers, which are randomly initialized. The self-attention parameters for segments are also randomly initialized. The window size is 1024 unless otherwise specified. These settings closely follow Longformer described in Beltagy et al., Longformer: The long-document transformer, arXiv preprint arXiv:2004.05150, 2020, which has 12 layers for the encoder and decoder, is initialized from BART, and uses a local window size of 1024. Thus, comparison with Longformer is a test of the effect of top-down correction for token representations. Standard train/validation/test splits are used for all datasets. Model performance is evaluated with ROUGE scores described in Lin, ROUGE: A package for automatic evaluation of summaries, in Text Summarization Branches Out, pages 74-81, 2004. Reported performance is based on the checkpoint with the best validation R-2 score.”) Gupta and Pang are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Gupta and Pang to incorporate the context-specific self-attention machine learning model is one of a plurality of context-specific self-attention machine learning models that are respectively associated with the plurality of distinct context window sizesof Pang. This allows for greater efficiency and greater capture of context as recognized by Pang [0015]. As to claim 6, Gupta in view of Pang teaches 6. The computer-implemented method of claim 5, Furthermore, Gupta teaches wherein the cross-context token representation is based at least in part at least one of a plurality of context-specific token representations respectively generated by the plurality of context-specific self-attention machine learning models. (See Gupta (15:65-166) “(80) In one embodiment, the ML model includes: an encoder component used to encode the user utterance and the contextual information for each turn into a plurality of turn vectors, a context fusion component including a multi-dimensional attention layer to be applied to the plurality of turn vectors, and an intent classification and slot label prediction component used to obtain the intent classification and the one or more slots labels.”) (84) In one embodiment, output from the ML model is provided to an ensemble model with output from one or more additional ML models to obtain the intent classification and the one or more slot labels for the message.”) (see Gupta (5:19-30) “(27) To learn this f.sub.context, in some embodiments, one of at least two different cNLU-based ML models described herein can be used. At a high level, each of these models takes contextual history derived from a turn history over K turns into account in addition to the current utterance itself. As described in separate sections below, these example models include a contextual self-attention model (cSA) and a contextual Memory network (cMemN), each used to perform joint IC and SL tasks.”) As to Claim 7 Gupta in view of Pang teaches 7. The computer-implemented method of claim 1, Furthermore, Pang teaches wherein the distinct context window size of the context-specific self-attention machine learning model is determined based at least in part on a sliding window size for the context-specific self-attention machine learning model. (see Pang [0013] “In one embodiment, in the bottom-up inference, contextual embeddings of the tokens are computed with a specified number of layers of local self-attention. In particular, each token only attends to nearby tokens within a window of a specified size. The computation complexity is thereby limited in contrast to full self-attention models.”) Gupta in view of Pang are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Gupta and Pang to incorporate wherein the distinct context window size of the context-specific self-attention machine learning model is determined based at least in part on a sliding window size for the context-specific self-attention machine learning model of Pang. This allows for greater efficiency and greater capture of context as recognized by Pang [0015]. As to dependent Claim 11, Claim 11 is a system claim with limitations similar to that of claim 2 and is rejected under the same rationale. As to dependent Claim 12, Claim 12 is a system claim with limitations similar to that of claim 3 and is rejected under the same rationale. As to dependent Claim 14, Claim 14 is a system claim with limitations similar to that of claim 5 and is rejected under the same rationale. As to dependent Claim 15, Claim 15 is a system claim with limitations similar to that of claim 6 and is rejected under the same rationale. As to dependent Claim 16, Claim 16 is a system claim with limitations similar to that of claim 7 and is rejected under the same rationale. As to dependent Claim 17, Claim 17 is a system claim with limitations similar to that of claim 8 and is rejected under the same rationale. Claims 4 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Gupta (U.S. Patent Number US-11580968-B1), in view of Pang (U.S. Patent Number US-11941346-B2) and further in view of Peng-Hsuan Li (NPL Why Attention? Analyze BiLSTM Deficiency and Its Remedies in the Case of NER). As to Claim 4, Gupta in view of Pang teaches: 4. The computer-implemented method of claim 3, Gupta in view of Pang do not specifically teach wherein the cross-context representation inference machine learning model further comprises an attention-based machine learning model and is configured to: generate a sequence representation for the input text token using the attention-based machine learning model and based at least in part on the context-specific token representation for the input text token; and generate the cross-context token representation for the input text token based at least in part on the sequence representation. (see Peng-Hsuan Li (Section 3.5-3.6) “As the forward and backward hidden states are interleaved between stacked LSTM layers, Cross-BiLSTM-CNN models cross-context patterns by computing representations of the whole sequence in a feed-forward, additive manner. Specifically, for the XOR cases introduced in Section 3.3, 3.4, although phrase 1 and phrase 3 still have the same past context for the middle token and hence the first… As the higher-level LSTMs of Cross-BiLSTM-CNN have interleaved input from forward and backward hidden states down below, their weight parameters double the size of the first-level LSTMs. Nevertheless, the cross formulation pro-vides the modeling capacity abs 3.6 Att-BiLSTM-CNN Another way to capture the interaction between past and future context per time step is to add a token-level self-attentive mechanism on top of the same BiLSTM formulation introduced in Section 3.2. Given the hidden features H of a whole sequence, the model projects each hidden state to different subspaces, depending on whether it is used as the query vector to consult other hidden states for each word token, the key vector to compute its dot-similarities with in-coming queries, or the value vector to be weighted and actually convey information to the querying token. As different aspects of a task can call for different attention, multiple attention”) Gupta in view of Pang and Peng-Hsuan Li are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Gupta and Pang to incorporate wherein the cross-context representation inference machine learning model further comprises an attention-based machine learning model and is configured to: generate a sequence representation for the input text token using the attention-based machine learning model and based at least in part on the context-specific token representation for the input text token; and generate the cross-context token representation for the input text token based at least in part on the sequence representation of Peng-Hsuan Li. This allows for improved sequence labeling and identification with multiple token mentions as recognized by Peng-Hsuan Li [Abstract]. As to dependent Claim 13, Claim 13 is a system claim with limitations similar to that of claim 4 and is rejected under the same rationale. Claims 8, 9, 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Gupta (U.S. Patent Number US-11580968-B1), in view of Pang (U.S. Patent Number US-11941346-B2) and further in view of BRADBURY (U.S. Patent Number US-20180129937-A1) As to Claim 8 Gupta in view of Pang teaches 8. The computer-implemented method of claim 1, Gupta in view of Pang do not specifically teach wherein the context-specific self-attention machine learning model comprises a convolutional neural network, is associated with a window size, and the distinct context window size is based at least in part on the window size. However, Bradbury does teach this limitation (see Bradbury [0059] “Regarding the state vector, a current state vector c.sub.t is the consolidation of a current activation vector z.sub.t with the past state vector c.sub.t−1. The current activation vector z.sub.t is identified by a current convolutional vector y.sub.t, which is derived from a convolution over a current time series window of input vectors x.sub.t, . . . , x.sub.t+k−1, where k is the convolutional filter size or width. Anthropomorphically, the current state vector c.sub.t knows the recipe of combining or mixing a currently convolved input vector window x.sub.t, . . . , x.sub.t+k−1 with the past state vector c.sub.t−1 so as to summarize the current input vector window x.sub.t, . . . , x.sub.t+k−1 in light of the contextual past. Thus the current activation vector z.sub.t and the past state vector c.sub.t−1 are used to generate the current state vector c.sub.t that includes aspects of the current input vector window x.sub.t, . . . , x.sub.t+k−1. “) (see Bradbury [0041] “QRNN convolutional layer 200 performs parallel convolutions tom time series windows over the input vectors x.sub.1, . . . , x.sub.6, . . . , x.sub.n with a bank of b filters to concurrently output a sequence Yϵcustom-character.sup.ζd×m of m convolutional vectors y.sub.1, . . . , y.sub.5, . . . , y.sub.m.Math.ζd is the dimensionality of each convolutional vector, where ζ identifies a dimensionality augmentation parameter. As used herein, “parallelism across the timestep or time series dimension” or “timestep or time series parallelism” refers to the QRNN convolutional layer 200 applying a convolutional filter bank in parallel to the input vectors x.sub.1, . . . , x.sub.6, . . . , x.sub.n over m time series windows to concurrently produce m convolutional vectors y.sub.1, . . . , y.sub.5, . . . , y.sub.m.”) Gupta in view of Pang and Bradury are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Gupta and Pang to incorporate wherein the context-specific self-attention machine learning model comprises a convolutional neural network, is associated with a window size, and the distinct context window size is based at least in part on the window size of Bradury. This allows high throughput and good scaling to long sequences as recognized by Bradury [0030]. As to Claim 9, Gupta and Pang teach 9. The computer-implemented method of claim 1, Gupta Pang Li do not specifically teach wherein the context-specific self-attention machine learning model comprises a convolutional neural network, is associated with a filter width, and the distinct context window size is based at least in part on the filter width. However, Bradbury does teach this limitation (see Bradbury [0059] “Regarding the state vector, a current state vector c.sub.t is the consolidation of a current activation vector z.sub.t with the past state vector c.sub.t−1. The current activation vector z.sub.t is identified by a current convolutional vector y.sub.t, which is derived from a convolution over a current time series window of input vectors x.sub.t, . . . , x.sub.t+k−1, where k is the convolutional filter size or width. Anthropomorphically, the current state vector c.sub.t knows the recipe of combining or mixing a currently convolved input vector window x.sub.t, . . . , x.sub.t+k−1 with the past state vector c.sub.t−1 so as to summarize the current input vector window x.sub.t, . . . , x.sub.t+k−1 in light of the contextual past. Thus the current activation vector z.sub.t and the past state vector c.sub.t−1 are used to generate the current state vector c.sub.t that includes aspects of the current input vector window x.sub.t, . . . , x.sub.t+k−1. “) (see Bradbury [0041] “QRNN convolutional layer 200 performs parallel convolutions tom time series windows over the input vectors x.sub.1, . . . , x.sub.6, . . . , x.sub.n with a bank of b filters to concurrently output a sequence Yϵcustom-character.sup.ζd×m of m convolutional vectors y.sub.1, . . . , y.sub.5, . . . , y.sub.m.Math.ζd is the dimensionality of each convolutional vector, where ζ identifies a dimensionality augmentation parameter. As used herein, “parallelism across the timestep or time series dimension” or “timestep or time series parallelism” refers to the QRNN convolutional layer 200 applying a convolutional filter bank in parallel to the input vectors x.sub.1, . . . , x.sub.6, . . . , x.sub.n over m time series windows to concurrently produce m convolutional vectors y.sub.1, . . . , y.sub.5, . . . , y.sub.m.”) Gupta in view of Pang and Bradury are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Gupta and Pang to incorporate teach wherein the context-specific self-attention machine learning model comprises a convolutional neural network, is associated with a filter width, and the distinct context window size is based at least in part on the filter width of Bradury. This allows high throughput and good scaling to long sequences as recognized by Bradury [0030]. As to dependent Claim 19, Claim 19 is a non-transitory computer-readable media claim with limitations similar to that of claim 9 and is rejected under the same rationale. As to dependent Claim 20, Claim 20 is a non-transitory computer-readable media claim with limitations similar to that of claim 8 and is rejected under the same rationale. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. JAGANATHAN (US Publication No. US 2020/0302297 A1) Artificial Intelligence-Based Base Calling Abstract: The technology disclosed processes input data through a neural network and produces an alternative representation of the input data. The input data includes per-cycle image data for each of one or more sequencing cycles of a sequencing run. The per-cycle image data depicts intensity emissions of one or more analytes and their surrounding background captured at a respective sequencing cycle. The technology disclosed processes the alternative representation through an output layer and producing an output and base calls one or more of the analytes at one or more of the sequencing cycles based on the output. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KRISTEN MICHELLE MASTERS whose telephone number is (703)756-1274. The examiner can normally be reached M-F 8:30 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KRISTEN MICHELLE MASTERS/Examiner, Art Unit 2659 /EDGAR X GUERRA-ERAZO/Primary Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Dec 13, 2024
Application Filed
Aug 11, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731600
METHOD AND DEVICE FOR MANAGING AUDIO BASED ON SPECTROGRAM
3y 5m to grant Granted Sep 08, 2026
Patent 12725628
VOICE IDENTIFICATION FOR OPTIMIZING VOICE SEARCH RESULTS
4y 8m to grant Granted Sep 01, 2026
Patent 12724986
METHODS, SYSTEMS, AND NON-TRANSITORY COMPUTER-READABLE RECORD MEDIA TO PROVIDE TRANSLATION RESULT OF CONVERSATION MESSAGE
4y 5m to grant Granted Sep 01, 2026
Patent 12707198
ACOUSTIC ECHO CANCELLATION SYSTEM AND ASSOCIATED METHOD
3y 0m to grant Granted Aug 11, 2026
Patent 12694889
PROFANITY FILTER FOR COLLABORATION SESSIONS IN HETEROGENOUS COMPUTING PLATFORMS
3y 10m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
65%
Grant Probability
89%
With Interview (+24.1%)
3y 0m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 51 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month