Prosecution Insights
Last updated: August 18, 2026
Application No. 18/424,572

SLANG USAGE DETECTION AND MITIGATION FOR LARGE LANGUAGE MODELS

Final Rejection §103
Filed
Jan 26, 2024
Examiner
TENGBUMROONG, NATHAN NARA
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Intuit Inc.
OA Round
2 (Final)
42%
Grant Probability
Moderate
3-4
OA Rounds
6m
Est. Remaining
75%
With Interview

Examiner Intelligence

Grants 42% of resolved cases
42%
Career Allowance Rate
10 granted / 24 resolved
-20.3% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
17 currently pending
Career history
51
Total Applications
across all art units

Statute-Specific Performance

§101
26.0%
-14.0% vs TC avg
§103
57.3%
+17.3% vs TC avg
§102
13.7%
-26.3% vs TC avg
§112
2.6%
-37.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 24 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment Claims 1, 4-12, 14, and 16-20 are amended. Claims 2-3 and 15 are cancelled. Claims 21-23 are newly added. As such, claims 1, 4-14, and 16-23 are presented for examination. Response to Arguments Rejection under 35 U.S.C. 103 Applicant’s arguments been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Allowable Subject Matter Claim 4 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 and 5-6 are rejected under 35 U.S.C. 103 as being unpatentable over Elisco et al. (US 20220300711 A1; hereinafter referred to as Elisco) in view of Esponda (US 20200201898 A1), Erb et al. (US 20250086440 A1; hereinafter referred to as Erb), and Pei et al. (Pei, Z., Sun, Z., & Xu, Y. (2019, November). Slang detection and identification. In Proceedings of the 23rd conference on computational natural language learning (CoNLL) (pp. 881-889); hereinafter referred to as Pei). Regarding claim 1, Elisco discloses: a method of training a machine learning (ML) model to detect slang usage, comprising… masking at least one token of the first plurality of tokens for each respective training data instance of the first subset and the second plurality of tokens for each respective training data instance of the second subset to create a plurality of masked training data instances ([0048] During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token, a second proportion such as 10% may be replaced with a random token, and/or remaining tokens may be kept the same. A task in training may be to predict an original masked token); training the ML model, using a masked-language modeling head, to predict the at least one token masked for each of the plurality of masked training data instances using the plurality of training data instances and the plurality of masked training data instances… ([0048] A task in training may be to predict an original masked token… Only masked tokens in a note may contribute to a loss function which influences how to change network parameters to improve performance. A model being trained may learn which tokens appear in a similar context for a given dataset). Elisco does not explicitly, but Esponda teaches: obtaining a first subset of training data instances of a plurality of training data instances, wherein each respective training data instance in the first subset comprises a first plurality of tokens that represents at least a definition associated with a slang token of a plurality of slang tokens but does not include the slang token… ([0037] the model is created using a statistical machine learning process that builds the model by learning mathematical relationships between features of the usage context and corresponding definitions. In an embodiment, the model is trained using a data set that includes previously-determined acronym definition-usage context pairs. The trained model is then used by the inference engine to predict a context-relevant definition for the current use of the acronym. Acronyms can represent slang.); and determining, by an entailment model, that an entailment score between the definition of the respective training data instance and the second plurality of tokens satisfies a threshold… ([0058] the confidence score is binary; for example, the confidence score has a value of 1 if a user has confirmed the definition and usage context pair, and the confidence score has a value of 0 if the definition and usage context pair has not been user-confirmed. In other embodiments, for example when machine learning is used, the confidence score may vary along a continuum, such as between 0 and 1). Elisco and Esponda are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco to combine the teachings of Esponda because doing so would allow for slang-definition training datasets to be used for detecting slang instances, such as abbreviations/acronyms, by incorporating contextual information to determine the definition of a slang term, leading to more accurate slang detection (Esponda [0019] The disclosed technologies improve upon these prior approaches by automatically inferring the relevant definition based on the context of an application or work process. Thus, the disclosed technologies do not require users to explicitly provide context information in order to suggest a definition for an abbreviation or an acronym that is likely to be relevant to the current context). The combination of Elisco and Esponda does not explicitly, but Erb teaches: generating a second subset of training data instances of the plurality of training data instances by, for each respective training data instance in the first subset: generating, by a large language model (LLM), a second plurality of tokens based on a prompt comprising the respective training data instance and the slang token associated with the respective training data instance ([0080] - the server determines a suitable input prompt for instructing the generative AI model to output a definition for the selected words. The input prompt includes at least a portion of the selected text and instructions for generating supplementary data (i.e., definitions, summary, etc.) associated with the selected text); determining the second plurality of tokens comprise the slang token… ([0082] the server receives an output of the generative AI model. The output may include a definition for a single word, a group of multiple words, and/or one or more phrases contained in the user-selected text. Additionally, or alternatively, the output may include a summary of a passage of text selected by the user). Elisco, Esponda, and Erb are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco and Esponda to combine the teachings of Erb because doing so would allow for a set of training data to be generated using a prompt to an LLM for a definition of a term (such as slang), leading to improved training of ML model for slang detection (Erb [0043] the definitions having the highest ratings/ranking may be prioritized in responding to user requests for definitions of selected terms. If a certain definition (or group of definitions) for a term has a statistically significant rating advantage, the alternative (i.e., non-preferred) definitions may be deleted from the database. The computing system may obtain new definitions of terms and store them in place of the deleted definitions (up to a defined limit on number of definitions). Additionally, or alternatively, the preferred definitions may be used as part of (i.e., included in) input prompts for obtaining new definitions of terms). The combination of Elisco, Esponda, and Erb does not explicitly, but Pei teaches: labeling each respective training data instance of the plurality of training data instances with at least one of a slang instance label, a non-slang instance label, one or more slang token labels, or one or more non-slang token labels to generate a plurality of labeled training data instances ([3.1] In the slang identification task, our models identify each token within the input sentence as ‘nonslang’ or ‘slang’ by sequence labeling, which determines the exact positions of slang usage); and training the ML model ([Fig. 2] For a specific token fire in the source sentence “she can cook some fire food”, the related linguistic features are represented as token vectors to concatenate the feature-based input for this token. Each randomly initialized vector is updated during training), using a classification head, to: classify each respective labeled training data instance of the plurality of labeled training data instances as one of a slang instance or a non-slang instance ([3.1] the models in the identification task encapsulate the detection task; an empty prediction that labels all tokens as ‘non-slang’ is equivalent to classifying the sentence as a non-slang sentence in the detection task, and vice versa) and thereby generate a classification output; or classify each token associated with each respective labeled training data instance of the plurality of labeled training data instances as one of a slang token or a non-slang token and thereby generate the classification output ([4.2.1] We evaluated our models to determine whether a given sentence contains at least one slang usage). Elisco, Esponda, Erb, and Pei are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Esponda, and Erb to combine the teachings of Pei because doing so would help improve slang detection and classification in machine learning models by utilizing tokenization and word embeddings to help locate slang (Pei [5] For unknown tokens, character-based convolutional embeddings improve the model in handling novel slang terms. We demonstrate that features combined with distributed word embeddings help machine detection of slang in general, and that Part-of-Speech among others is a prominent feature of slang usage. Our work provides a basis for locating slang from its flexible and unconventional syntactic word uses and offers opportunities for slang processing in downstream tasks in natural language processing). Regarding claim 5, the combination of Elisco, Esponda, Erb, and Pei teaches: the method of claim 1. Elisco further teaches: wherein masking the at least one token of the plurality of tokens for each training data instance of the plurality of training data instances to create the plurality of masked training data instances comprises masking the at least one slang token included in the plurality of tokens for each training data instance in the second subset of the plurality of training data instances ([0048] During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token... a resulting model may be catered towards clinically-specific text and language used by case workers. Model may be able to learn slang, acronyms, synonyms, misspellings, jargon, and more which may be otherwise absent from generalized models). Regarding claim 6, the combination of Elisco, Esponda, Erb, and Pei teaches: the method of claim 1. Elisco further teaches: wherein the ML model comprises an encoder only transformer architecture ([0024] As used in this disclosure a “transformer model” is a deep learning model for processing sequential data, such as natural language, for tasks such as translation and text summarization. As a non-limiting example a transformer model may include pre-trained systems such as Bidirectional Encoder Representations from Transformers (BERT). BERT is an example of an encoder only transformer.). Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Elisco in view of Esponda, Erb, and Pei, as applied to claims 1 and 5-6 above, and further in view of Sewak et al. (US 20220414137 A1; hereinafter referred to as Sewak). Regarding claim 7, the combination of Elisco, Esponda, Erb, and Pei teaches: the method of claim 6. The combination of Elisco, Esponda, Erb, and Pei does not explicitly, but Sewak teaches: wherein the encoder-only transformer architecture comprises a decoding-enhanced bidirectional encoder representations from transformers with disentangled attention (DeBERTa) model ([0081] For better results, larger and more expressive models may be used. The models may preferably be pre-trained (concept of transfer learning where models learn partially from large unsupervised and unlabelled data) and further fine-tuned with data, preferably from similar domains as application requirements. Some examples of similar models could be (but not limited to) GPT-3, Microsoft DeBerta etc, preferably models with a good zero-shot generative capabilities mode (a mode in which the model could generate text without fine-tuning with specific type of data)). Elisco, Esponda, Erb, Pei, and Sewak are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Esponda, Erb, and Pei to combine the teachings of Sewak because doing so would allow for the use of a DeBERTa model, which would improve the efficiency for training a ML slang detection model (Sewak [0081] Generative NLP models available in a repository 162 are loaded (or remains pre-loaded throughout). For better results, larger and more expressive models may be used. The models may preferably be pre-trained (concept of transfer learning where models learn partially from large unsupervised and unlabelled data) and further fine-tuned with data, preferably from similar domains as application requirements). Claims 8 are rejected under 35 U.S.C. 103 as being unpatentable over Elisco in view of Pei , Esponda, and Sackett et al. (US 20230147359 A1; hereinafter referred to as Sackett). Regarding claim 8, Elisco discloses: a method of training a machine learning (ML) model to mitigate slang usage, comprising: masking at least one token of a plurality of tokens for each respective first training data instance of a plurality of first training data instances to create a plurality of masked training data instances… ([0048] During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token, a second proportion such as 10% may be replaced with a random token, and/or remaining tokens may be kept the same. A task in training may be to predict an original masked token); training the ML model, using a masked-language modeling head, to predict the at least one token masked for each of the plurality of masked training data instances using the plurality of first training data instances and the plurality of masked training data instances… ([0048] A task in training may be to predict an original masked token… Only masked tokens in a note may contribute to a loss function which influences how to change network parameters to improve performance. A model being trained may learn which tokens appear in a similar context for a given dataset); and for each of the plurality of second training data instances, training the ML model, using a causal language modeling head ([0024] a transformer model may include pre-trained systems such as Bidirectional Encoder Representations from Transformers (BERT) and/or Generative Pre-trained Transformer (GPT). GPT inherently contains a causal language modeling head.), to predict the second plurality of tokens from the first plurality of tokens ([0048] During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token, a second proportion such as 10% may be replaced with a random token, and/or remaining tokens may be kept the same. A task in training may be to predict an original masked token). Elisco does not explicitly, but Pei teaches: wherein: the plurality of tokens for each respective first training data instance in a first subset of the plurality of first training data instances ([4.1] We consider datasets that are composed of sentences in two distinct categories, standard (slangless) and slang-specific) does not include any slang token of a plurality of slang tokens ([4.1] The sentences from Wall Street News are taken to be non-slang sentences since the news-based sentences were typically standard English conformed and reviewed before publication. In order to construct an even more trustworthy negative set for standard English, we filtered the sentences from Wall Street News based on the proportion of unknown tokens within the sentences), and the plurality of tokens for each respective first training data instance in a second subset of the plurality of first training data instances includes at least one slang token of the plurality of slang tokens… ([4.1] We collect positive examples from lexical entries in the Online Slang Dictionary (OSD) where example usage sentences are available). Elisco and Pei are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco to combine the teachings of Pei because doing so would help improve slang detection and classification in machine learning models by utilizing tokenization and word embeddings to help locate slang (Pei [5] For unknown tokens, character-based convolutional embeddings improve the model in handling novel slang terms. We demonstrate that features combined with distributed word embeddings help machine detection of slang in general, and that Part-of-Speech among others is a prominent feature of slang usage. Our work provides a basis for locating slang from its flexible and unconventional syntactic word uses and offers opportunities for slang processing in downstream tasks in natural language processing). The combination of Elisco and Pei does not explicitly, but Esponda teaches: obtaining a plurality of second training data instances, wherein each respective second training data instance of the plurality of second training data instances comprises: a training input comprising a first plurality of tokens including at least one slang token of the plurality of slang tokens ([0046] Contextual definition inference engine 56 is programmed to receive an acronym 52 and a usage context 54 from online system 50, while online system 50 is displaying a first user interface screen, UI.sub.A. UI.sub.A is, for example, a document editing screen. Contextual definition inference engine 56 is programmed to execute operations of process 100A, described above, to resolve the acronym definition. In doing so, contextual definition inference engine is programmed to interface with digital dictionary 60 to determine candidate definitions and the associated usage contexts. An acronym can be slang.); and a training output comprising a second plurality of tokens that does not include the any slang token of the plurality of slang tokens… ([0037] In an embodiment, the model is trained using a data set that includes previously-determined acronym definition-usage context pairs. The trained model is then used by the inference engine to predict a context-relevant definition for the current use of the acronym). Elisco, Pei, and Esponda are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco and Pei to combine the teachings of Esponda because doing so would allow for slang-definition training datasets to be used for detecting slang instances, such as abbreviations/acronyms, by incorporating contextual information to determine the definition of a slang term, leading to more accurate slang detection (Esponda [0019] The disclosed technologies improve upon these prior approaches by automatically inferring the relevant definition based on the context of an application or work process. Thus, the disclosed technologies do not require users to explicitly provide context information in order to suggest a definition for an abbreviation or an acronym that is likely to be relevant to the current context). The combination of Elisco, Pei, and Esponda does not explicitly, but Sackett teaches: wherein an entailment score between the training input and the training output of the respective second training data instance is greater than a threshold… ([0067] Each entailment value 218 indicates whether the input phrase 202 entails, contradicts, or is neutral with respect to the respective example phrase… the highest confidence score or a number of highest confidence scores (e.g., top three scores, all scores above a threshold score, etc.) can be used to generate the entailment classification data 222). Elisco, Pei, Esponda, and Sackett are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Pei, and Esponda to combine the teachings of Sackett because doing so would allow for better classification of slang terms using entailment scores, leading to better detection of slang terms in training (Sackett [0044] Certain aspects of the present disclosure, including the use of an entailment classifier alone or combined with at least one of a pattern matching classifier (e.g., RegEx classifier) and an SML classifier (e.g., a BERT classifier), provide specific improvements to the technological process of interpreting user input and generating appropriate responses in natural language processing, especially with respect to chatbots and automated psychological therapy, and especially with respect to fields where subject-matter granularity is important). Regarding claim 11, the combination of Elisco, Pei, Esponda, and Sackett teaches: the method of claim 8. Elisco further teaches: wherein masking the at least one token of the plurality of tokens for each respective first training data instance of the plurality of first training data instances to create the plurality of masked training data instances comprises randomly masking at least a first percentage of the plurality of tokens for each respective first training data instance ([0048] Language modeling 320 may involve learning a probability distribution over a sequence of words, which probability distribution may be used to characterize relationships between words, for instance and without limitation as captured by geometric relationships between vectors as described in this disclosure. During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token). Regarding claim 12, the combination of Elisco, Pei, Esponda, and Sackett teaches: the method of claim 8. Elisco further teaches: wherein masking the at least one token of the plurality of tokens for each respective first training data instance of the plurality of first training data instances to create the plurality of masked training data instances comprises masking the at least one slang token included in the plurality of tokens for each respective first training data instance in the second subset of the plurality of first training data instances ([0048] During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token... a resulting model may be catered towards clinically-specific text and language used by case workers. Model may be able to learn slang, acronyms, synonyms, misspellings, jargon, and more which may be otherwise absent from generalized models). Regarding claim 13, the combination of Elisco, Pei, Esponda, and Sackett teaches: the method of claim 8. Elisco further teaches: wherein the ML model comprises an encoder-decoder transformer architecture ([0073] Encoder element 704 first receives an input 716, wherein an “input” is any textual representation, audiographic representation, and/or videographic representation. Input 716 is entered to a multi-head-attention 708 such that encoder 704 produces output encodings that are provided to the next encoder element and/or a decoder element 720. As used in this disclosure a decoder element 720 is an element that decodes the output encodings from encoder element 704 to provide output probabilities 724). Claims 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Elisco in view of Pei, Esponda, and Sackett, as applied to claims 8 and 11-13 above, and further in view of Erb. Regarding claim 9, the combination of Elisco, Pei, Esponda, and Sackett teaches: the method of Claim 8. Esponda further teaches: wherein: each respective first training data instance in the first subset of the plurality of first training data instances comprises at least a definition associated with a slang token in the plurality of slang tokens… ([0037] the model is created using a statistical machine learning process that builds the model by learning mathematical relationships between features of the usage context and corresponding definitions. In an embodiment, the model is trained using a data set that includes previously-determined acronym definition-usage context pairs. The trained model is then used by the inference engine to predict a context-relevant definition for the current use of the acronym. Acronyms can represent slang.); Sackett further teaches: and determining, by an entailment model, that an entailment score between the definition of the respective first training data instance and the plurality of output tokens satisfies a threshold ([0067] Each entailment value 218 indicates whether the input phrase 202 entails, contradicts, or is neutral with respect to the respective example phrase… the highest confidence score or a number of highest confidence scores (e.g., top three scores, all scores above a threshold score, etc.) can be used to generate the entailment classification data 222). The combination of Elisco, Pei, Esponda, and Sackett does not explicitly, but Erb teaches: and the method further comprises, generating the second subset of the plurality of first training data instances by, for each respective first training data instance in the first subset of the plurality of first training data instances: generating, by a large language model (LLM), a plurality of output tokens based on a prompt comprising the respective first training data instance and the slang token associated with the respective first training data instance ([0080] the server determines a suitable input prompt for instructing the generative AI model to output a definition for the selected words. The input prompt includes at least a portion of the selected text and instructions for generating supplementary data (i.e., definitions, summary, etc.) associated with the selected text); determining the plurality of output tokens comprise the slang token… ([0082] the server receives an output of the generative AI model. The output may include a definition for a single word, a group of multiple words, and/or one or more phrases contained in the user-selected text. Additionally, or alternatively, the output may include a summary of a passage of text selected by the user). Elisco, Pei, Esponda, Sackett, and Erb are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Pei, Esponda, and Sackett to combine the teachings of Erb because doing so would allow for a set of training data to be generated using a prompt to an LLM for a definition of a term (such as slang), leading to improved training of ML model for slang detection (Erb [0043] the definitions having the highest ratings/ranking may be prioritized in responding to user requests for definitions of selected terms. If a certain definition (or group of definitions) for a term has a statistically significant rating advantage, the alternative (i.e., non-preferred) definitions may be deleted from the database. The computing system may obtain new definitions of terms and store them in place of the deleted definitions (up to a defined limit on number of definitions). Additionally, or alternatively, the preferred definitions may be used as part of (i.e., included in) input prompts for obtaining new definitions of terms). Regarding claim 10, the combination of Elisco, Pei, Esponda, and Sackett teaches: the method of Claim 8. Esponda further teaches: each respective first training data instance in the second subset of the plurality of first training data instances comprises at least a definition of a slang token in the plurality of slang tokens and a first sentence comprising the slang token… ([0037] the model is created using a statistical machine learning process that builds the model by learning mathematical relationships between features of the usage context and corresponding definitions. In an embodiment, the model is trained using a data set that includes previously-determined acronym definition-usage context pairs. The trained model is then used by the inference engine to predict a context-relevant definition for the current use of the acronym. Acronyms can represent slang.) Sackett further teaches: and determining, by an entailment model, that an entailment score between the definition of the respective first training data instance and the plurality of output tokens satisfies a threshold ([0067] Each entailment value 218 indicates whether the input phrase 202 entails, contradicts, or is neutral with respect to the respective example phrase… the highest confidence score or a number of highest confidence scores (e.g., top three scores, all scores above a threshold score, etc.) can be used to generate the entailment classification data 222). The combination of Elisco, Pei, Esponda, and Sackett does not explicitly, but Erb teaches: and the method further comprises, generating the first subset of the plurality of first training data instances by, for each respective first training data instance in the second subset of the plurality of first training data instances: generating, by a large language model (LLM), a plurality of output tokens based on a prompt comprising the first sentence of the respective first training data instance comprising the slang token ([0080] the server determines a suitable input prompt for instructing the generative AI model to output a definition for the selected words. The input prompt includes at least a portion of the selected text and instructions for generating supplementary data (i.e., definitions, summary, etc.) associated with the selected text); determining the plurality of output tokens do not comprise the slang token… ([0082] the server receives an output of the generative AI model. The output may include a definition for a single word, a group of multiple words, and/or one or more phrases contained in the user-selected text. Additionally, or alternatively, the output may include a summary of a passage of text selected by the user). Claims 14, 20, and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Elisco in view of Pei , Walia (US 20170116177 A1), and Sackett. Regarding claim 14, Elisco discloses: a method of slang detection and mitigation, comprising: receiving an input sentence comprising a plurality of tokens… ([0046] textual input 304 such as a document from a current document sequence as described below may be received by neural network 108. Computing device 104 may tokenize textual input 304 prior to provision to base network 204). Elisco does not explicitly, but Pei teaches: processing, with a first machine learning (ML) model trained for slang classification, the input sentence comprising the plurality of tokens and thereby generating at least one of: a first classification output for the input sentence, the first classification output comprising a slang instance classification ([4.2.1.] We evaluated our models to determine whether a given sentence contains at least one slang usage); or a second classification output for each of the plurality of tokens of the input sentence, at least one second classification output comprising a slang token classification… ([3.1] In the slang identification task, our models identify each token within the input sentence as ‘nonslang’ or ‘slang’ by sequence labeling, which determines the exact positions of slang usage). Elisco and Pei are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco to combine the teachings of Pei because doing so would help improve slang detection and classification in machine learning models by utilizing tokenization and word embeddings to help locate slang (Pei [5] For unknown tokens, character-based convolutional embeddings improve the model in handling novel slang terms. We demonstrate that features combined with distributed word embeddings help machine detection of slang in general, and that Part-of-Speech among others is a prominent feature of slang usage. Our work provides a basis for locating slang from its flexible and unconventional syntactic word uses and offers opportunities for slang processing in downstream tasks in natural language processing). The combination of Elisco and Pei does not explicitly, but Walia teaches: and based on at least one of the first classification output comprising the slang instance classification or the at least one second classification output comprising the slang token classification ([0052] The short form replacement module 214 is configured to replace abbreviations, slangs and acronyms in the textual data with corresponding full-word representations. For example, words “good”, “gd” or “gooood” may be normalized to “good” and words like “I'll” may be normalized to “I will”. Further, abbreviation substitutions may include substituting the word “account” for “acc”, “credit card” for “cc”, and so on. The normalization module includes a shortform replacement module that determines and classifies slang.), processing with a second ML model ([0045] The normalization module 140 is depicted to include a regularly used expression module 202, a character removal module 204, a symbol substitution module 206, a word class substitution module 208, a stemming module 210, a stop-word removal module 212, a short form replacement module 214, a white space removal module 216 and a spell checker module 218) trained for slang mitigation ([0058] the flow 300 includes removing non-English characters from the textual data. At operation 308, the flow 300 includes substituting symbols, abbreviations, slangs and acronyms with equivalent word representations. The removal of non-English characters and the substitution of symbols, abbreviations, slangs and acronyms may be performed), the input sentence to generate an output sentence… ([0067] the flow 400 includes correcting at least one word in the one or more sentences of the textual data corresponding to the natural language communication. In an embodiment, the correction can be based on the comparison of the SLM log probabilities of a word with those of suggestions for the word. In at least one example embodiment, the correction of a word may involve replacing the word with the highest scored suggestion); and the output sentence does not comprise any slang tokens ([0064, 0068] The natural language communication may include one or more sentences of textual data in partially normalized form (for example, the one or more sentences in the natural language communication may have differently expressed regularly used expressions replaced; slangs, abbreviations, symbols, acronyms substituted… the flow 400 may further include outputting the normalized one or more sentences of the textual data). Elisco, Pei, and Walia are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco and Pei to combine the teachings of Walia because doing so would allow for the detection and removal of slang from an input sentence based on a text normalization model that classifies slang words, leading to more accurate slang detection and mitigation in sentences (Walia [0055] the normalization of textual data may be performed based on a variety of functions including standardized functions and those based on client specification/preference. For example, enterprises may provide a specific word list related to client products and/or services to be exempted during spell checking and so on and so forth. In at least one example embodiment, a default ordering of operations to be performed for normalization of textual data can be defined. One example sequence of operations for normalization of textual data includes: replace email addresses, replace URLs, replace special symbols, replace regular expressions (time, date, dollar amount, etc.), replace string-lookup based word classes, abbreviations, and symbols, remove white spaces, and spell checking. The order in which the processing operations for normalization of textual data are sequenced may be fixed or may be customized by a user of the apparatus). The combination of Elisco, Pei, and Walia does not explicitly, but Sackett teaches: wherein: an entailment score between the input sentence and the output sentence satisfies a threshold… ([0067] Each entailment value 218 indicates whether the input phrase 202 entails, contradicts, or is neutral with respect to the respective example phrase… the highest confidence score or a number of highest confidence scores (e.g., top three scores, all scores above a threshold score, etc.) can be used to generate the entailment classification data 222). Elisco, Pei, Walia, and Sackett are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Pei, and Walia to combine the teachings of Sackett because doing so would allow for better classification of slang terms using entailment scores, leading to better detection of slang terms in training (Sackett [0044] Certain aspects of the present disclosure, including the use of an entailment classifier alone or combined with at least one of a pattern matching classifier (e.g., RegEx classifier) and an SML classifier (e.g., a BERT classifier), provide specific improvements to the technological process of interpreting user input and generating appropriate responses in natural language processing, especially with respect to chatbots and automated psychological therapy, and especially with respect to fields where subject-matter granularity is important). Regarding claim 20, the combination of Elisco, Pei, Walia, and Sackett teaches: the method of claim 14. Elisco further teaches: prior to using the first ML model to process the input sentence to generate at least one of the first classification output or the second classification output: training the first ML model, using a masked-language modeling head, to predict at least one token masked for each respective masked training data instance of a plurality of masked training data instances… ([0048] During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token, a second proportion such as 10% may be replaced with a random token, and/or remaining tokens may be kept the same. A task in training may be to predict an original masked token); and training the first ML model, using a classification head and a plurality of non-masked training data instances ([0048] some sub-network processing may include token-level classification 312, which may classify tokens to outputs of interest, and/or sentence-level classification 316, which may classify sentences to outputs of interest… During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token, a second proportion such as 10% may be replaced with a random token, and/or remaining tokens may be kept the same), to generate an instance-level slang classification or a token-level slang classification ([0048] a resulting model may be catered towards clinically-specific text and language used by case workers. Model may be able to learn slang, acronyms, synonyms, misspellings, jargon, and more which may be otherwise absent from generalized models). Pei further teaches: wherein: each respective masked training data instance of a first subset of the plurality of masked training data instances does not include any slang token of a plurality of slang tokens ([4.1] The sentences from Wall Street News are taken to be non-slang sentences since the news-based sentences were typically standard English conformed and reviewed before publication. In order to construct an even more trustworthy negative set for standard English, we filtered the sentences from Wall Street News based on the proportion of unknown tokens within the sentences), and each respective masked training data instance of a second subset of the plurality of masked training data instances comprises at least one slang token of the plurality of slang tokens… ([4.1] We collect positive examples from lexical entries in the Online Slang Dictionary (OSD) where example usage sentences are available). Regarding claim 23, the combination of Elisco, Pei, Walia, and Sackett teaches: the method of claim 20. Elisco further teaches: wherein the at least one token masked for each respective masked training data instance of a plurality of masked training data instances is randomly selected among a plurality of tokens associated each respective masked training data instance of a plurality of masked training data instances ([0048] Language modeling 320 may involve learning a probability distribution over a sequence of words, which probability distribution may be used to characterize relationships between words, for instance and without limitation as captured by geometric relationships between vectors as described in this disclosure. During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token). Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Elisco in view of Pei, Walia, and Sackett, as applied to claims 14, 20, and 23 above, and further in view of Lancioni et al. (US 20250117593 A1; hereinafter referred to as Lancioni). Regarding claim 16, the combination of Elisco, Pei, Walia, and Sackett teaches: the method of claim 14. The combination of Elisco, Pei, Walia, and Sackett does not explicitly, but Lancioni teaches: further comprising using the output sentence to perform one or more tasks ([0022] Moreover, in some examples, the output data may undergo post-processing after it is generated by the model to transform the output into a useful result (e.g., a display of data, an instruction to be executed by a machine, etc.)). Elisco, Pei, Walia, Sackett, and Lancioni are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Pei, Walia, and Sackett to combine the teachings of Lancioni because doing so would allow for specific training data to generated by an LLM for the purpose of training and improving an ML model for slang detection (Lancioni [0124] methods have been disclosed that enable large language models to be utilized in various contexts while providing guardrails for the responses that are provided the LLM. Disclosed systems, apparatus, articles of manufacture, and methods improve the efficiency of using a computing device by limiting the amount of additional anti-hypothesis that are added to a prompt for generation of a response). Claims 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Elisco in view of Pei, Walia, and Sackett, as applied to claims 14, 20, and 23 above, and further in view of Tensmeyer et al. (US 20230334244 A1; hereinafter referred to as Tensmeyer). Regarding claim 17, the combination of Elisco, Pei, Walia, and Sackett teaches: the method of claim 14. Elisco further teaches: wherein: the first ML model comprises an encoder-only transformer architecture… ([0024] a transformer model may include pre-trained systems such as Bidirectional Encoder Representations from Transformers (BERT)). The combination of Elisco, Pei, Walia, and Sackett does not explicitly, but Tensmeyer teaches: and the second ML model comprises an encoder-decoder transformer architecture ([0049] a sequence-to-sequence training is used, where the fixer module 126 uses an encoder-decoder model to predict the probability of the next token given a context from a previous token. The output of the fixer module 126 can then be used to generate modified training sentence 612 with the assertion in the training sentence corresponding to masked training sentence 606). Elisco, Pei, Lancioni, Walia, and Tensmeyer are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Pei, Walia, and Sackett to combine the teachings of Tensmeyer because doing so would allow for use of an encoder-decoder transformer for more flexible training of a ML model using masked tokens (Tensmeyer [0036] the interaction of the user with the system can be used as a training signal to further improve the fact correction system 102. For example, when a user selects a suggested sentence from a ranked list, the selection can be used to further train the fact correction system 102 to rank that sentence first, or higher, in subsequent suggested corrections. In some embodiments, the data table used to modify the masked tokens in natural language sentence 105 can also be provided in output 130). Regarding claim 18, the combination of Elisco, Pei, Walia, and Sackett teaches: the method of claim 14. Pei further teaches: wherein: each respective masked training data instance of a first subset of the plurality of masked training data instances does not include any slang token of a plurality of slang tokens ([4.1] The sentences from Wall Street News are taken to be non-slang sentences since the news-based sentences were typically standard English conformed and reviewed before publication. In order to construct an even more trustworthy negative set for standard English, we filtered the sentences from Wall Street News based on the proportion of unknown tokens within the sentences), and each respective masked training data instance of a second subset of the plurality of masked training data instances comprises at least one slang token of the plurality of slang tokens ([4.1] We collect positive examples from lexical entries in the Online Slang Dictionary (OSD) where example usage sentences are available). The combination of Elisco, Pei, Walia, and Sackett does not explicitly, but Tensmeyer teaches: prior to using the second ML model to process the input sentence to generate the output sentence: training the second ML model, using a masked-language modeling head, to predict at least one token masked for each respective masked training data instance of a plurality of masked training data instances… ([0004] The fact correction system then predicts a new token for each of the one or more masked tokenized element based on the input sentence with the one or more masked tokenized element and the identified data table using a third machine learning model). Claims 19 is rejected under 35 U.S.C. 103 as being unpatentable over Elisco in view of Pei, Walia, Sackett, and Tensmeyer, as applied to claims 17-18 above, and further in view of Lancioni. Regarding claim 19, the combination of Elisco, Pei, Walia, Sackett, and Tensmeyer teaches: the method of claim 18. The combination of Elisco, Pei, Walia, Sackett, and Tensmeyer does not explicitly, but Lancioni teaches: training the second ML model, using a causal language modeling head ([0015] Many different types of machine learning models and/or machine learning architectures exist. In examples disclosed herein, a Large Language Model (LLM) such as ChatGPT is used. GPT models inherently contain a causal language modeling head.), to generate a second output sentence from a second output sentence including the at least one slang token of the plurality of slang tokens ([0061-0062] the message validator circuitry 170 may seek to determine whether inappropriate content is included in the response message, whether offensive language is included in the response message, whether the response message contains slang or unprofessionally written language.. a positive identification of the offensive language and slang wording may result in both anti-hypothesis being utilized to modify the original prompt). Elisco, Pei, Walia, Sackett, Tensmeyer, and Lancioni are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Pei, Walia, Sackett, and Tensmeyer to combine the teachings of Lancioni because doing so would allow for specific training data to generated by an LLM for the purpose of training and improving an ML model for slang detection (Lancioni [0124] methods have been disclosed that enable large language models to be utilized in various contexts while providing guardrails for the responses that are provided the LLM. Disclosed systems, apparatus, articles of manufacture, and methods improve the efficiency of using a computing device by limiting the amount of additional anti-hypothesis that are added to a prompt for generation of a response). Claims 21-22 is rejected under 35 U.S.C. 103 as being unpatentable over Elisco in view of Pei, Walia, and Sackett, as applied to claims 14, 20, and 23 above, and further in view of Esponda and Erb. Regarding claim 21, the combination of Elisco, Pei, Walia, and Sackett teaches: the method of Claim 20. Elisco further teaches: and masking one or more tokens of the plurality of output tokens ([0048] During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token, a second proportion such as 10% may be replaced with a random token, and/or remaining tokens may be kept the same. A task in training may be to predict an original masked token). Sackett further teaches: determining, by an entailment model, that an entailment score between the definition, without the at least one token masked, of the respective masked training data instance and the plurality of output tokens satisfies a threshold… ([0067] Each entailment value 218 indicates whether the input phrase 202 entails, contradicts, or is neutral with respect to the respective example phrase… the highest confidence score or a number of highest confidence scores (e.g., top three scores, all scores above a threshold score, etc.) can be used to generate the entailment classification data 222). The combination of Elisco, Pei, Walia, and Sackett does not explicitly, but Esponda teaches: wherein: each respective masked training data instance in the first subset of the plurality of masked training data instances comprises at least a definition associated with a slang token in the plurality of slang tokens, wherein at least one token of the definition is masked… ([0037] the model is created using a statistical machine learning process that builds the model by learning mathematical relationships between features of the usage context and corresponding definitions. In an embodiment, the model is trained using a data set that includes previously-determined acronym definition-usage context pairs. The trained model is then used by the inference engine to predict a context-relevant definition for the current use of the acronym. Acronyms can represent slang.). Elisco, Pei, Walia, Sackett, and Esponda are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Pei, Walia, and Sackett to combine the teachings of Esponda because doing so would allow for slang-definition training datasets to be used for detecting slang instances, such as abbreviations/acronyms, by incorporating contextual information to determine the definition of a slang term, leading to more accurate slang detection (Esponda [0019] The disclosed technologies improve upon these prior approaches by automatically inferring the relevant definition based on the context of an application or work process. Thus, the disclosed technologies do not require users to explicitly provide context information in order to suggest a definition for an abbreviation or an acronym that is likely to be relevant to the current context). The combination of Elisco, Pei, Walia, Sackett, and Esponda does not explicitly, but Erb teaches: generating, by a large language model (LLM), a plurality of output tokens based on a prompt comprising the definition, without the at least one token masked, for the respective masked training data instance and the slang token associated with the respective masked training data instance ([0080] the server determines a suitable input prompt for instructing the generative AI model to output a definition for the selected words. The input prompt includes at least a portion of the selected text and instructions for generating supplementary data (i.e., definitions, summary, etc.) associated with the selected text); determining the plurality of output tokens comprise the slang token… ([0082] the server receives an output of the generative AI model. The output may include a definition for a single word, a group of multiple words, and/or one or more phrases contained in the user-selected text. Additionally, or alternatively, the output may include a summary of a passage of text selected by the user). Elisco, Pei, Walia, Sackett, Esponda, and Erb are considered analogous in the field of natural language processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Elisco, Pei, Walia, Sackett, and Esponda to combine the teachings of Erb because doing so would allow for a set of training data to be generated using a prompt to an LLM for a definition of a term (such as slang), leading to improved training of ML model for slang detection (Erb [0043] the definitions having the highest ratings/ranking may be prioritized in responding to user requests for definitions of selected terms. If a certain definition (or group of definitions) for a term has a statistically significant rating advantage, the alternative (i.e., non-preferred) definitions may be deleted from the database. The computing system may obtain new definitions of terms and store them in place of the deleted definitions (up to a defined limit on number of definitions). Additionally, or alternatively, the preferred definitions may be used as part of (i.e., included in) input prompts for obtaining new definitions of terms). Regarding claim 22, the combination of Elisco, Pei, Walia, and Sackett teaches: the method of Claim 20. Elisco further teaches: and masking one or more tokens of the plurality of output tokens ([0048] During training, a first proportion such as 15% of tokens in a textual input may be replaced with a special mask token, a second proportion such as 10% may be replaced with a random token, and/or remaining tokens may be kept the same. A task in training may be to predict an original masked token). Sackett further teaches: determining, by an entailment model, that an entailment score between the first sentence, without the at least one token masked, and the plurality of output tokens satisfies a threshold… ([0067] Each entailment value 218 indicates whether the input phrase 202 entails, contradicts, or is neutral with respect to the respective example phrase… the highest confidence score or a number of highest confidence scores (e.g., top three scores, all scores above a threshold score, etc.) can be used to generate the entailment classification data 222). The combination of Elisco, Pei, Walia, and Sackett does not explicitly, but Esponda teaches: wherein: each respective masked training data instance in the second subset of the plurality of masked training data instances comprises at least a first sentence comprising a slang token in the plurality of slang tokens, wherein at least one token of the first sentence is masked… ([0037] the model is created using a statistical machine learning process that builds the model by learning mathematical relationships between features of the usage context and corresponding definitions. In an embodiment, the model is trained using a data set that includes previously-determined acronym definition-usage context pairs. The trained model is then used by the inference engine to predict a context-relevant definition for the current use of the acronym. Acronyms can represent slang.). The combination of Elisco, Pei, Walia, Sackett, and Esponda does not explicitly, but Erb teaches: and the method further comprises, generating the first subset of the plurality of masked training data instances by, for each respective masked training data instance in the second subset of the plurality of masked training data instances: generating, by a large language model (LLM), a plurality of output tokens based on a prompt comprising the first sentence comprising the slang token, without the at least one token masked, of the respective masked training data instance ([0080] the server determines a suitable input prompt for instructing the generative AI model to output a definition for the selected words. The input prompt includes at least a portion of the selected text and instructions for generating supplementary data (i.e., definitions, summary, etc.) associated with the selected text); determining the plurality of output tokens do not comprise the slang token… ([0082] the server receives an output of the generative AI model. The output may include a definition for a single word, a group of multiple words, and/or one or more phrases contained in the user-selected text. Additionally, or alternatively, the output may include a summary of a passage of text selected by the user). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Nathan Tengbumroong whose telephone number is (703)756-1725. The examiner can normally be reached Monday - Friday, 11:30 am - 8:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NATHAN TENGBUMROONG/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Jan 26, 2024
Application Filed
Jan 09, 2026
Non-Final Rejection mailed — §103
Mar 30, 2026
Applicant Interview (Telephonic)
Mar 30, 2026
Examiner Interview Summary
Apr 09, 2026
Response Filed
Jul 29, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682177
GENERATING GOAL-ORIENTED DIALOGUES FROM DOCUMENTS
4y 0m to grant Granted Jul 14, 2026
Patent 12675510
SYSTEMS AND METHODS FOR PROVIDING USER INTERFACES TO CONVERSE WITH A CORPUS OF ELECTRONIC DOCUMENTS VIA A LARGE LANGUAGE MODEL
3y 1m to grant Granted Jul 07, 2026
Patent 12658181
INTERACTIVE DECODING OF WORDS FROM PHONEME SCORE DISTRIBUTIONS
2y 8m to grant Granted Jun 16, 2026
Patent 12640161
METHOD AND APPARATUS FOR PROCESSING AUDIO FOR SCENE CLASSIFICATION
3y 0m to grant Granted May 26, 2026
Patent 12530536
Mixture-Of-Expert Approach to Reinforcement Learning-Based Dialogue Management
2y 11m to grant Granted Jan 20, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
42%
Grant Probability
75%
With Interview (+33.3%)
3y 0m (~6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 24 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month