Prosecution Insights
Last updated: August 17, 2026
Application No. 18/668,086

CONSISTENCY EVALUATION FOR DOCUMENT SUMMARIES USING LANGUAGE MODEL NEURAL NETWORKS

Non-Final OA §101§103
Filed
May 17, 2024
Priority
May 18, 2023 — provisional 63/467,568
Examiner
HAN, KYU HYUNG
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
50%
Grant Probability
Moderate
1-2
OA Rounds
1y 10m
Est. Remaining
79%
With Interview

Examiner Intelligence

Grants 50% of resolved cases
50%
Career Allowance Rate
7 granted / 14 resolved
-10.0% vs TC avg
Strong +29% interview lift
Without
With
+29.2%
Interview Lift
resolved cases with interview
Typical timeline
4y 1m
Avg Prosecution
24 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
31.9%
-8.1% vs TC avg
§103
58.9%
+18.9% vs TC avg
§102
2.7%
-37.3% vs TC avg
§112
6.5%
-33.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 14 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections – 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Claims 1-14 are method claims. Claims 15-20 are machine/system/product claims. Therefore, claims 1-20 are directed to either a process, machine, manufacture or composition of matter. With respect to claim 1: Step 2A – Prong 1: … … for each document, processing content from the document using a respective plurality of generative summarization models to generate a plurality of summaries of the documents; (mental process – a person can manually process content from the document using a respective plurality of generative summarization models to generate a plurality of summaries of the documents with the assistance of a pen/paper.) for each summary of each of the plurality of documents: generating a respective input sequence that comprises (i) a natural language instruction to evaluate a consistency of the summary with the document, (mental process – a person can manually generate a respective input sequence that comprises a natural language instruction to evaluate a consistency of the summary with the document with the assistance of a pen/paper.) (ii) the content from the document, (mental process – a person can manually generate a respective input sequence that comprises the content from the document with the document with the assistance of a pen/paper.) and (iii) the summary of the document; (mental process – a person can manually generate a respective input sequence that comprises the summary of the document with the assistance of a pen/paper.) processing the input sequence using a language model neural network to generate an output that rates the consistency of the summary with the document; (mental process – a person can manually process the input sequence using a language model neural network to generate an output that rates the consistency of the summary with the document with the assistance of a pen/paper.) and generating a training example that includes the content from the document, the summary of the document, and the output that rates the consistency of the summary with the document. (mental process – a person can manually generate a training example that includes the content from the document, the summary of the document, and the output that rates the consistency of the summary with the document with the assistance of a pen/paper.) Step 2A – Prong 2: This judicial exception is not integrated into a practical application. A method performed by one or more computers, the method comprising: (mere instructions to apply the exception using a generic computer component – computer applies exception.) obtaining a plurality of documents; (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)). Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. A method performed by one or more computers, the method comprising: (mere instructions to apply the exception using a generic computer component – computer applies exception.) obtaining a plurality of documents; (MPEP 2106.05(d)(II) indicate that merely “Receiving or transmitting data over a network, e.g., using the Internet to gather data” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim – the documents are merely received). Thereby, a conclusion that the claimed distribute step is well-understood, routine, conventional activity is supported under Berkheimer.) With respect to claim 2: Step 2A – Prong 1: … … and generate an output that rates the consistency of the input summary with the input document, (mental process – a person can manually generate an output that rates the consistency of the input summary with the input document with the assistance of a pen/paper.) and wherein the training comprises, for each training example, using the output that rates the consistency of the summary with the document as a target output for the consistency evaluation neural network. (mental process – a person can recognize that the training comprises, for each training example, using the output that rates the consistency of the summary with the document as a target output for the consistency evaluation neural network.) Step 2A – Prong 2: This judicial exception is not integrated into a practical application. The method of claim 1, further comprising: training a consistency evaluation neural network on the training examples for the summaries of the plurality of documents, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: High level recitation of training a neural network to create summaries of documents.); wherein the consistency evaluation neural network is configured to receive an input that comprises content from an input document and an input summary of the input document (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)). … … Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. The method of claim 1, further comprising: training a consistency evaluation neural network on the training examples for the summaries of the plurality of documents, (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: High level recitation of training a neural network to create summaries of documents.); wherein the consistency evaluation neural network is configured to receive an input that comprises content from an input document and an input summary of the input document (MPEP 2106.05(d)(II) indicate that merely “Receiving or transmitting data over a network, e.g., using the Internet to gather data” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim – the input content from document are merely received). Thereby, a conclusion that the claimed distribute step is well-understood, routine, conventional activity is supported under Berkheimer.) … … With respect to claim 3: Step 2A – Prong 1: The method of claim 1, wherein the output that rates the consistency of the summary with the document is a score that rates the consistency of the summary with the document. (mental process – a person can recognize that the output that rates the consistency of the summary with the document is a score that rates the consistency of the summary with the document.) With respect to claim 4: Step 2A – Prong 1: The method of claim 3, wherein the score that rates the consistency of the summary with the document is a binary score that assigns a first value to being consistent with the document and a second value to being inconsistent with the document. (mental process – a person can recognize that the score that rates the consistency of the summary with the document is a binary score that assigns a first value to being consistent with the document and a second value to being inconsistent with the document.) With respect to claim 5: Step 2A – Prong 1: The method of claim 1, wherein the language model neural network has been trained on a language modeling objective. (mental process – a person can recognize that the language model neural network has been trained on a language modeling objective.) With respect to claim 6: Step 2A – Prong 1: The method of claim 5, wherein the language model neural network has been fine-tuned on one or more natural language instruction following objectives. (mental process – a person can recognize that the language model neural network has been fine-tuned on one or more natural language instruction following objectives.) With respect to claim 7: Step 2A – Prong 1: The method of claim 6, wherein the language model neural network has not been fine-tuned or trained on any objective that requires evaluating consistency of summaries with corresponding documents. (mental process – a person can recognize that the language model neural network has not been fine-tuned or trained on any objective that requires evaluating consistency of summaries with corresponding documents.) With respect to claim 8: Step 2A – Prong 1: The method of claim 1, wherein the language model neural network is a large language model neural network that has more than 10 billion parameters. (mental process – a person can recognize that the language model neural network is a large language model neural network that has more than 10 billion parameters.) With respect to claim 9: Step 2A – Prong 1: The method of claim 8, wherein the language model neural network has more than 100 billion parameters. (mental process – a person can recognize that the language model neural network is a large language model neural network that has more than 100 billion parameters.) With respect to claim 10: Step 2A – Prong 1: The method of claim 8, wherein the language model neural network has more than 500 billion parameters. (mental process – a person can recognize that the language model neural network is a large language model neural network that has more than 500 billion parameters.) With respect to claim 11: Step 2A – Prong 1: The method of claim 1, wherein the plurality of documents include one or more documents in each of a plurality of natural languages. (mental process – a person can recognize that the plurality of documents include one or more documents in each of a plurality of natural languages.) With respect to claim 12: Step 2A – Prong 1: The method of claim 1, wherein, for each document, the respective plurality of generative summarization models comprise (i) at least two models that have been trained on different data sets, (ii) at least two models that different numbers of parameters, or (iii) both. (mental process – a person can recognize that for each document, the respective plurality of generative summarization models comprise (i) at least two models that have been trained on different data sets, (ii) at least two models that different numbers of parameters, or (iii) both.) With respect to claim 13: Step 2A – Prong 1: The method of claim 1, … (mental process from claim 1) … … Step 2A – Prong 2: This judicial exception is not integrated into a practical application. … wherein the respective plurality of generative summarization models for each document comprise at least a subset of generative summarization models in a set of generative summarization models, the method further comprising: generating the set of generative summarization models, comprising: obtaining a plurality of training data sets; (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)). obtaining data specifying a plurality of un-trained generative summarization models, each of the un-trained generative summarization models having a different number of parameters; (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g)). and for each un-trained generative summarization model, training a respective instance of the un-trained generative summarization model on each of the plurality of training data sets to generate a corresponding trained generative summarization model. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: High level recitation of training a generative summarization model on each of the plurality of training data sets.); Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception. … wherein the respective plurality of generative summarization models for each document comprise at least a subset of generative summarization models in a set of generative summarization models, the method further comprising: generating the set of generative summarization models, comprising: obtaining a plurality of training data sets; (MPEP 2106.05(d)(II) indicate that merely “Receiving or transmitting data over a network, e.g., using the Internet to gather data” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim – training data sets are merely received). Thereby, a conclusion that the claimed distribute step is well-understood, routine, conventional activity is supported under Berkheimer.) obtaining data specifying a plurality of un-trained generative summarization models, each of the un-trained generative summarization models having a different number of parameters; (MPEP 2106.05(d)(II) indicate that merely “Receiving or transmitting data over a network, e.g., using the Internet to gather data” is a well‐understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim – the data specifying a plurality of un-trained generative summarization models are merely received). Thereby, a conclusion that the claimed distribute step is well-understood, routine, conventional activity is supported under Berkheimer.) and for each un-trained generative summarization model, training a respective instance of the un-trained generative summarization model on each of the plurality of training data sets to generate a corresponding trained generative summarization model. (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: High level recitation of training a generative summarization model on each of the plurality of training data sets.); With respect to claim 14: Step 2A – Prong 1: The method of claim 1, wherein the respective plurality of generative summarization models for each document are encoder-decoder Transformer neural networks. (mental process – a person can recognize that the respective plurality of generative summarization models for each document are encoder-decoder Transformer neural networks.) Claim 15 is substantially similar to claim 1, but has the following additional elements: With respect to claim 15: Step 2A – Prong 2: This judicial exception is not integrated into a practical application. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one more computers to perform operations comprising: (mere instructions to apply the exception using a generic computer component – computer applies exception.) Claims 16-18 are rejected on the same grounds under 35 U.S.C. 101 as claims 2-4 as they are substantially similar, respectively. Mutatis mutandis. Claim 19 is rejected on the same grounds under 35 U.S.C. 101 as claim 13 as they are substantially similar. Mutatis mutandis. Claim 20 is substantially similar to claim 1, but has the following additional elements: With respect to claim 20: Step 2A – Prong 1: One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one more computers to perform operations comprising: (mere instructions to apply the exception using a generic computer component – computer applies exception.) Claim Rejections – 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7, 11, 12, 14-18, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Kryscinski (US20210124876A1) hereinafter known as Kryscinski in view of Luo (“ChatGPT as a Factual Inconsistency Evaluator for Text Summarization”) hereinafter known as Luo. Regarding independent claim 1, Kryscinski teaches: A method performed by one or more computers, the method comprising: obtaining a plurality of documents; (Kryscinski [0045]: “training data generation module 130 receives an unannotated collection or set S of source documents” Kryscinski teaches obtaining documents.) for each document, processing content from the document using a respective plurality of generative summarization models to generate a plurality of summaries of the documents; (Kryscinski [0060]: “the manually annotated dataset utilizes summaries output by state-of-the-art summarization models, including extractive, abstractive, and hybrid approaches” Kryscinski teaches that summarization models process the documents to output summaries of the documents.) … … … … and generating a training example that includes the content from the document, the summary of the document, and the output that rates the consistency of the summary with the document. (Kryscinski [0055]: “the paraphrasing transformation example is a semantically invariant transformation, whereas the transformation examples for sentence negation, pronoun swap, entity swap, and number swap are semantically variant transformations … a novel claim sentence is labeled as CONSISTENT or CORRECT” Kryscinski teaches both semantically invariant and variant transformations, which are in regard to the summaries. Given the summaries and the content of the document, there may be a label of consistent.) Kryscinski does not explicitly teach: for each summary of each of the plurality of documents: generating a respective input sequence that comprises (i) a natural language instruction to evaluate a consistency of the summary with the document, (ii) the content from the document, and (iii) the summary of the document; processing the input sequence using a language model neural network to generate an output that rates the consistency of the summary with the document; However, Luo teaches: for each summary of each of the plurality of documents: generating a respective input sequence that comprises (i) a natural language instruction to evaluate a consistency of the summary with the document, (Luo [Page 3, Col. 2, Paragraph 4]: “We experiment with two different zero shot prompts in the NLI setting. The first one is based on direct assessment by directly asking ChatGPT to answer yes or no given the question. … Decide if the following summary is consistent with the corresponding article.” Luo teaches a generation of a prompt that requests a yes/no answer to the question of whether the summary is consistent with the document contents.) (ii) the content from the document, (Luo [Page 3, Col. 2, Paragraph 4]: “The first one is based on direct assessment by directly asking ChatGPT to answer yes or no given the question. … Article: [Article]” Luo teaches that in addition to generating the prompt, the LLM is provided the article, or content from the document.) and (iii) the summary of the document; (Luo [Page 3, Col. 2, Paragraph 4]: “The first one is based on direct assessment by directly asking ChatGPT to answer yes or no given the question. … Summary: [Summary]” Luo teaches that in addition to generating the prompt, the LLM is provided the summary of the document.) processing the input sequence using a language model neural network to generate an output that rates the consistency of the summary with the document; (Luo [Page 4, Col. 1, Paragraph 4]: “When processing the responses, we only consider solid judgment like "the summary is consistent with the article" as consistency, claims such as ‘partially consistent’ or ‘mostly consistent’ are all deemed as inconsistent.” Luo teaches that the LLM generated an output that rated the consistency as a gradient of consistency. Note that Luo did simplify further into consistent/nonconsistent categories.) Kryscinski and Luo are in the same field of endeavor as the present invention, as the references are directed to summarizing documents and determining the consistency of the summarization with the original documents using language models. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine generating a summary of a document as taught in Kryscinski with using a large language model to prompt and determine the consistency of the summary with the document as taught in Luo. Luo provides this additional functionality. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Kryscinski to include teachings of Luo because the combination would allow for language models to be prompted with determining the consistency of the summary. This has the potential benefit of determining whether the summaries of documents can be viable in important settings, such as in healthcare. Regarding dependent claim 2, Kryscinski and Luo teach: The method of claim 1, further comprising: training a consistency evaluation neural network on the training examples for the summaries of the plurality of documents, wherein the consistency evaluation neural network is configured to receive an input that comprises content from an input document and an input summary of the input document and generate an output that rates the consistency of the input summary with the input document, and wherein the training comprises, for each training example, using the output that rates the consistency of the summary with the document as a target output for the consistency evaluation neural network. (Kryscinski [0037]: “factual consistency module 150 can be trained—using the artificially generated training data set output from the training data generation module 130 and the annotated test set output from the data annotation module 140—for one or more tasks related to factual consistency verification. In some embodiments, these tasks include: 1) identifying whether sentences remain factually consistent after transformation” Kryscinski teaches that the factual consistency module takes in as input the document and the summary and outputs a label of whether the sentences remain consistent after transformation.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 3, Kryscinski and Luo teach: The method of claim 1, wherein the output that rates the consistency of the summary with the document is a score that rates the consistency of the summary with the document. (Luo [Page 5, Col. 2, Paragraph 4]: “we apply the consistency rating task on ChatGPT by asking it to mark the consistency of a summary with the reference to its source document on a scale from 1-10 points, where 1 point stands for total inconsistency, and 10 represents full consistency.” Luo teaches a rating of consistency of the summary with the document on a 10 point scale.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 4, Kryscinski and Luo teach: The method of claim 3, wherein the score that rates the consistency of the summary with the document is a binary score that assigns a first value to being consistent with the document and a second value to being inconsistent with the document. (Luo [Page 4, Col. 1, Paragraph 4]: “When processing the responses, we only consider solid judgment like "the summary is consistent with the article" as consistency, claims such as ‘partially consistent’ or ‘mostly consistent’ are all deemed as inconsistent.” Luo teaches that the LLM generated an output that rated the consistency as a gradient of consistency. Note that Luo did simplify further into binary consistent/nonconsistent categories.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 5, Kryscinski and Luo teach: The method of claim 1, wherein the language model neural network has been trained on a language modeling objective. (Luo [Page 1, Col. 2, Paragraph 2]: “large language models (LLMs), such as GPT-3” Luo teaches that the training is done on GPT, which is an LLM.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 6, Kryscinski and Luo teach: The method of claim 5, wherein the language model neural network has been fine-tuned on one or more natural language instruction following objectives. (Luo [Page 6, Col. 1, Paragraph 5]: “We conduct our experiments using the API of Chat GPT(gpt-3.5-turbo-0301) which is trained based on InstructGPT (Ouyang et al., 2022) with reinforce learning from human feedback (RLHF).” Luo teaches that the GPT is tuned using human feedback.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 7, Kryscinski and Luo teach: The method of claim 6, wherein the language model neural network has not been fine-tuned or trained on any objective that requires evaluating consistency of summaries with corresponding documents. (Luo [Page 6, Col. 1, Paragraph 6]: “ChatGPT is able to achieve comparable performance or even bet ter performance compared to the previous state of-the-art evaluation models without training on relevant tasks” Luo teaches that the GPT that was used was not previously trained on the task of evaluating consistency of summaries.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 11, Kryscinski and Luo teach: The method of claim 1, wherein the plurality of documents include one or more documents in each of a plurality of natural languages. (Kryscinski [0048]: “paraphrases are produced by backtranslation using Neural Machine Translation (NMT) systems, … an original sentence in English language is translated to an intermediate, non-English language” Kryscinski teaches that the documents may consist of more than one language, with a combination of English and non-English languages.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 12, Kryscinski and Luo teach: The method of claim 1, wherein, for each document, the respective plurality of generative summarization models comprise (i) at least two models that have been trained on different data sets, (ii) at least two models that different numbers of parameters, or (iii) both. (Kryscinski [0060]: In some embodiments, the manually annotated dataset utilizes summaries output by state-of-the-art summarization models, including extractive, abstractive, and hybrid approaches” Kryscinski teaches that the summary output may be a hybrid approach, which is training on two different models.) The reasons to combine are substantially similar to those of claim 1. Regarding dependent claim 14, Kryscinski and Luo teach: The method of claim 1, wherein the respective plurality of generative summarization models for each document are encoder-decoder Transformer neural networks. (Chowdhery [Page 3, Paragraph 45: “we trained a 540B parameter language model on 6144 TPU v4 chips at efficiency levels that could not be reached before for models of this scale” Chowdhery teaches that the language model is a transformer neural network.) The reasons to combine are substantially similar to those of claim 1. Claim 15 is substantially similar to claim 1, but has the following additional elements: Regarding independent claim 15, Kryscinski and Luo teach: A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one more computers to perform operations comprising: (Kryscinski [0031]: “Memory 120 may be used to store software executed by computing device” Kryscinski teaches memory and computer.) The reasons to combine are substantially similar to those of claim 1. Claims 16-18 are rejected on the same grounds under 35 U.S.C. 103 as claims 2-4 as they are substantially similar, respectively. Mutatis mutandis. Claim 20 is substantially similar to claim 1, but has the following additional elements: Regarding independent claim 20, Kryscinski and Luo teach: One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one more computers to perform operations comprising: (Kryscinski [0031]: “Memory 120 may be used to store software executed by computing device” Kryscinski teaches memory and computer.) The reasons to combine are substantially similar to those of claim 1. Claims 8-10, 13, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Kryscinski in view of Luo in view of Chowdhery et al. (“PaLM: Scaling Language Modeling with Pathways”) hereinafter known as Chowdhery. Regarding dependent claim 8, Kryscinski and Luo teach: The method of claim 1, Kryscinski and Luo do not explicitly teach: wherein the language model neural network is a large language model neural network that has more than 10 billion parameters. However, Chowdhery teaches: wherein the language model neural network is a large language model neural network that has more than 10 billion parameters. (Chowdhery [Page 3, Paragraph 45: “we trained a 540B parameter language model on 6144 TPU v4 chips at efficiency levels that could not be reached before for models of this scale” Chowdhery teaches that the transformer language model trained had 540 billion parameters.) Chowdhery is in the same field as the present invention, since it is directed to using high numbers of parameters on language models to train on specific tasks. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine prompting language models to determine the consistency of summaries of documents as taught in Kryscinski as modified by Luo with using models with high (over 500 billion) parameters as taught in Chowdhery. Chowdhery provides this additional functionality. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Kryscinski as modified by Luo to include teachings of Chowdhery because the combination would allow for the model to be weighted on a greater number of factors. This has the potential benefit of providing a more accurate judgement on whether a summary is consistent with its document or not. Regarding dependent claim 9, Kryscinski and Luo teach: The method of claim 8, wherein the language model neural network has more than 100 billion parameters. (Chowdhery [Page 3, Paragraph 45: “we trained a 540B parameter language model on 6144 TPU v4 chips at efficiency levels that could not be reached before for models of this scale” Chowdhery teaches that the transformer language model trained had 540 billion parameters.) The reasons to combine are substantially similar to those of claim 8. Regarding dependent claim 9, Kryscinski, Luo, and Chowdhery teach: The method of claim 8, wherein the language model neural network has more than 100 billion parameters. (Chowdhery [Page 3, Paragraph 45: “we trained a 540B parameter language model on 6144 TPU v4 chips at efficiency levels that could not be reached before for models of this scale” Chowdhery teaches that the transformer language model trained had 540 billion parameters.) The reasons to combine are substantially similar to those of claim 8. Regarding dependent claim 10, Kryscinski, Luo, and Chowdhery teach: The method of claim 9, wherein the language model neural network has more than 500 billion parameters. (Chowdhery [Page 3, Paragraph 45: “we trained a 540B parameter language model on 6144 TPU v4 chips at efficiency levels that could not be reached before for models of this scale” Chowdhery teaches that the transformer language model trained had 540 billion parameters.) The reasons to combine are substantially similar to those of claim 8. Regarding dependent claim 13, Kryscinski, Luo, and Chowdhery teach: The method of claim 1, wherein the respective plurality of generative summarization models for each document comprise at least a subset of generative summarization models in a set of generative summarization models, the method further comprising: generating the set of generative summarization models, comprising: obtaining a plurality of training data sets; (Kryscinski [0045]: “training data generation module 130 receives an unannotated collection or set S of source documents” Kryscinski teaches obtaining documents.) obtaining data specifying a plurality of un-trained generative summarization models, each of the un-trained generative summarization models having a different number of parameters; (Chowdhery [Page 4, Paragraph 3]: “we present results at three different parameter scales: 8B, 62B, and 540B. Typically, scaling from 62B to 540B results in similar performance as scaling from 8B to 62B, which is consistent with the “power law” rule of thumb often observed in neural network scaling” Chowdhery teaches using different models of varying number of parameters: 8, 62, and 540 billion.) and for each un-trained generative summarization model, training a respective instance of the un-trained generative summarization model on each of the plurality of training data sets to generate a corresponding trained generative summarization model. (Kryscinski [0060]: “the manually annotated dataset utilizes summaries output by state-of-the-art summarization models, including extractive, abstractive, and hybrid approaches” Kryscinski teaches that summarization models process the documents to output summaries of the documents.) The reasons to combine are substantially similar to those of claim 8. Claim 19 is rejected on the same grounds under 35 U.S.C. 103 as claim 13 as they are substantially similar. Mutatis mutandis. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. [Pagnoni et al. (US 20240202461 A1)]. This prior art describes deconstructing masked words/sentences from input document by prompting a language model with questions. This is relevant because the masked words/sentences are described as a summary of the document itself and thus the prompt questions reveals some level of consistency between the masked words/sentences and the input document. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYU HYUNG HAN whose telephone number is (703) 756-5529. The examiner can normally be reached on MF 9-5. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached on (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Kyu Hyung Han/ Examiner Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123
Read full office action

Prosecution Timeline

May 17, 2024
Application Filed
Jul 30, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12651157
METHODS AND SYSTEMS FOR GENERATING THE GRADIENTS OF A LOSS FUNCTION WITH RESPECT TO THE WEIGHTS OF A CONVOLUTION LAYER
4y 2m to grant Granted Jun 09, 2026
Patent 12585928
HARDWARE ARCHITECTURE FOR INTRODUCING ACTIVATION SPARSITY IN NEURAL NETWORK
4y 10m to grant Granted Mar 24, 2026
Patent 12387101
SYSTEMS AND METHODS FOR PRUNING BINARY NEURAL NETWORKS GUIDED BY WEIGHT FLIPPING FREQUENCY
4y 3m to grant Granted Aug 12, 2025
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
50%
Grant Probability
79%
With Interview (+29.2%)
4y 1m (~1y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 14 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month