Prosecution Insights
Last updated: October 02, 2026
Application No. 18/348,759

COMPUTER-READABLE RECORDING MEDIUM STORING MACHINE LEARNING PROGRAM, MACHINE LEARNING METHOD, AND INFORMATION PROCESSING APPARATUS FOR PERFORMING EFFICIENT TRAINING ON LANGUAGE MODEL BY INCORPORATING NON-FUNCTIONAL PERFORMANCE IN LOSS FUNCTION

Non-Final OA §103
Filed
Jul 07, 2023
Priority
Jul 14, 2022 — JP 2022-113423
Examiner
WITHEY, THEODORE JOHN
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Fujitsu Limited
OA Round
4 (Non-Final)
43%
Grant Probability
Moderate
4-5
OA Rounds
0m
Est. Remaining
83%
With Interview

Examiner Intelligence

Grants 43% of resolved cases
43%
Career Allowance Rate
13 granted / 30 resolved
-18.7% vs TC avg
Strong +39% interview lift
Without
With
+39.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
25 currently pending
Career history
68
Total Applications
across all art units

Statute-Specific Performance

§101
17.5%
-22.5% vs TC avg
§103
57.1%
+17.1% vs TC avg
§102
14.7%
-25.3% vs TC avg
§112
9.4%
-30.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 30 resolved cases

Office Action

§103
DETAILED ACTION This office action is in response to Applicant’s Amendment/Request for Reconsideration, received on 06/03/2026. Claims 1, 4, 5, 7 have been amended. Claims 1, 3-7 are pending and have been considered. The examiner would like to note that a certified copy of the foreign priority document has been received. The “Priority” section of this action has been updated to reflect these changes. Applicant’s claim to foreign priority has been perfected. The examiner would like to further note that, due to Applicant’s arguments against the examiner’s incomplete explanatory statement in the rejection of the independent claims in the non-final action mailed 03/03/2026, this action will be a second non-final. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 06/03/2026, see pgs. 7-11, have been fully considered but they are not persuasive. Applicant’s representative asserts, “As explained below, at least these features of claim 1 distinguishes over each of Lester, Rafferty, and Peleg, and thus over their combination. Firstly, the Office Action asserts on pages 5-6 that: ‘Regarding claim 1, Lester discloses: [...] measuring, for each of a plurality of pieces of data included in the corpus ([0079] the training data 162 can include a plurality of training examples and a plurality of respective labels, [0085] machine-learned model(s) can process the text or natural language data to generate an output), a non-functional performance that represents a performance for a requirement that excludes a function of each of the plurality of pieces of data ([0031] the prompt gradient can be determined by evaluating a loss function that is evaluated based on a difference between the training output and the one or more training labels. The loss function can include a perceptual loss or another loss function. In some implementations, the labels can include ground truth outputs for the respective training examples [Determining perceptual loss, i.e. accuracy, between training examples and ground truth output indicates a non-functional performance measure, see [0057] of the instant application, wherein prediction accuracy excludes functionality, i.e. masking or other prediction method. The examiner asserts that this is ]).’ The phrase ‘[The examiner asserts that this is ]’ at the end of this statement in the Office Action is incomplete and unclear. Applicant cannot fully grasp the Examiner's intended meaning or assertion from this incomplete statement, which hinders a complete and precise response to this specific point. The Office Action acknowledges on page 6 that: ‘Lester does not disclose: the non-functional performance excluding objective accuracy for prediction task.’ However, the Examiner asserts that: ‘Rafferty discloses: the non-functional performance excluding objective accuracy for prediction task ([0061] A training criterion may include a number of epochs, a training time, [A training time tracks to a non-functional performance, see [0022] of instant application 'program execution speed' Wherein the training is in the context of prediction, see abstract]). Lester and Rafferty are considered analogous art within user data prediction. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lester to incorporate the teachings of Rafferty, because of the novel way to develop a filtering management system configured to determine an optimal time to deliver incoming notifications based on the intensity of the current user's task, improving the accuracy and applicability of the filtering management system (Rafferty, [0007]).’ ‘Rafferty further discloses: as a weight term, a parameter determined according to a measurement result of the non-functional performance that is the parameter that indicates a ratio of reflecting the non-functional performance in the language model ([0061] In some embodiments, training of a model may terminate when a training criterion is satisfied. A training criterion may include... a performance metric (e.g., an estimate of accuracy in reproducing test data, a loss function metric), or the like. Model trainer 436 may be configured to adjust model parameters during training. Model parameters may include weights, coefficients, offsets, or the like, [In view of the previously disclosed training time criterion of Rafferty which is a measure of non-functional performance. Considering the plurality of training criterion disclosed in Rafferty, in view of the loss function of Lester which can include a variety of losses ([0035]), indicating each disclosed training criterion of Rafferty could be implemented using the multi-level loss function of Lester. Further, representing non-functional performance as a ratio reflecting the non- functional performance, i.e. time, is a form of loss function, i.e. 50/35, 50 seconds taken over the target criterion 35 seconds, giving a loss of 15 seconds. The ratio will not be satisfied until the training criterion is. Further, representing time as something which "indicates a ratio "is a generic configuration which does not add patentable weight to the claims and does not necessarily have to be a ratio itself]).’ Assuming arguendo that Raffety discloses 'A training criterion may include a number of epochs, a training time,’ this feature of Raffety is totally different from ‘the non-functional performance excluding objective accuracy for prediction task’ in the context of ‘measuring, for each of a plurality of pieces of data included in the corpus, a non-functional performance that represents a performance for a requirement that excludes a function of the each of the plurality of pieces of data, the non-functional performance excluding objective accuracy for prediction task,’ as claimed. Further, this feature of Raffety fails to teach or suggest ‘as a weight term, a parameter determined according to a measurement result of the non-functional performance that is the parameter that indicates a ratio of reflecting the non- functional performance in the language model’ in the context of ‘the loss function includes: as a loss term, a difference between the correct answer data and a prediction result, and as a weight term, a parameter determined according to a measurement result of the non-functional performance that is the parameter that indicates a ratio of reflecting the non-functional performance in the language model’, as claimed.” With regard to Applicant’s comments about the examiner’s incomplete statement, the examiner apologizes for the error, and the instant action is not made final in order to give Applicant a fair opportunity to respond to the complete statement, but asserts that this statement was merely conclusory and does not change the functionality of the rejection of the action, i.e. “The examiner asserts that this is a non-functional performance as currently claimed”. This statement will be amended below. With regard to Applicant’s arguments against the training criterion of Rafferty being “totally different from ‘the non-functional performance excluding objective accuracy for prediction task’ in the context of ‘measuring…a non-functional performance that represents a performance for a requirement that excludes a function of each of the plurality of pieces of data’”, the examiner respectfully disagrees. The instant application defines Fig. 9 to be a depiction of explaining how non-functional performance is measured. This appears to be prediction accuracy of generated code. Further, [0022] defines the non-functional performance to be “…a program execution speed, accuracy of a machine learning model generated by the program, or the like.” Applicant has amended the claims to have the non-functional performance excluding accuracy. This indicates the non-functional performance to be “a program execution speed…or the like”. The examiner would now like to refer to Rafferty. [0061] discloses terminating training of a model when a training criterion is satisfied, wherein that criterion can be a training time. The examiner asserts a training time to be a non-functional performance metric under the BRI of the “non-functional performance” as claimed by Applicant. Checking to see if a training time has surpassed a threshold necessarily indicates monitoring said time value to know when it has surpassed the threshold, ending training. This is a measurement of time as compared to the threshold, not a measurement of performance occurring during said training time. Applicant is merely asserting that this feature is different from the claimed non-functional performance. Further, with regard to the “as a weight term…”, the examiner respectfully asserts that the Applicant is merely asserting novelty without explicitly stating where said novelty lies in the claims as currently constructed. Applicant's arguments fail to comply with 37 CFR 1.111(b) because they amount to a general allegation that the claims define a patentable invention without specifically pointing out how the language of the claims patentably distinguishes them from the references. Applicant's arguments do not comply with 37 CFR 1.111(c) because they do not clearly point out the patentable novelty which he or she thinks the claims present in view of the state of the art disclosed by the references cited or the objections made. Further, they do not show how the amendments avoid such references or objections. Applicant’s representative continues, “Applicant respectfully disagrees that Rafferty [0061] discloses ‘the non- functional performance excluding objective accuracy for prediction task’ in the context of ‘measuring, for each of a plurality of pieces of data included in the corpus, a non-functional performance that represents a performance for a requirement that excludes a function of the each of the plurality of pieces of data, the non-functional performance excluding objective accuracy for prediction task,’ as claimed. Rafferty [0061] states that: [0061] In some embodiments, training of a model may terminate when a training criterion is satisfied. A training criterion may include a number of epochs, a training time, a performance metric (e.g., an estimate of accuracy in reproducing test data, a loss function metric), or the like. Model trainer 436 may be configured to adjust model parameters during training. Model parameters may include weights, coefficients, offsets, or the like. Training may be supervised or unsupervised. This disclosure refers to conditions for terminating model training, and ‘a training time’ is listed merely as one such criterion. That is, Rafferty's ‘training time’ in the context of a ‘training criterion’ is a temporal threshold that defines the upper limit for the duration of a training process, not ‘the non-functional performance excluding objective accuracy for prediction task’ as a measurement result of ‘measuring, for each of a plurality of pieces of data included in the corpus, a non-functional performance...’ step. As such, Rafferty fails to teach or suggest ‘the non-functional performance excluding objective accuracy for prediction task’ in the context of ‘measuring, for each of a plurality of pieces of data included in the corpus, a non- functional performance that represents a performance for a requirement that excludes a function of the each of the plurality of pieces of data, the non-functional performance excluding objective accuracy for prediction task,’ as claimed.” In response, the examiner would like to refer to their previous analysis of the broadest reasonable interpretation of the non-functional performance as currently claimed in view of the training times of Rafferty. As previously explained, in order to know when the training time criterion is satisfied, this time must be monitored throughout multiple epochs of training (Lester discloses a plurality of training text datasets, [0006], Peleg discloses 20 training epochs, [0186]), indicating the training time of each epoch to be a non-functional performance metric of the model being trained, which is accumulated after each epoch to be compared to the threshold. The examiner respectfully asserts that this interpretation is under the broadest reasonable interpretation of a non-functional performance. Applicant’s representative continues, “ Moreover, the Office Action asserts on pages 8-9 that: ‘Rafferty further discloses: as a weight term, a parameter determined according to a measurement result of the non-functional performance that is the parameter that indicates a ratio of reflecting the non-functional performance in the language model ([0061] In some embodiments, training of a model may terminate when a training criterion is satisfied. A training criterion may include... a performance metric (e.g., an estimate of accuracy in reproducing test data, a loss function metric), or the like. Model trainer 436 may be configured to adjust model parameters during training. Model parameters may include weights, coefficients, offsets, or the like, [In view of the previously disclosed training time criterion of Rafferty which is a measure of non-functional performance. Considering the plurality of training criterion disclosed in Rafferty, in view of the loss function of Lester which can include a variety of losses ([0035]), indicating each disclosed training criterion of Rafferty could be implemented using the multi-level loss function of Lester. Further, representing non-functional performance as a ratio reflecting the non- functional performance, i.e. time, is a form of loss function, i.e. 50/35, 50 seconds taken over the target criterion 35 seconds, giving a loss of 15 seconds. The ratio will not be satisfied until the training criterion is. Further, representing time as something which ‘indicates a ratio’ is a generic configuration which does not add patentable weight to the claims and does not necessarily have to be a ratio itself])." Applicant respectfully disagrees with this assertion. As established above, Rafferty's ‘training time,’ in the context of a ‘training criterion,’ is a temporal threshold for the training process itself, rather than a measurement result of the non-functional performance measured for each of a plurality of pieces of data included in the corpus. Consequently, this feature of Rafferty fundamentally fails to teach or suggest the claimed limitation of a loss function that ‘includes: [...] as a weight term, a parameter determined according to a measurement result of the non-functional performance that is the parameter that indicates a ratio of reflecting the non- functional performance in the language model.’ In addition, Rafferty's mention of ‘adjusting model parameters during training’ [0061] is a general statement and does not disclose the specific mechanism of using a weight term based on non- functional performance measurement for this adjustment, nor does Lester[0035] provide such a specific teaching. Among other things, a prima facie case of obviousness must establish that the asserted combination of references teaches or suggests each and every element of the claimed invention. In view of the distinction of claim 1 noted above, at least one claimed element is not present in the asserted combination of references. Hence, the Office Action fails to establish a prima facie case of obviousness vis-a-vis claim 1.” In response, the examiner agrees with Applicant’s assertions that the training time of Rafferty is a “temporal threshold for the training process itself”, but it is unclear to the examiner why this does not track to the non-functional performance metric as currently claimed. A time taken to perform an epoch of training is a measurement result of the non-functional performance, i.e. time, to be combined with other epochs to generate an overall training time for determining when to end training. Each epoch will be represented with its own time. As the examiner asserts that Rafferty teaches the non-functional performance as currently claimed, Applicant’s subsequent arguments with respect to the “weight term” claim language are unpersuasive as they merely allege patentability. Applicant's arguments fail to comply with 37 CFR 1.111(b) because they amount to a general allegation that the claims define a patentable invention without specifically pointing out how the language of the claims patentably distinguishes them from the references. Applicant's arguments do not comply with 37 CFR 1.111(c) because they do not clearly point out the patentable novelty which he or she thinks the claims present in view of the state of the art disclosed by the references cited or the objections made. Further, they do not show how the amendments avoid such references or objections. With regard to Applicant’s arguments against establishing a prima facie case of obviousness, the examiner respectfully asserts that the requirements for establishing obviousness as disclosed in MPEP 2142 have been satisfied in view of the above arguments maintaining that the combination of Lester, Rafferty, and Peleg disclose every element of the claimed invention. Applicant's arguments filed 06/03/2026, see pg. 11, with respect to dependent claims 3-5, 6-7 have been fully considered but they are not persuasive. In view of the examiner arguments for maintaining the previously cited combination of art, Applicant’s arguments for withdrawal of said rejections is not persuasive. Priority Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed for the parent Application No. JP2022-113423, filed on 07/14/2022. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lester et al. (US-20230325725-A1), hereinafter Lester, in view of Rafferty et al. (US-20210297376-A1), hereinafter Rafferty, further in view of Peleg et al. (US-20220198136-A1), hereinafter Peleg. Regarding claim 1, Lester discloses: a non-transitory computer-readable recording medium ([0066] one or more non-transitory computer-readable storage mediums) storing a machine learning program ([0095] Each application contains its own machine learning library and machine-learned model(s)) of performing training on a language model that is a machine learning model having at least a plurality of parameters to be trained on a machine learning processing using, as a training data set, a corpus that is language resources ([0028] a large pre-trained language model, [0079] the model trainer 160 can train the pre-trained machine-learned models, [A large language model is machine learning having a plurality, i.e. large amount, of parameters, wherein a language model is indicating to be trained on language]), the machine learning program comprising instructions which, when executed by a computer ([0072] instructions 138 which are executed by the processor 132 to cause the server computing system 130 to perform operations), cause the computer to execute processing comprising: measuring, for each of a plurality of pieces of data included in the corpus ([0079] the training data 162 can include a plurality of training examples and a plurality of respective labels, [0085] machine-learned model(s) can process the text or natural language data to generate an output), a non-functional performance that represents a performance for a requirement that excludes a function of the each of the plurality of pieces of data ([0031] the prompt gradient can be determined by evaluating a loss function that is evaluated based on a difference between the training output and the one or more training labels. The loss function can include a perceptual loss or another loss function. In some implementations, the labels can include ground truth outputs for the respective training examples [Determining perceptual loss, i.e. accuracy, between training examples and ground truth output indicates a non-functional performance measure, see [0057] of the instant application, wherein prediction accuracy excludes functionality, i.e. masking or other prediction method. The examiner asserts that this is a non-functional performance measure as currently claimed]). Lester does not disclose: the non-functional performance excluding objective accuracy for prediction task. Rafferty discloses: the non-functional performance excluding objective accuracy for prediction task ([0061] A training criterion may include a number of epochs, a training time, [A training time tracks to a non-functional performance, see [0022] of instant application “program execution speed”. Wherein the training is in the context of prediction, see abstract]). Lester and Rafferty are considered analogous art within user data prediction. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lester to incorporate the teachings of Rafferty, because of the novel way to develop a filtering management system configured to determine an optimal time to deliver incoming notifications based on the intensity of the current user’s task, improving the accuracy and applicability of the filtering management system (Rafferty, [0007]). Lester in view of Rafferty does not disclose: performing, based on using divided data obtained by dividing each of the plurality of pieces of data into a first portion of the data and a second portion that is correct answer data as training data. Peleg discloses: performing, based on using divided data obtained by dividing each of the plurality of pieces of data into a first portion of the data and a second portion that is correct answer data as training data ([0129] In order to synthesize text from at least the first and second text passages, the writing assistant may change the order of content in the text passages, merge sentences, split sentences… [0158] Such capabilities may be provided by training a model to predict text within a document from a large corpus conditioned upon the preceding text [In view of the training examples and ground truth outputs for text prediction of Lester, indicating that the text splitting of Peleg could be used to determine accuracy of predictions using the second portion of split sentences based on preceding, i.e. first portion, as correct answer data in view of the loss function comparing ground truth to expected output for predictions of Lester which indicates a correct answer data to compare the prediction to for determining accuracy]). Lester are considered analogous art within textual prediction analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lester in view of Rafferty to incorporate the teachings of Peleg, because of the novel way to generate text predictions based on context and/or word sense in addition to just words themselves, improving the rate of meaningful text prediction generation (Peleg, [0002]). Lester further discloses: the machine learning processing on the language model to predict the second portion of the data in response to an input of the first portion of the data ([0030] the pre-trained machine-learned model can include a model adapted to generate a text prediction output for text that follows an input text (e.g., the input text can include “the sky is” and the output can be “blue”). Alternatively and/or additionally, the pre-trained machine-learned model may have been trained with text masking (e.g., the input text can include “The man old” and the output can be “is”) [In view of the text splitting of Peleg, further in view of the ground truth outputs and training examples of Lester, indicates a system that is predicting ends of sentences, i.e. second portions of data, based on opening statements, i.e. first portions of data, using the text splitting of Peleg, wherein a ground truth output of Lester would be “the sky is blue” or “blue” for the split segment “the sky is” as would be determined in Peleg. Further, in view of the loss function of Lester ([0054]), there is indicated that there is an original undivided sentence so loss between target and prediction can be determined, further indicating division at some point to make the prediction]), wherein the machine learning processing includes updating the plurality of parameters included in the language model based on using a loss function ([0077] a loss can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function)), the loss function includes: as a loss term, a difference between the correct answer data and a prediction result ([0054] The model's prediction can be compared to the target to calculate a loss, and the error can be back-propagated to calculate gradients, however the system may only apply these gradient updates to our new learnable vectors [Wherein a target tracks to correct answer data in view of the ground truth outputs for training examples disclosed in [0031], [0117]]). Rafferty further discloses: as a weight term, a parameter determined according to a measurement result of the non-functional performance that is the parameter that indicates a ratio of reflecting the non-functional performance in the language model ([0061] In some embodiments, training of a model may terminate when a training criterion is satisfied. A training criterion may include… a performance metric (e.g., an estimate of accuracy in reproducing test data, a loss function metric), or the like. Model trainer 436 may be configured to adjust model parameters during training. Model parameters may include weights, coefficients, offsets, or the like, [In view of the previously disclosed training time criterion of Rafferty which is a measure of non-functional performance. Considering the plurality of training criterion disclosed in Rafferty, in view of the loss function of Lester which can include a variety of losses ([0035]), indicating each disclosed training criterion of Rafferty could be implemented using the multi-level loss function of Lester. Further, representing non-functional performance as a ratio reflecting the non-functional performance, i.e. time, is a form of loss function, i.e. 50/35, 50 seconds taken over the target criterion 35 seconds, giving a loss of 15 seconds. The ratio will not be satisfied until the training criterion is. Further, representing time as something which “indicates a ratio” is a generic configuration which does not add patentable weight to the claims and does not necessarily have to be a ratio itself]). Regarding claim 3, Lester in view of Rafferty, further in view of Peleg discloses: the non-transitory computer-readable recording medium of claim 1. Lester further discloses: wherein as the loss term of the loss function for the machine learning processing, the loss term based on an appearance probability of a superficial character of each of the plurality of pieces of data is used ([0035] the prompt training can include training the model conditioned by the prompt to output the most probable label. Training can involve a perceptual loss and/or a variety of other losses… [0134] Instead of modeling classification as the probability of an output class given some input, p(y|X), where X is a series of tokens and y is a single class label, the systems and methods can model the function as conditional generation, where Y is a sequence of tokens that represent a class label [Training based on probabilities of predicted labels, wherein training also considers loss (consider the loss function between training output and training example labels [0117]), indicates a loss term to be minimized based on probability of appearance, i.e. superficially, of a label and the associated loss based on accuracy/probability the selected label, comprising word(s), e.g. comprised of superficial characters, in a prediction, in view of the plurality of pieces of data of Lester as previously disclosed. Further, determining a most probable label indicates analysis of words and/or characters (superficially) of preceding text to know what is most probable to come next]). Regarding claim 4, Lester in view of Rafferty, further in view of Peleg discloses: the non-transitory computer-readable recording medium according to claim 1. Rafferty further discloses: wherein the plurality of pieces of data in the corpus includes a plurality of programs ([Fig. 4, Programs 434 within Memory 430], [Plurality of programs emphasized]). Lester further discloses: wherein the measuring includes measuring, for each of the plurality of programs ([0037] The prompt database can include a plurality of prompts associated with a plurality of different tasks [Differing prompts and associated tasks indicates the programs are different, i.e. the outputs will be vastly different dependent on the prompt/task]), the non-functional performance of an output obtained by executing the each of the plurality of programs ([In view of the plurality of programs of Rafferty which could each correspond to its own prompt/task of Lester as [0081] of Lester discloses a model trainer 160 which contains program files]), the non-functional performance excluding a function that defines an operation of the each of the plurality of programs ([0031] the prompt gradient can be determined by evaluating a loss function that is evaluated based on a difference between the training output and the one or more training labels. The loss function can include a perceptual loss or another loss function. In some implementations, the labels can include ground truth outputs for the respective training examples [Determining perceptual loss, i.e. accuracy, between training examples and ground truth output indicates a non-functional performance measure, see [0057] of the instant application, wherein prediction accuracy excludes functionality, i.e. masking or other prediction method]), and, the executing the machine learning processing includes executing, through machine learning that uses divided data obtained by dividing each of the plurality of programs into a head portion and a subsequent portion that is correct answer data as training data ([Fig. 1A, Training Computing System 150, Data 156], [0030] In some implementations, the pre-trained machine-learned model can include a model adapted to generate a text prediction output for text that follows an input text (e.g., the input text can include “the sky is” and the output can be “blue”) … [0054] The model's prediction can be compared to the target to calculate a loss [Wherein “the sky is” represents a head portion and “blue” represents a subsequent portion. In view of the target/prediction comparison to calculate loss, there is an indication that “the sky is blue” is an original training data divided with “blue” removed as a correct answer target so a loss between the prediction and target can be accurately determined. Further, consider the previously disclosed text splitting of Peleg in view of the target/predictions of Lester]), machine learning processing of training the language model that predicts the subsequent portion of the program according to an input of the head portion of the program ([0030] the input text can include “The man old” and the output can be “is” [In this example, the input head portion represents “The man old” with a predicted subsequent portion “is”, wherein Lester further discloses predicting additional text, [0112]]). Regarding claim 5, Lester in view of Rafferty, further in view of Peleg discloses: the non-transitory computer-readable recording medium according to claim 1. Peleg further discloses: wherein the plurality of pieces of data in the corpus includes a plurality of pieces of document data ([0155] In some cases, such documents used for training may include just a few sentences or paragraphs. In other cases, however, such documents may be thousands or hundreds of thousands of pages long and may offer many examples of word usages, context dependencies, etc. When constructing a training set using a training document, portions of the document may be labeled to obtain two parts (e.g., a prefix and a suffix). In some cases, such splits may be introduced at the end of a sentence within the training document). Lester further discloses: wherein the measuring includes measuring, for each of the plurality of pieces of document data ([0037] The prompt database can include a plurality of prompts associated with a plurality of different tasks [Differing prompts and associated tasks indicates the documents/text are different, i.e. the outputs will be vastly different dependent on the prompt/task]), the non- functional performance that indicates evaluation for an indirect function from a direct function of each of the plurality of pieces of document data ([0031] the prompt gradient can be determined by evaluating a loss function that is evaluated based on a difference between the training output and the one or more training labels. The loss function can include a perceptual loss or another loss function. In some implementations, the labels can include ground truth outputs for the respective training examples [Determining perceptual loss, i.e. accuracy, between training examples and ground truth output indicates a non-functional performance measure, see [0057] of the instant application, wherein prediction accuracy is an indirect function based on the direct function of sentence/word masking/prediction/etc.]), that excludes a function that defines a direct usage when each of the plurality of pieces of document data is used ([0217] The prediction can then be used in order to evaluate a loss function (e.g., the loss function may be evaluated by comparing the prediction and a respective label for the example [Determining loss based solely on prediction accuracy indicates the direct usage, i.e. word/sentence masking/predicting method, is excluded from the measuring]); and, the executing the machine learning processing includes executing, through machine learning that uses divided data obtained by dividing each of the plurality of pieces of document data into the first portion and the second portion that is correct answer data as training data ([Fig. 1A, Training Computing System 150, Data 156], [0030] In some implementations, the pre-trained machine-learned model can include a model adapted to generate a text prediction output for text that follows an input text (e.g., the input text can include “the sky is” and the output can be “blue”) … [0054] The model's prediction can be compared to the target to calculate a loss [Wherein “the sky is” represents a first portion and “blue” represents a second portion. In view of the target/prediction comparison to calculate loss, there is an indication that “the sky is blue” is an original training data divided with “blue” removed as a correct target answer so a loss between the prediction and target can be accurately determined. Further, consider the previously disclosed text splitting of Peleg in view of the target/predictions of Lester]), machine learning processing of training the language model that predicts the second portion according to an input of the first portion of the document data ([0030] the input text can include “The man old” and the output can be “is” [In this example, the input first portion represents “The man old” with a predicted second portion “is”, wherein Lester further discloses predicting additional text, [0112]]). Regarding claim 6, Lester discloses: a computer-implemented machine learning method ([0095] Each application contains its own machine learning library and machine-learned model(s), [Machine learning indicates a computer implementation as the machine]) of performing training on a language model that is a machine learning model having at least a plurality of parameters trained to be on a machine learning processing using ([0028] a large pre-trained language models), as a training data set, a corpus that is language resources, ([0079] the model trainer 160 can train the pre-trained machine-learned models, [A large language model is machine learning having a plurality, i.e. large amount, of parameters, wherein a language model is indicating to be trained on language]), the machine learning method comprising: measuring, for each of a plurality of pieces of data included in the corpus ([0079] the training data 162 can include a plurality of training examples and a plurality of respective labels, [0085] machine-learned model(s) can process the text or natural language data to generate an output), a non-functional performance that represents a performance for a requirement that excludes a function of the each of the plurality of pieces of data ([0031] the prompt gradient can be determined by evaluating a loss function that is evaluated based on a difference between the training output and the one or more training labels. The loss function can include a perceptual loss or another loss function. In some implementations, the labels can include ground truth outputs for the respective training examples [Determining perceptual loss, i.e. accuracy, between training examples and ground truth output indicates a non-functional performance measure, see [0057] of the instant application, wherein prediction accuracy excludes functionality, i.e. masking or other prediction method. The examiner asserts that this is a non-functional performance measure as currently claimed]). Lester does not disclose: the non-functional performance excluding objective accuracy for prediction task. Rafferty discloses: the non-functional performance excluding objective accuracy for prediction task ([0061] A training criterion may include a number of epochs, a training time, [A training time tracks to a non-functional performance, see [0022] of instant application “program execution speed”. Wherein the training is in the context of prediction, see abstract]). Lester and Rafferty are considered analogous art within user data prediction. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lester to incorporate the teachings of Rafferty, because of the novel way to develop a filtering management system configure to determine an optimal time to deliver incoming notifications based on the intensity of the current user’s task, improving the accuracy and applicability of the filtering management system (Rafferty, [0007]). Lester in view of Rafferty does not disclose: performing, based on using divided data obtained by dividing each of the plurality of pieces of data into a first portion of the data and a second portion that is correct answer data as training data. Peleg discloses: performing, based on using divided data obtained by dividing each of the plurality of pieces of data into a first portion of the data and a second portion that is correct answer data as training data ([0129] In order to synthesize text from at least the first and second text passages, the writing assistant may change the order of content in the text passages, merge sentences, split sentences… [0158] Such capabilities may be provided by training a model to predict text within a document from a large corpus conditioned upon the preceding text [In view of the training examples and ground truth outputs for text prediction of Lester, indicating that the text splitting of Peleg could be used to determine accuracy of predictions using the second portion of split sentences based on preceding, i.e. first portion, as correct answer data in view of the loss function comparing ground truth to expected output for predictions of Lester which indicates a correct answer data to compare the prediction to for determining accuracy]). Lester are considered analogous art within textual prediction analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lester in view of Rafferty to incorporate the teachings of Peleg, because of the novel way to generate text predictions based on context and/or word sense in addition to just words themselves, improving the rate of meaningful text prediction generation (Peleg, [0002]). Lester further discloses: the machine learning processing on the language model to predict the second portion of the data in response to an input of the first portion of the data ([0030] the pre-trained machine-learned model can include a model adapted to generate a text prediction output for text that follows an input text (e.g., the input text can include “the sky is” and the output can be “blue”). Alternatively and/or additionally, the pre-trained machine-learned model may have been trained with text masking (e.g., the input text can include “The man old” and the output can be “is”) [In view of the text splitting of Peleg, further in view of the ground truth outputs and training examples of Lester, indicates a system that is predicting ends of sentences, i.e. second portions of data, based on opening statements, i.e. first portions of data, using the text splitting of Peleg, wherein a ground truth output of Lester would be “the sky is blue” or “blue” for the split segment “the sky is” as would be determined in Peleg. Further, in view of the loss function of Lester ([0054]), there is indicated that there is an original undivided sentence so loss between target and prediction can be determined, further indicating division at some point to make the prediction]), wherein the machine learning processing includes updating the plurality of parameters included in the language model based on using a loss function ([0077] a loss can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function)), the loss function includes: as a loss term, a difference between the correct answer data and a prediction result ([0054] The model's prediction can be compared to the target to calculate a loss, and the error can be back-propagated to calculate gradients, however the system may only apply these gradient updates to our new learnable vectors [Wherein a target tracks to correct answer data in view of the ground truth outputs for training examples disclosed in [0031], [0117]]). Rafferty further discloses: as a weight term, a parameter determined according to a measurement result of the non-functional performance that is the parameter that indicates a ratio of reflecting the non-functional performance in the language model ([0061] In some embodiments, training of a model may terminate when a training criterion is satisfied. A training criterion may include… a performance metric (e.g., an estimate of accuracy in reproducing test data, a loss function metric), or the like. Model trainer 436 may be configured to adjust model parameters during training. Model parameters may include weights, coefficients, offsets, or the like, [In view of the previously disclosed training time criterion of Rafferty which is a measure of non-functional performance. Considering the plurality of training criterion disclosed in Rafferty, in view of the loss function of Lester which can include a variety of losses ([0035]), indicating each disclosed training criterion of Rafferty could be implemented using the multi-level loss function of Lester. Further, representing non-functional performance as a ratio reflecting the non-functional performance, i.e. time, is a form of loss function, i.e. 50/35, 50 seconds taken over the target criterion 35 seconds, giving a loss of 15 seconds. The ratio will not be satisfied until the training criterion is. Further, representing time as something which “indicates a ratio” is a generic configuration which does not add patentable weight to the claims and does not necessarily have to be a ratio itself]). Regarding claim 7, Lester discloses: an information processing apparatus (Abstract, Systems and methods for natural language processing) of performing training on a language model that is a machine learning model having at least a plurality of parameters to be trained on a machine learning processing ([0028] a large pre-trained language models, [Large language models necessarily have a plurality of parameters]) using, as a training data set, a corpus that is language resources, ([0079] the model trainer 160 can train the pre-trained machine-learned models), the information processing apparatus comprising: a memory ([0066] a memory 114); and, a processor coupled to the memory ([0066] one or more processors 112 [In view of the system of Fig. 1A, where the processors and memories are clearly coupled]), the processor being configured to: measure, for each of a plurality of pieces of data included in the corpus ([0079] the training data 162 can include a plurality of training examples and a plurality of respective labels, [0085] machine-learned model(s) can process the text or natural language data to generate an output), a non-functional performance that represents a performance for a requirement that excludes a function of each of the plurality of pieces of data ([0031] the prompt gradient can be determined by evaluating a loss function that is evaluated based on a difference between the training output and the one or more training labels. The loss function can include a perceptual loss or another loss function. In some implementations, the labels can include ground truth outputs for the respective training examples [Determining perceptual loss, i.e. accuracy, between training examples and ground truth output indicates a non-functional performance measure, see [0057] of the instant application, wherein prediction accuracy excludes functionality, i.e. masking or other prediction method. The examiner asserts that this is ]). Lester does not disclose: the non-functional performance excluding objective accuracy for prediction task. Rafferty discloses: the non-functional performance excluding objective accuracy for prediction task ([0061] A training criterion may include a number of epochs, a training time, [A training time tracks to a non-functional performance, see [0022] of instant application “program execution speed”. Wherein the training is in the context of prediction, see abstract]). Lester and Rafferty are considered analogous art within user data prediction. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lester to incorporate the teachings of Rafferty, because of the novel way to develop a filtering management system configure to determine an optimal time to deliver incoming notifications based on the intensity of the current user’s task, improving the accuracy and applicability of the filtering management system (Rafferty, [0007]). Lester in view of Rafferty does not disclose: perform, based on using divided data obtained by dividing each of the plurality of pieces of data into a first portion of the data and a second portion that is correct answer data as training data. Peleg discloses: perform, based on using divided data obtained by dividing each of the plurality of pieces of data into a first portion of the data and a second portion that is correct answer data as training data ([0129] In order to synthesize text from at least the first and second text passages, the writing assistant may change the order of content in the text passages, merge sentences, split sentences… [0158] Such capabilities may be provided by training a model to predict text within a document from a large corpus conditioned upon the preceding text [In view of the training examples and ground truth outputs for text prediction of Lester, indicating that the text splitting of Peleg could be used to determine accuracy of predictions using the second portion of split sentences based on preceding, i.e. first portion, as correct answer data in view of the loss function comparing ground truth to expected output for predictions of Lester which indicates a correct answer data to compare the prediction to for determining accuracy]). Lester are considered analogous art within textual prediction analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lester in view of Rafferty to incorporate the teachings of Peleg, because of the novel way to generate text predictions based on context and/or word sense in addition to just words themselves, improving the rate of meaningful text prediction generation (Peleg, [0002]). Lester further discloses: the machine learning processing on the language model to predict the second portion of the data in response to an input of the first portion of the data ([0030] the pre-trained machine-learned model can include a model adapted to generate a text prediction output for text that follows an input text (e.g., the input text can include “the sky is” and the output can be “blue”). Alternatively and/or additionally, the pre-trained machine-learned model may have been trained with text masking (e.g., the input text can include “The man old” and the output can be “is”) [In view of the text splitting of Peleg, further in view of the ground truth outputs and training examples of Lester, indicates a system that is predicting ends of sentences, i.e. second portions of data, based on opening statements, i.e. first portions of data, using the text splitting of Peleg, wherein a ground truth output of Lester would be “the sky is blue” or “blue” for the split segment “the sky is” as would be determined in Peleg. Further, in view of the loss function of Lester ([0054]), there is indicated that there is an original undivided sentence so loss between target and prediction can be determined, further indicating division at some point to make the prediction]), wherein the machine learning processing includes updating the plurality of parameters included in the language model based on using a loss function ([0077] a loss can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function)), the loss function includes: as a loss term, a difference between the correct answer data and a prediction result ([0054] The model's prediction can be compared to the target to calculate a loss, and the error can be back-propagated to calculate gradients, however the system may only apply these gradient updates to our new learnable vectors [Wherein a target tracks to correct answer data in view of the ground truth outputs for training examples disclosed in [0031], [0117]]). Rafferty further discloses: as a weight term, a parameter determined according to a measurement result of the non-functional performance that is the parameter that indicates a ratio of reflecting the non-functional performance in the language model ([0061] In some embodiments, training of a model may terminate when a training criterion is satisfied. A training criterion may include… a performance metric (e.g., an estimate of accuracy in reproducing test data, a loss function metric), or the like. Model trainer 436 may be configured to adjust model parameters during training. Model parameters may include weights, coefficients, offsets, or the like, [In view of the previously disclosed training time criterion of Rafferty which is a measure of non-functional performance. Considering the plurality of training criterion disclosed in Rafferty, in view of the loss function of Lester which can include a variety of losses ([0035]), indicating each disclosed training criterion of Rafferty could be implemented using the multi-level loss function of Lester. Further, representing non-functional performance as a ratio reflecting the non-functional performance, i.e. time, is a form of loss function, i.e. 50/35, 50 seconds taken over the target criterion 35 seconds, giving a loss of 15 seconds. The ratio will not be satisfied until the training criterion is. Further, representing time as something which “indicates a ratio” is a generic configuration which does not add patentable weight to the claims and does not necessarily have to be a ratio itself]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Chen et al. (US-20200285899-A1) discloses “Mechanisms are provided for training a computer implemented model. The mechanisms perform multiple instances of training of the computer implemented model, where each instance of training of the computer implemented model comprises training the computer implemented model using a different training data set to generate a different instance of a trained computer implemented model. The mechanisms generate computer implemented model results after each instance of training by executing the corresponding instance of the trained computer implemented model. The mechanisms record differences in the instances of training of the computer implemented model in association with corresponding identifiers of the instances of trained computer implemented model and corresponding computer implemented model results. The mechanisms analyze the recorded differences and the corresponding computer implemented model results, and generate an output indicating a correlation between recorded differences and corresponding computer implemented model results.” (abstract). Specifically, [0054] discloses a performance/accuracy metric of training a model. See entire document. Mueller et al. (US-20210326717-A1) discloses “Techniques for code-free automated machine learning (ML) are described. Users can train high-quality ML models and pipelines without necessarily needing to write code by providing a training dataset to a code-free machine learning service. The service may deploy an ML orchestration function and a storage location on behalf of a user. When a modification is made to the storage bucket, such as by the user providing a training dataset, the orchestration function is invoked and can automatically initiate an AutoML process using at least the training data to train multiple ML model variants. The resultant ML model(s) and associated metrics can be provided to the user, deployed behind an endpoint, and/or used to generate inferences” (abstract). Specifically, [0113] discloses model metric including execution latency. See entire document. Jain et al. (US-20240007414-A1) discloses “Methods, apparatus, systems, and articles of manufacture are disclosed to optimize resources in edge networks. An example apparatus includes agent managing circuitry to invoke an exploration agent to identify platform resource devices, select a first one of the identified platform resource devices, and generate first optimization metrics for the workload corresponding to the first one of the identified platform resource devices, the first optimization metrics corresponding to a first path. The example agent is to further select a second one of the identified platform resource devices, generate second optimization metrics for the workload corresponding to the second one of the identified platform resource devices, the second optimization metrics corresponding to a second path. The example apparatus also includes benchmark managing circuitry to embed second semantic information to the workload, the second semantic information including optimized graph information and platform structure information corresponding to the second one of the identified platform resource devices, and reconfiguration managing circuitry to select the first path or the second path during runtime based on (a) service level agreement (SLA) information and (b) utilization information corresponding to the first and second identified platform resource devices.” (abstract). [0220] discloses accuracy metrics, speed metrics, power consumption metrics of models. See entire document. Any inquiry concerning this communication or earlier communications from the examiner should be directed to THEODORE JOHN WITHEY whose telephone number is (703)756-1754. The examiner can normally be reached Monday - Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571) 272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /THEODORE WITHEY/Examiner, Art Unit 2655 /JESSE S PULLIAS/Primary Examiner, Art Unit 2655 08/24/26
Read full office action

Prosecution Timeline

Show 1 earlier event
May 22, 2025
Non-Final Rejection mailed — §103
Aug 25, 2025
Response Filed
Oct 17, 2025
Final Rejection mailed — §103
Jan 20, 2026
Request for Continued Examination
Jan 29, 2026
Response after Non-Final Action
Mar 03, 2026
Non-Final Rejection mailed — §103
Jun 03, 2026
Response Filed
Aug 27, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724985
MACHINE-LEARNING BASED IRRELEVANT SENTENCE CLASSIFIER
3y 11m to grant Granted Sep 01, 2026
Patent 12657389
TECHNOLOGIES FOR ERROR REDUCTION IN INTENT CLASSIFICATION
3y 5m to grant Granted Jun 16, 2026
Patent 12646499
METHOD, DEVICE, AND COMPUTER PROGRAM PRODUCT FOR PROCESSING INFORMATION
3y 3m to grant Granted Jun 02, 2026
Patent 12632670
Natural Language Processing for Identifying Bias in a Span of Text
3y 2m to grant Granted May 19, 2026
Patent 12591744
METHOD FOR TRAINING SEMANTIC REPRESENTATION MODEL, DEVICE AND STORAGE MEDIUM
4y 0m to grant Granted Mar 31, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

4-5
Expected OA Rounds
43%
Grant Probability
83%
With Interview (+39.3%)
2y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 30 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month