DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Notice for all US Patent Applications filed on or after March 16, 2013
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Status of the Claims
This communication is in response to communications received on 4/30/24. Claim(s) none is/are amended, claim(s) none is/are cancelled, claim(s) none is/are new, and applicant does not provide any information on where support for the amendments can be found in the instant specification. Therefore, Claims 1-20 is/are pending and have been addressed below.
Information Disclosure Statement
The information disclosure statement(s) (IDS) submitted on 8/1/25 was/were considered by the examiner.
Response to Arguments
There are no arguments.
Claims Without Prior Art Rejections
Claim(s) 9-10 do/does not have prior art rejections. The remaining rejections are 101 as noted below.
Closest prior art to the invention include
Michael et al. (US 2024/0320450 A1) in view of Majumder et al. (WO 2024/261154 A1) and Peng et al. published April 17, 2024 (reference U on the Notice of References Cited) for claim(s) 9-10 and 17.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter as noted below.
The limitation(s) below for representative claim(s) 1, 12, and 18 that, under its broadest reasonable interpretation, is directed to fine tuning large language models.
Step 1: The claim(s) as drafted, is/are a process (claim(s) 1-11 recites a series of steps) and system (claim(s) 12-20 recites a series of components).
Step 2A – Prong 1: The claimed invention is directed to an abstract idea without significantly more. The claim(s) recite(s) (emphasis added):
Claim 1: accessing a fine-tuning input in a target language for fine-tuning a pre- trained large language model (LLM);
obtaining labeled data based on the fine-tuning input in the target language as a fine-tuning output;
obtaining a fine-tuning input in a reference language corresponding to the fine-tuning input in the target language; and
fine-tuning the pre-trained LLM for a generative task based on the fine- tuning input in the target language, the fine-tuning input in the reference language, and the fine-tuning output to obtain a first fine-tuned LLM.
Claim(s) 12 and 18: same analysis as claim(s) 1.
Dependent claims 2-11, 13-17, and 19-20 recite the same or similar abstract idea(s) as independent claim(s) 1, 12, and 18 with merely a further narrowing of the abstract idea(s): .
The identified limitations of the independent and dependent claims above fall well-within the groupings of subject matter identified by the courts as being abstract concepts of:
mathematical relationships, mathematical formulas or equations, or mathematical calculations because the invention is directed to the application of mathematical processes as they are associated with fine tuning large language models.
Step 2A – Prong 2: This judicial exception is not integrated into a practical application because:
The additional elements unencompassed by the abstract idea include LLM, (claim(s) 1, 12, 18), a system comprising: a communications interface; a non-transitory computer-readable medium; and one or more processors (claim(s) 12), a non-transitory computer-readable medium, one or more processors (claim(s) 18), (claim(s) 18), generative model (claim(s) 4, 14), translation model (claim(s) 5), LLM (claim(s) 6-11), frozen model (claim(s) 8, 10, 11, 16, 17).
The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements as described above with respect to Step 2A Prong 2 fails to describe:
Improvements to the functioning of a computer, or to any other technology or technical field - see MPEP 2106.05(a)
Applying or using a judicial exception to effect a particular treatment or prophylaxis for a disease or medical condition – see Vanda Memo
Applying the judicial exception with, or by use of, a particular machine – see MPEP 2106.05(b)
Effecting a transformation or reduction of a particular article to a different state or thing - see MPEP 2106.05(c)
Applying or using the judicial exception in some other meaningful way beyond generally linking the use of the judicial exception to a particular technological environment, such that the claim as a whole is more than a drafting effort designed to monopolize the exception - see MPEP 2106.05(e) and Vanda Memo.
Thus the additional elements as described above with respect to Step 2A Prong 2 are merely invoked as a tool and/or general purpose computer to apply instructions of an abstract idea in a particular technological environment, and/or mere application of an abstract idea in a particular technological environment and merely limiting the use of an abstract idea to a particular technological field do not integrate an abstract idea into a practical application (MPEP 2106.05(f)&(h)).
Step 2B: The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus the additional elements as described above with respect to Step 2A Prong 2 are merely invoked as a tool and/or a general purpose computer to apply instructions of an abstract idea in a particular technological environment, and/or mere application of an abstract idea in a particular technological environment and merely limiting the use of an abstract idea to a particular technological field do not integrate an abstract idea into a practical application and thus similarly the combination and arrangement of the above identified additional elements when analyzed under Step 2B also fails to necessitate a conclusion that the claims amount to significantly more than the abstract idea for the same reasons as set forth above (MPEP 2106.05(f)&(h)).
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-2, 6-7, 12-13, 15, and 18-19 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Michael et al. (US 2024/0320450 A1).
Regarding claim 1, 12, and 18, Michael teaches a method comprising:
{a non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to: - claim 12}
{a system comprising: a communications interface; a non-transitory computer-readable medium; and one or more processors communicatively coupled to the communications interface and the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non- transitory computer-readable medium to: - claim 18}
accessing a fine-tuning input in a target language for fine-tuning a pre- trained large language model (LLM);
obtaining labeled data based on the fine-tuning input in the target language as a fine-tuning output;
obtaining a fine-tuning input in a reference language corresponding to the fine-tuning input in the target language [for the limitations above, see at least Fig. 1 and [0026-0027] “In an embodiment of the present disclosure, the system is implemented in an electronic device 102. Examples of the electronic device 102 may include, but are not limited to, a smartphone, a laptop, a camera device, a smartwatch, and the like.
The system 100 may include one or more processors/controllers 104, an Input/Output (I/O) interface 106, a plurality of modules 108, and a memory 110.”;
[0053] fine tuning is performed using target language thus data is accessed “The process begins with pre-training the LLM-based language translation models on large-scale multilingual corpora, allowing them to learn linguistic patterns, syntactic structures, and semantic relationships. … To make these LLM-based language translation models task-specific, fine-tuning is performed.”;
[0052] target and reference language is obtained via fine-tuning input “In an embodiment of the present disclosure, the obtaining module 206 may be configured to receive a source language text and a set of target language labels. The source language text is the training data required for the AI-based LTN. In an embodiment of the present disclosure, a set of sentences in the source language/languages and its translation to the target language comprise the training data. Further, one or more sentences in the source language is the “source language text. In an embodiment of the present disclosure, the set of target language labels are translations of the source language text in the target language.”]; and
fine-tuning the pre-trained LLM for a generative task based on the fine- tuning input in the target language, the fine-tuning input in the reference language, and the fine-tuning output to obtain a first fine-tuned LLM [see at least [0053] “In an embodiment of the present disclosure, the AI-based LTN and the AI-based IDN (i.e., Large Language Model (LLM)-based language translation models) are fine-tuned. Further, fine-tuning the LLM-based language translation models is a technique that involves adapting pre-trained models specifically designed for language translation tasks. This approach is used in the field of machine translation due to its ability to achieve state-of-the-art performance and generalize across various language pairs. The process begins with pre-training the LLM-based language translation models on large-scale multilingual corpora, allowing them to learn linguistic patterns, syntactic structures, and semantic relationships. … To make these LLM-based language translation models task-specific, fine-tuning is performed. Task-specific labeled data, comprising sentence pairs in the source and target languages, is used to further train the LLM-based language translation models. This fine-tuning process adapts the pre-trained models to specific language pairs, enabling them to capture the idiosyncrasies and nuances of translation in the target domain. … By combining pre-training knowledge with task-specific training, fine-tuning enables rapid development of high-performance translation models, advancing the capabilities of AI in multilingual communication and facilitating accurate and efficient language translation.”].
Regarding claim 2 and 13, Michael teaches the method of claim 1, wherein the reference language is English, and the target language is not English [see at least [0066] “For example, the “source language text” refers to the English sentences, which act as the input for the AI-based LTN. The “target language labels” represent the corresponding French translations, which act as the desired output or target for the AI-based LTN to learn from. Accordingly, the set of target language labels may correspond to the desired or expected translations of the input sentences. The set of target language labels may serve as a reference for the translator to understand how to convert the source language text accurately.”].
Regarding claim 6 and 19, Michael teaches the method of claim 1, wherein fine-tuning the pre-trained LLM for a generative task comprises:
using the pre-trained LLM to generate an interim output in the target language for the generative task based on the fine-tuning input in the target language and the fine-tuning input in the reference language [see at least [0067] “In an embodiment of the present disclosure, the AI-based LTN is trained using a transformer trainer class. The source language text is processed through the encoder-attention-decoder layers, such that the text in the target language is generated.”;
[0053] “To make these LLM-based language translation models task-specific, fine-tuning is performed. Task-specific labeled data, comprising sentence pairs in the source and target languages, is used to further train the LLM-based language translation models. This fine-tuning process adapts the pre-trained models to specific language pairs, enabling them to capture the idiosyncrasies and nuances of translation in the target domain.”];
minimizing a cross entropy loss associated with the interim output in the target language and the fine-tuning output to obtain one or more optimized weights for the pre-trained LLM; and
providing the first fine-tuned LLM comprising the one or more optimized weights [for the limitations above, see at least [0067-0068] “In an embodiment of the present disclosure, the AI-based LTN is trained using a transformer trainer class. The source language text is processed through the encoder-attention-decoder layers, such that the text in the target language is generated. At step 510, training arguments are created for optimizing hyperparameters. In step 510, training arguments for optimizing hyperparameters are created. Further, the parameters like epoch and learning rate are optimized so that the training loss is minimized.
Further, at step 512, the weights of the AI-based LTN are updated. Further, the trained AI-based LTN is obtained at step 514. In an embodiment of the present disclosure, the evaluation data is used for calculating the accuracy of the prediction of the intent category. In an embodiment of the present disclosure, the AI-based LTN receives a user input (user utterance) 514 and generates the user utterance in the base language, and then the AI-based IDN processes the data for intent classification 516 and action selection 518. In the intent classification, the AI-based IDN outputs the predicted intent from the user's message which is translated to the base language by the AI-based LTN. In the action selection, system 100 selects the appropriate action or response to be taken by the chatbot based on the predicted intent. At step 520, the action execution is performed.”].
Regarding claim 7 and 15, Michael teaches the method of claim 1, further comprising:
receiving an input in the target language;
obtaining an input in the reference language corresponding to the input in the target language; and
implementing the first fine-tuned LLM to generate an output in the target language for the generative task based on the input in the target language and the input in the reference language [for the limitations above, see at least [0052] target and reference language is obtained via fine-tuning input “In an embodiment of the present disclosure, the obtaining module 206 may be configured to receive a source language text and a set of target language labels. The source language text is the training data required for the AI-based LTN. In an embodiment of the present disclosure, a set of sentences in the source language/languages and its translation to the target language comprise the training data. Further, one or more sentences in the source language is the “source language text. In an embodiment of the present disclosure, the set of target language labels are translations of the source language text in the target language.”;
[0053] “In an embodiment of the present disclosure, the AI-based LTN and the AI-based IDN (i.e., Large Language Model (LLM)-based language translation models) are fine-tuned. Further, fine-tuning the LLM-based language translation models is a technique that involves adapting pre-trained models specifically designed for language translation tasks. This approach is used in the field of machine translation due to its ability to achieve state-of-the-art performance and generalize across various language pairs. The process begins with pre-training the LLM-based language translation models on large-scale multilingual corpora, allowing them to learn linguistic patterns, syntactic structures, and semantic relationships. … To make these LLM-based language translation models task-specific, fine-tuning is performed. Task-specific labeled data, comprising sentence pairs in the source and target languages, is used to further train the LLM-based language translation models. This fine-tuning process adapts the pre-trained models to specific language pairs, enabling them to capture the idiosyncrasies and nuances of translation in the target domain. … By combining pre-training knowledge with task-specific training, fine-tuning enables rapid development of high-performance translation models, advancing the capabilities of AI in multilingual communication and facilitating accurate and efficient language translation.”].
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
It has been held that a prior art reference must either be in the field of applicant’s endeavor or, if not, then be reasonably pertinent to the particular problem with which the applicant was concerned, in order to be relied upon as a basis for rejection of the claimed invention. See In re Oetiker, 977 F.2d 1443, 24 USPQ2d 1443 (Fed. Cir. 1992).
Claim(s) 3-5 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Michael et al. (US 2024/0320450 A1) in view of Majumder et al. (WO 2024/261154 A1).
Regarding claim 3, Michael teaches the method of claim 1, as well as the labeled data.
Michael doesn’t/don’t explicitly teach however Majumder discloses
(original vs citation) wherein the labeled data is provided by a human annotator [see at least [0092, 0098] “In examples where the moderation result is indicative of a content item requiring human review, the content item may be held until it has been reviewed by a moderator, for example via a suitable dashboard.”].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Michael with Majumder to include the limitation(s) above as disclosed by Majumder. Doing so would improve Michael’s (Michael) translation via an additional step [see at least Majumder [0001-0025, 0086] ].
Furthermore, all of the claimed elements were known in the prior arts of a) Michael and b) Majumder and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention.
Regarding claim 4 and 14, Michael teaches the method of claim 1, as well as
generate the labeled data based on the fine-tuning input in the target language.
Michael doesn’t/don’t explicitly teach however Majumder discloses
(original vs citation) further comprising using a generative model to generate the labeled data based on the fine-tuning input in the target language [see at least [0042] “As discussed hereinabove, the generative model 200 may generate text, audio, images, video or any other suitable content, or a combination thereof. For simplicity of explanation, in the following discussion it will be assumed that the model is a text generation model. For example the model 200 may be a large language model (LLM) such as GPT-4, provided by Open Al®.”;
[0014] “However, a wide variety of generative models may be employed in conjunction with the present techniques, including video generation models, audio generation models, multimedia generation models and so on. Example models include GPT-3, GPT-3.5 turbo, GPT-4, GPT-4o, ChatGPT, and Dall-E.”;
[0085] “In some embodiments, one or more of the Al models 533 may be based on one or more generic language models. These may be pretrained on large volumes of general text data, and thus are suitable for a wide variety of language processing tasks. The Al model is then subsequently tailored to a particular task by further training (or fine-tuning), based on a tailored training set comprising application-specific training data. Examples of such generic language models (also referred to in the art as "large language models") include BERT, GPT-3, and cohere.”;
[0086] “In one example, the plurality of Al models 533 includes a nudity detector, such as NudeNet (https://pypi.org/project/NudeNet/). The plurality of Al models 533 may include one or more machine translation models, configured to translate the content item. The translation models may be applied as a pre-processing step, with the output of the translation models forming input to the other models 533.”;
[0037] “The trained machine learning model 110 is trained to generate data items to include in a prompt 121 for submission to a generative model 200. In the main, the data items comprise text strings for inclusion in a text prompt submitted to a generative model 200. However, it is to be understood that in some examples, the data items may be images, audio, video, structured data (e.g. tabular data) or any other suitable data. In this context, generating the data items may include selecting suitable data items from a plurality of stored data items. The training of the model 110 will be discussed in more detail below with respect to Figure 2. [0038] In some examples, the data items 111 are "shots" - i.e. labelled examples of content and associated emotion. The technique may therefore be a "one-shot" or "few-shot" learning technique in which one or a small number (e.g. < 10) labelled examples are included in a prompt for a generative model 200. The shots may be selected from a store of shots, or generated from scratch by the model 110.”].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Michael with Majumder to include the limitation(s) above as disclosed by Majumder. Doing so would improve Michael’s (Michael) translation via an additional pre-processing step [see at least Majumder [0001-0025, 0086] ].
Furthermore, all of the claimed elements were known in the prior arts of a) Michael and b) Majumder and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention.
Regarding claim 5, Michael teaches the method of claim 1, as well as generate the labeled data based on the fine-tuning input in the target language.
Michael doesn’t/don’t explicitly teach however Majumder discloses
(original vs citation) further comprising using a translation model to generate the labeled data based on the fine-tuning input in the target language [see at least [0042] “As discussed hereinabove, the generative model 200 may generate text, audio, images, video or any other suitable content, or a combination thereof. For simplicity of explanation, in the following discussion it will be assumed that the model is a text generation model. For example the model 200 may be a large language model (LLM) such as GPT-4, provided by Open Al®.”;
[0014] “However, a wide variety of generative models may be employed in conjunction with the present techniques, including video generation models, audio generation models, multimedia generation models and so on. Example models include GPT-3, GPT-3.5 turbo, GPT-4, GPT-4o, ChatGPT, and Dall-E.”;
[0085] “In some embodiments, one or more of the Al models 533 may be based on one or more generic language models. These may be pretrained on large volumes of general text data, and thus are suitable for a wide variety of language processing tasks. The Al model is then subsequently tailored to a particular task by further training (or fine-tuning), based on a tailored training set comprising application-specific training data. Examples of such generic language models (also referred to in the art as "large language models") include BERT, GPT-3, and cohere.”;
[0086] “In one example, the plurality of Al models 533 includes a nudity detector, such as NudeNet (https://pypi.org/project/NudeNet/). The plurality of Al models 533 may include one or more machine translation models, configured to translate the content item. The translation models may be applied as a pre-processing step, with the output of the translation models forming input to the other models 533.”;
[0037] “The trained machine learning model 110 is trained to generate data items to include in a prompt 121 for submission to a generative model 200. In the main, the data items comprise text strings for inclusion in a text prompt submitted to a generative model 200. However, it is to be understood that in some examples, the data items may be images, audio, video, structured data (e.g. tabular data) or any other suitable data. In this context, generating the data items may include selecting suitable data items from a plurality of stored data items. The training of the model 110 will be discussed in more detail below with respect to Figure 2. [0038] In some examples, the data items 111 are "shots" - i.e. labelled examples of content and associated emotion. The technique may therefore be a "one-shot" or "few-shot" learning technique in which one or a small number (e.g. < 10) labelled examples are included in a prompt for a generative model 200. The shots may be selected from a store of shots, or generated from scratch by the model 110.”].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Michael with Majumder to include the limitation(s) above as disclosed by Majumder. Doing so would improve Michael’s (Michael) translation via an additional pre-processing step [see at least Majumder [0001-0025, 0086] ].
Furthermore, all of the claimed elements were known in the prior arts of a) Michael and b) Majumder and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention.
Claim(s) 8, 11, 16, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Michael et al. (US 2024/0320450 A1) in view of Majumder et al. (WO 2024/261154 A1) and Peng et al. published April 17, 2024 (reference U on the Notice of References Cited).
Regarding claim 8, 16, and 20, Michael teaches the method of claim 1, as well as and; fine-tuning the pre-trained LLM for the generative task based on the fine- tuning input in the reference language, the first output in the target language, the first output in the reference language, and the fine-tuning output to obtain a fine-tuned LLM (fine-tuning the pre-trained LLM for the generative task based on the fine- tuning input in the reference language, the first output in the target language, the first output in the reference language, and the fine-tuning output to obtain a first fine-tuned LLM - claim 1).
Michael teaches (original vs citation) further comprising: using a frozen model to generate a first output in the target language for the generative task based on the fine-tuning input in the target language;
using the frozen model to generate a first output in the reference language for the generative task based on the fine-tuning input in the reference language; and [for the limitations above, see at least Fig. 1 and [0026-0027] “In an embodiment of the present disclosure, the system is implemented in an electronic device 102. Examples of the electronic device 102 may include, but are not limited to, a smartphone, a laptop, a camera device, a smartwatch, and the like.
The system 100 may include one or more processors/controllers 104, an Input/Output (I/O) interface 106, a plurality of modules 108, and a memory 110.”;
[0053] fine tuning is performed using target language thus data is accessed “The process begins with pre-training the LLM-based language translation models on large-scale multilingual corpora, allowing them to learn linguistic patterns, syntactic structures, and semantic relationships. … To make these LLM-based language translation models task-specific, fine-tuning is performed.”;
[0052] target and reference language is obtained via fine-tuning input “In an embodiment of the present disclosure, the obtaining module 206 may be configured to receive a source language text and a set of target language labels. The source language text is the training data required for the AI-based LTN. In an embodiment of the present disclosure, a set of sentences in the source language/languages and its translation to the target language comprise the training data. Further, one or more sentences in the source language is the “source language text. In an embodiment of the present disclosure, the set of target language labels are translations of the source language text in the target language.”].
Michael doesn’t/don’t explicitly teach however Majumder discloses
(original vs citation) fine-tuning the pre-trained LLM for the generative task based on the fine- tuning input in the reference language, the first output in the target language, the first output in the reference language, and the fine-tuning output to obtain a second fine-tuned LLM [see at least [0042] “As discussed hereinabove, the generative model 200 may generate text, audio, images, video or any other suitable content, or a combination thereof. For simplicity of explanation, in the following discussion it will be assumed that the model is a text generation model. For example the model 200 may be a large language model (LLM) such as GPT-4, provided by Open Al®.”;
[0014] “However, a wide variety of generative models may be employed in conjunction with the present techniques, including video generation models, audio generation models, multimedia generation models and so on. Example models include GPT-3, GPT-3.5 turbo, GPT-4, GPT-4o, ChatGPT, and Dall-E.”;
[0085] “In some embodiments, one or more of the Al models 533 may be based on one or more generic language models. These may be pretrained on large volumes of general text data, and thus are suitable for a wide variety of language processing tasks. The Al model is then subsequently tailored to a particular task by further training (or fine-tuning), based on a tailored training set comprising application-specific training data. Examples of such generic language models (also referred to in the art as "large language models") include BERT, GPT-3, and cohere.”;
[0086] “In one example, the plurality of Al models 533 includes a nudity detector, such as NudeNet (https://pypi.org/project/NudeNet/). The plurality of Al models 533 may include one or more machine translation models, configured to translate the content item. The translation models may be applied as a pre-processing step, with the output of the translation models forming input to the other models 533.”;
[0037] “The trained machine learning model 110 is trained to generate data items to include in a prompt 121 for submission to a generative model 200. In the main, the data items comprise text strings for inclusion in a text prompt submitted to a generative model 200. However, it is to be understood that in some examples, the data items may be images, audio, video, structured data (e.g. tabular data) or any other suitable data. In this context, generating the data items may include selecting suitable data items from a plurality of stored data items. The training of the model 110 will be discussed in more detail below with respect to Figure 2. [0038] In some examples, the data items 111 are "shots" - i.e. labelled examples of content and associated emotion. The technique may therefore be a "one-shot" or "few-shot" learning technique in which one or a small number (e.g. < 10) labelled examples are included in a prompt for a generative model 200. The shots may be selected from a store of shots, or generated from scratch by the model 110.”].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Michael with Majumder to include the limitation(s) above as disclosed by Majumder. Doing so would improve Michael’s (Michael) translation via an additional pre-processing step [see at least Majumder [0001-0025, 0086] ].
Furthermore, all of the claimed elements were known in the prior arts of a) Michael and b) Majumder and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention.
Michael in view of Majumder doesn’t/don’t explicitly teach however Lester discloses
further comprising: using a frozen model to generate a first output in the target language for the generative task based on the fine-tuning input in the target language;
using the frozen model to generate a first output in the reference language for the generative task based on the fine-tuning input in the reference language [for the limitations above, see at least [pg 1] “An appealing alternative is to share across all downstream tasks a single frozen pre-trained language model, in which all weights are fixed. In an exciting development, GPT-3 showed convincingly that a frozen model can be conditioned to perform different tasks through “in-context” learning. With this approach, a user primes the model for a given task through prompt design, i.e., hand-crafting a text prompt with a description or examples of the task at hand. For instance, to condition a model for sentiment analysis, one could attach the prompt, “Is the following movie review positive or negative?” before the input sequence, “This movie was amazing!” ”].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Michael in view of Majumder with Lester to include the limitation(s) above as disclosed by Lester. Doing so would improve Michael in view of Majumder’s (Michael) translation via an additional pre-processing step [see at least Lester [pg 1] ].
Furthermore, all of the claimed elements were known in the prior arts of a) Michael in view of Majumder and b) Lester and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention.
Regarding claim 11, modified Michael teaches the method of claim 8, .
Modified Michael doesn’t/don’t explicitly teach however Lester discloses
wherein the frozen model comprises the pre-trained LLM or a different LLM [for the limitations above, see at least [pg 1] “An appealing alternative is to share across all downstream tasks a single frozen pre-trained language model, in which all weights are fixed. In an exciting development, GPT-3 showed convincingly that a frozen model can be conditioned to perform different tasks through “in-context” learning. With this approach, a user primes the model for a given task through prompt design, i.e., hand-crafting a text prompt with a description or examples of the task at hand. For instance, to condition a model for sentiment analysis, one could attach the prompt, “Is the following movie review positive or negative?” before the input sequence, “This movie was amazing!” ”].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify modified Michael with Lester to include the limitation(s) above as disclosed by Lester. Doing so would improve modified Michael’s (Michael) translation via an additional pre-processing step [see at least Lester [pg 1] ].
Furthermore, all of the claimed elements were known in the prior arts of a) modified Michael and b) Lester and c) one skilled in the art could have combined the elements as claimed by known methods with no change in their respective functions, and the combination would have yielded predictable results to one of ordinary skill in the art before the effective filing date of the claimed invention.
Conclusion
When responding to the office action, any new claims and/or limitations should be accompanied by a reference as to where the new claims and/or limitations are supported in the original disclosure.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Schmidt et al. – Don’t Stop Fine-Tuning: On Training Regimes for Few-Shot Cross-Lingual Transfer with Multilingual Language Models (relevant because it teaches fine tuning LLMs for reference and target languages) as noted in IDS dated 8/1/25
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES WEBB whose telephone number is (313)446-6615. The examiner can normally be reached on M-F 10-3.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jerry O’Connor can be reached on (571) 272-6787. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAMES WEBB/Examiner, Art Unit 3624