DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mingye Zhu et al. (hereinafter Zhu) (“LIRE: listwise reward enhancement for preference alignment,” 06/04/2024) in view of Benjamin Newman (hereinafter Newman) (US 20230083512 A1, 03/16/2023).
Regarding claim 1, Zhu teaches;
A method of fine-tuning (NOTE: iteratively training the pretrained LLM πθ on specific datasets De constitutes fine-tuning, see pg.4 algorithm 1 below) a neural network based model ([pg. 5] we initialize the target policy πθ as the pretrained LLM), the method comprising:
[pg. 4]
PNG
media_image1.png
387
496
media_image1.png
Greyscale
receiving, ([pg. 3] a set of queries Q = {x(i)} is given, i ∈ {1,··· ,N})
generating ([pg. 4] for each query x(i), sample M responses A(i) ∼ πθ(y|x(i)), via a pre-trained neural network based model: ([pg. 5] we initialize the target policy πθ as the pretrained LLM)
a first response (NOTE: y(i)1, see below) based on a first input sample (NOTE: x(i), see below) of the plurality of input samples (NOTE: Q, see below), and a second response (NOTE: y(i)2, see below) based on the first input sample; ([pg. 3] set of queries Q = {x(i)} … each query is associated with M responses A(i) = {y(i)1,··· ,y(i)M})
generating, via a trained reward model: ([pg. 2] well-trained reward function) a first reward score (NOTE: R(x(i),y(i)1), see below) based on the first input sample (NOTE: x(i), see below) and the first response (NOTE: y(i)1, see below), and a second reward score (NOTE: R(x(i),y(i)2), see below) based on the first input sample (NOTE: x(i), see below) and the second response (NOTE: y(i)2) ([pg. 3] each response y(i)j for query x(i) is paired with a score R(x(i),y(i)j) by some Reward Model RM)
computing a loss function (NOTE: listwise loss, see below) based on the first prompt (NOTE: x, see below), the first response (NOTE: y1, see below), the second response (NOTE: y2, see below), the first reward score (NOTE: R(x, y1), see below), and the second reward score (NOTE: R(x, y2), see below)
[pg. 3]
PNG
media_image2.png
185
542
media_image2.png
Greyscale
and updating parameters of the neural network based model (NOTE: πθ, see below) based on the loss function (NOTE: J(θ))
[pg. 4, algorithm 1]
PNG
media_image3.png
62
465
media_image3.png
Greyscale
Zhu fails to explicitly teach but Newman teaches;
NOTE: From the applicant’s specification – “[0049] The data interface 215 may comprise a communication interface.”
receiving, via a data interface, a training dataset including a plurality of input samples; ([0032] In an example, the system 100 may receive, via a communication interface, a training data set, the training dataset including a plurality of sets of similar natural language queries)
OBVIOUSNESS TO COMBINE NEWMAN:
Newman is analogous art to the present disclosure as it pertains to training a language model.
Zhu already takes input queries as inputs to its training algorithm while Newman teaches a closely related neural-network / language model training system that receives a training dataset through a communication interface.
Zhu further states that;
([0058] The data interface 715 may be … a communication interface that may receive or retrieve a previously stored question from the database.)
Therefore, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to modify Zhu to receive its training dataset, including the input queries, via a data interface as taught by Newman, in order to provide a conventional mechanism for receiving or retrieving the training data from a data source for use by the model training system.
Regarding claim 2, Zhu teaches;
wherein the first reward score is a value between 0 and 1 ([pg. 3] we apply softmax to the reward scores)
NOTE: Softmax outputs are between 0 and 1.
Regarding claim 3, Zhu teaches;
wherein the first reward score and the second reward score ([pg. 3] each query is associated with M responses … each response y(i)j for query x(i) is paired with a [reward] score … M descends to 2) sum to 1 ([pg. 3] we apply softmax to the reward scores of a single query)
NOTE: A softmax function normalizes a set of values into values between 0 and 1 whose sum is 1. Thus, when M = 2, the softmax function normalizes the first (R(x, y1)) and second (R(x, y2)) reward scores such that the first and second reward scores sum to 1.
Regarding claim 4, Zhu teaches;
wherein the computing the loss function includes summing values based on:
[pg. 3]
PNG
media_image4.png
186
542
media_image4.png
Greyscale
a first comparison ([pg. 3] we reformulate the response probability distribution against the entire response set A as: Pπθ (y1|x, A)) of a first probability that the neural network based model generates the first response ([pg. 3] probability of the sentence y1 … πθ(y1|x)) and a second probability that the neural network based model generates the second response ([pg. 3] probability of the sentence y2 … πθ(y2|x)), scaled by the first reward score (R(x, y1), see below); and a second comparison ([pg. 3] Pπθ (y2|x, A)) of the first probability that the neural network based model generates the first response ([pg. 3] πθ(y1|x))) and the probability that the neural network based model generates the second response ([pg. 3] πθ(y2|x)), scaled by the second reward score (R(x, y2), see below)
[pg. 3]
PNG
media_image5.png
91
456
media_image5.png
Greyscale
PNG
media_image6.png
90
514
media_image6.png
Greyscale
Regarding claim 5,
generating ([pg. 4] for each query x(i), sample M responses A(i) ∼ πθ(y|x(i)), via the pre-trained neural network based model ([pg. 5] initialize the target policy πθ as the pretrained LLM), a third response (NOTE: response y(i)3, see below) based on the first input sample (each query is associated with M responses A(i) = {y(i)1 ,··· ,y(i)M})
and generating, via the trained reward model ([pg. 2] well-trained reward function), a third reward score (NOTE: R(x(i),y(i)3), see below) based on the first input sample and the third response ([pg. 3] each response y(i)j for query x(i) is paired with a score R(x(i),y(i)j) by some Reward Model RM.)
wherein the computing the loss function is further based on the third response (NOTE: y3, see below) and the third reward score (NOTE: R(x, y3), see below)
[pg. 3]
PNG
media_image7.png
189
552
media_image7.png
Greyscale
Regarding claim 6, Zhu teaches;
wherein the neural network based model is initialized with a same set of parameters ([pg. 3] language model parameterized by θ) as the pre-trained neural network based model ([pg. 5] we initialize the target policy πθ as the pretrained LLM)
Regarding claim 7, Zhu teaches;
wherein the neural network based model is the pre-trained neural network based model. ([pg. 5] we initialize the target policy πθ as the pretrained LLM and start to optimize the objective J(θ))
Regarding claim 8,
Claim 8 is a system claim that is substantially similar to method claim 1, with the addition of the following limitations, which are taught by Newman:
a memory that stores the neural network based model ([0026] the memory 112 may store one or more pre-trained language models) and a plurality of processor ([0055] computing device 700 includes a processor) executable instructions ([0056] Memory 720 may be used to store software executed by computing device); a communication interface that receives a training dataset including a plurality of input samples ([0032] receive, via a communication interface, a training data set, … including a plurality of sets of similar natural language queries); and one or more hardware processors ([0055] computing device 700 includes a processor) that read and execute the plurality of processor-executable instructions from the memory to perform operations comprising ([0056] Memory 720 may be used to store software executed by computing device):
OBVIOUSNESS:
It would have been obvious to one of ordinary skill in the art, before the effective filing date, to store instructions to perform Zhu’s model training method using the computing device architecture of Newman, including the communication interface, in order to provide conventional computing resources for storing and executing the large language training software and for receiving the data used by the training process. Such a modification would amount to implementing Zhu’s software based model training technique using Newman’s known computer architecture for training and operating large language models, yielding the predictable result of computer implementation of Zhu’s training method.
The remaining limitations are taught using the same reasoning as in claim 1.
Regarding claims 9-14,
Claims 9-14 are system claims that are substantially similar to method claims 2-7, respectively, and are taught using the same reasoning.
Regarding claim 15,
Claim 15 is a non-transitory machine-readable medium claim that is substantially similar to method claim 1, with the addition of the following limitations, which are taught by Newman:
A non-transitory machine-readable medium comprising a plurality of machine-executable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform operations comprising: ([0058] memory 720 may include non-transitory, tangible, machine readable media that includes executable code that when run by one or more processors (e.g., processor 710) may cause the one or more processors to perform the methods described in further detail herein)
OBVIOUSNESS:
It would have been obvious to one of ordinary skill in the art, before the effective filing date, to store instructions for performing Zhu’s model training method on a non-transitory machine readable medium as taught by Newman, so that the disclosed model training operations can be stored and executed by the one or more processors. This would constitute use of a known software storage technique for its established purpose and would predictably enable processor based execution of Zhu’s training method.
The remaining limitations are taught using the same reasoning as in claim 1.
Regarding claims 16-20,
Claims 16-20 are non-transitory machine-readable medium claims that are substantially similar to method claims 2-6, respectively, and are taught using the same reasoning.
CONCLUSION
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Matthew Alan Cady whose telephone number is (571) 272-7229. The examiner can normally be reached Monday - Friday, 7:30 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Cesar Paula can be reached on (571)272-4128. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC)
at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW ALAN CADY/ Examiner, Art Unit 2145
/CESAR B PAULA/Supervisory Patent Examiner, Art Unit 2145