DETAILED ACTION
1. This office action is in response to application 18/636,971 filed on 4/16/2024. Claims 1-15 are pending in this office action.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
3. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-9 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by US 2021/0303925 (hereinafter Hofmann).
As for claim 1 Hofmann discloses: A method performed by one or more computers and for training a population of adversarial neural networks using a base neural network (See paragraphs 0040 and 0053), wherein training the population comprises, at each of a plurality of training iterations and for each adversarial neural network in the population: receiving an adversarial input (See paragraphs 0032 and 0053) and; processing the adversarial input using the adversarial neural network to generate one or more adversarial base network inputs for the base neural network, wherein each adversarial base network input is generated in accordance with a respective training task that requires generating adversarial base network inputs that violate the one or more downstream task criteria (See paragraphs 0013 and 0048 note the loss function determines violations); processing the one or more adversarial base network inputs using the base neural network to generate one or more respective adversarial base network outputs for each adversarial base network input; determining one or more adversarial rewards for the adversarial base network outputs that each measure a likelihood that the outputs violate a corresponding set of one or more downstream task criteria (See paragraphs 0079-0082 note predictions are likelihoods of an occurrence); and training the adversarial neural network in accordance with the respective training task by optimizing an adversarial reinforcement learning loss function based at least on the adversarial reward (See paragraphs 0003 and 0006).
As for claim 2 the rejection of claim 1 is incorporated and further Hofmann discloses: wherein the adversarial input comprises data that the adversarial neural network can process to generate adversarial base network inputs in accordance with causing the base neural network to generate base network outputs that violate at least one of the downstream task criteria (See paragraph 0003).
As for claim 3 the rejection of claim 1 is incorporated and further Hofmann discloses: wherein each adversarial neural network in the population has been assigned a respective training task comprising causing the base neural network to violate a respective first downstream task criterion, and wherein the adversarial reward for the adversarial neural network comprises a respective measure of a likelihood that the one or more base network outputs violate the respective first downstream task criterion (See paragraph 0090 note any training task moves downstream).
As for claim 4 the rejection of claim 1 is incorporated and further Hofmann discloses: training the base neural network at each of a second plurality of training iterations using one or more adversarial base network inputs generated by at least one adversarial neural network of the population (See paragraph 0106).
As for claim 5 the rejection of claim 4 is incorporated and further Hofmann discloses: wherein training the base neural network comprises training the base neural network using reinforcement learning, comprising, at each of the second plurality of training iterations: receiving a plurality of inputs comprising the one or more generated adversarial base network inputs; generating one or more base network outputs for each of the plurality of inputs; determining one or more base rewards for the base network outputs that each measure a likelihood that the outputs violate a corresponding set of one or more downstream task criteria; and training the base neural network by optimizing a base reinforcement learning loss function based at least on the base reward (See paragraphs 0037 and 0048).
As for claim 6 the rejection of claim 5 is incorporated and further Hofmann discloses: wherein the plurality of inputs further comprises one or more base network inputs that were not generated by the population of adversarial neural networks (See paragraph 0048 note the user input is not generated by the system).
As for claim 7 the rejection of claim 5 is incorporated and further Hofmann discloses: generating the adversarial reward and the base reward using one or more reward models (See paragraph 0106 note baseline).
As for claim 8 the rejection of claim 6 is incorporated and further Hofmann discloses: wherein each of the one or more reward models comprise a reward language processing neural network that has been trained to score text samples with respect to one or more criteria (See paragraph 0045 and claim 6 note embedded terms).
As for claim 9 the rejection of claim 1 is incorporated and further Hofmann discloses: wherein using one or more reward models comprises: using a rule reward model to determine a respective probability of the adversarial base network output violating a set of rules corresponding with the one or more downstream task criteria; and using a preference reward model to determine a preference score as a measure of one or more human preference criteria (See paragraph 0151).
Claim Rejections - 35 USC § 103
4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 10-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hofmann as applied to claim 1 above, and further in view of US 20230081171 (hereinafter Zhang).
As for claim 10 the rejection of claim 1 is incorporated and further Zhang discloses: wherein the base neural network and each adversarial neural network in the population comprise language processing models (See paragraph 0038). It would have been obvious to an artisan of ordinary skill in the pertinent at the time the instantly claimed invention was filed to have incorporated the teaching of Zhang into the system of Hofmann. The modification would have been obvious because the two references are concerned with the solution to problem of processing data via training models, therefore there is an implicit motivation to combine these references (i.e. motivation from the references themselves). In other words, the ordinary skilled artisan, during his/her quest for a solution to the cited problem, would look to the cited references at the time the invention was made. Consequently, the ordinary skilled artisan would have been motivated to combine the cited references since Zhang’s teaching would enable users of the Hofmann system to have more efficient processing.
As for claim 11 the rejection of claim 10 is incorporated and further Zhang discloses: wherein the base neural network and each adversarial neural network in the population comprise attention-based language models(See paragraph 0037). It would have been obvious to an artisan of ordinary skill in the pertinent at the time the instantly claimed invention was filed to have incorporated the teaching of Zhang into the system of Hofmann. The modification would have been obvious because the two references are concerned with the solution to problem of processing data via training models, therefore there is an implicit motivation to combine these references (i.e. motivation from the references themselves). In other words, the ordinary skilled artisan, during his/her quest for a solution to the cited problem, would look to the cited references at the time the invention was made. Consequently, the ordinary skilled artisan would have been motivated to combine the cited references since Zhang’s teaching would enable users of the Hofmann system to have more efficient processing.
.
As for claim 12 the rejection of claim 10 is incorporated and further Zhang discloses: wherein each attention-based language model comprises: a shared core having one or more pretrained parameters configured to process an input and generate an intermediate output; and one or more heads, each comprising a set of fine-tunable parameters, configured to process the intermediate output of the shared core. (See paragraphs 0069-0071). It would have been obvious to an artisan of ordinary skill in the pertinent at the time the instantly claimed invention was filed to have incorporated the teaching of Zhang into the system of Hofmann. The modification would have been obvious because the two references are concerned with the solution to problem of processing data via training models, therefore there is an implicit motivation to combine these references (i.e. motivation from the references themselves). In other words, the ordinary skilled artisan, during his/her quest for a solution to the cited problem, would look to the cited references at the time the invention was made. Consequently, the ordinary skilled artisan would have been motivated to combine the cited references since Zhang’s teaching would enable users of the Hofmann system to have more efficient processing.
As for claim 13 the rejection of claim 12 is incorporated and further Zhang discloses: wherein the one or more heads comprise: a policy head configured to process the intermediate output to generate a prompt response comprising a probability distribution over next text tokens in a sequence of text tokens as the base network output; a value head configured to process the intermediate output to generate a value comprising a prediction of a maximal reward that can be achieved by the prompt response; a teacher policy head configured to process the intermediate output to generate a reference prompt response comprising a pre-fine-tuned comparison baseline for the prompt response; and one or more reward heads configured to determine one or more respective rewards for the base network output (See paragraph 0054). It would have been obvious to an artisan of ordinary skill in the pertinent at the time the instantly claimed invention was filed to have incorporated the teaching of Zhang into the system of Hofmann. The modification would have been obvious because the two references are concerned with the solution to problem of processing data via training models, therefore there is an implicit motivation to combine these references (i.e. motivation from the references themselves). In other words, the ordinary skilled artisan, during his/her quest for a solution to the cited problem, would look to the cited references at the time the invention was made. Consequently, the ordinary skilled artisan would have been motivated to combine the cited references since Zhang’s teaching would enable users of the Hofmann system to have more efficient processing.
.
As for claim 14 the rejection of claim 13 is incorporated and further Zhang discloses: wherein the one or more reward heads generate the one or more adversarial rewards. (See paragraph 0148). It would have been obvious to an artisan of ordinary skill in the pertinent at the time the instantly claimed invention was filed to have incorporated the teaching of Zhang into the system of Hofmann. The modification would have been obvious because the two references are concerned with the solution to problem of processing data via training models, therefore there is an implicit motivation to combine these references (i.e. motivation from the references themselves). In other words, the ordinary skilled artisan, during his/her quest for a solution to the cited problem, would look to the cited references at the time the invention was made. Consequently, the ordinary skilled artisan would have been motivated to combine the cited references since Zhang’s teaching would enable users of the Hofmann system to have more efficient processing.
.
As for claim 15 the rejection of claim 13 is incorporated and further Zhang discloses: wherein the one or more reward heads generate the one or more base rewards (See paragraph 0148). It would have been obvious to an artisan of ordinary skill in the pertinent at the time the instantly claimed invention was filed to have incorporated the teaching of Zhang into the system of Hofmann. The modification would have been obvious because the two references are concerned with the solution to problem of processing data via training models, therefore there is an implicit motivation to combine these references (i.e. motivation from the references themselves). In other words, the ordinary skilled artisan, during his/her quest for a solution to the cited problem, would look to the cited references at the time the invention was made. Consequently, the ordinary skilled artisan would have been motivated to combine the cited references since Zhang’s teaching would enable users of the Hofmann system to have more efficient processing.
.
Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ELIYAH STONE HARPER whose telephone number is (571)272-0759. The examiner can normally be reached on Monday-Friday 10:00 am - 6:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sanjiv Shah can be reached on (571) 272-4098. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Eliyah S. Harper/Primary Examiner, Art Unit 2166 July 19, 2026