DETAILED ACTION
This Office Action is in response to Applicant's Communication received on 11/13/2023 for application number 18/507,605.
Claims 1-20 are presented for examination. Claims 1, 8, and 15 are independent claims.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 11/13/2023 has been considered by the Examiner.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 6-7, 8, and 13-14 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Gulcehre et al. (US 2025/0036958 A1 hereinafter Gulcehre).
Regarding Claim 1, Gulcehre teaches a system ([0018] neural network training system 100 implemented as computer programs on one or more computers), comprising:
a memory that stores computer-executable components ([0121] the subject matter implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus; [0127] computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit; central processing unit will receive instructions and data from a read only memory or a random access memory); and
a processor that executes the computer-executable components stored in the memory ([0127] computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit; central processing unit will receive instructions and data from a read only memory or a random access memory), wherein the computer-executable components comprise:
a training component that trains a semantic parser to predict one or more parses for an input text ([0019]-[0021] the system that obtains data specifying an initial, pre-trained generative neural network and further trains the pre-trained neural network; the neural network generates a new output example conditioned on a context input; generates output sequences of tokens from a vocabulary conditioned on a context sequence; [0111] neural network models a computer language; the context input may represent a sequence of text in a natural language that specifies the operation of a computer program or other computer code, and the output example may represent a sequence of text in the computer language; [0065] the system can generate multiple different output examples for each context input using the neural network by making use of the stochastic sampling - thus, the system (i.e., training component) trains the neural network (i.e., semantic parser) to output sequences of text in computer language (i.e., parses) for an input text) using offline reinforcement learning ([0039] the system uses the reward function to train the generative neural network in an offline manner through reinforced self-training; [0096] the system can train the neural network using the reward scores to optimize any appropriate offline reinforcement learning objective) based on parallelizable offline sampling ([0008] the described techniques use offline learning and divide the training into two parts, a “grow” part that involves sampling from the generative neural network model, and an “improve” part that involves further training the model, in a way that allows the computationally intensive “grow” part to be parallelized; [0010] because the “grow step” in which the system needs to sample from the generative neural network is performed offline and separately from the “improve” step, the system can generate a large number of model samples in parallel by distributing the sampling across multiple sets of one or more hardware devices; [0071] the system perform multiple iterations of this parallelized sampling during each grow step; [0072] by parallelizing the sampling in this manner, the system can significantly decrease the time required to perform a grow step).
As to dependent Claim 6, Gulcehre teaches all the limitations of claim 1. Gulcehre further teaches wherein the offline reinforcement learning based on the parallelizable offline sampling generates a self-annotated dataset ([0041] the training is referred to as “self-training” because, at each training stage, the system generates the training data for the training using the current version of the neural network as of the training stage; [0052] at each grow step, the system increases the size of the data set by generating additional outputs for some or all of the context inputs in the data set; [0071] the system perform multiple iterations of parallelized sampling during each grow step; [0078] the system generates an expanded training data set that includes a plurality of training examples and a respective reward score for each training example; when the system has access to one or more ground truth output examples for some or all of the context inputs, the system can also include the ground truth output examples for the context inputs in the expanded training data set (i.e., self-annotated dataset)).
As to dependent Claim 7, Gulcehre teaches all the limitations of claim 6. Gulcehre further teaches wherein a text generator is trained on the self-annotated dataset ([0041] the training is referred to as “self-training” because, at each training stage, the system generates the training data for the training using the current version of the neural network as of the training stage; [0078] the system generates an expanded training data set that includes a plurality of training examples and a respective reward score for each training example; when the system has access to one or more ground truth output examples for some or all of the context inputs, the system can also include the ground truth output examples for the context inputs in the expanded training data set; [0080] at each improve step, the system trains the neural network using the expanded data set generated by performing the grow step to update the values of the parameters of the neural network).
Claims 8 and 13-14 are method claims corresponding to the system claims 1 and 6-7 respectively and therefore, rejected for the same reasons.
Claim 15 is a product claim corresponding to the system claim 1 above and therefore, rejected for the same reasons. Gulcehre further teaches a computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform steps ([0121] the subject matter implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus; [0127] computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit; central processing unit will receive instructions and data from a read only memory or a random access memory).
As to dependent Claim 20, Gulcehre teaches all the limitations of claim 15. Gulcehre further teaches wherein generate, by the processor, a self-annotated training dataset based on the offline reinforcement learning based on the parallelizable offline sampling ([0121] the subject matter implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus; [0127] computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit; central processing unit will receive instructions and data from a read only memory or a random access memory; [0041] the training is referred to as “self-training” because, at each training stage, the system generates the training data for the training using the current version of the neural network as of the training stage; [0052] at each grow step, the system increases the size of the data set by generating additional outputs for some or all of the context inputs in the data set; [0071] the system perform multiple iterations of parallelized sampling during each grow step; [0078] the system generates an expanded training data set that includes a plurality of training examples and a respective reward score for each training example; when the system has access to one or more ground truth output examples for some or all of the context inputs, the system can also include the ground truth output examples for the context inputs in the expanded training data set (i.e., self-annotated dataset)); and train, by the processor, a text generator based on the self-annotated training dataset ([0121] the subject matter implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus; [0127] computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit; central processing unit will receive instructions and data from a read only memory or a random access memory; [0041] the training is referred to as “self-training” because, at each training stage, the system generates the training data for the training using the current version of the neural network as of the training stage; [0078] the system generates an expanded training data set that includes a plurality of training examples and a respective reward score for each training example; when the system has access to one or more ground truth output examples for some or all of the context inputs, the system can also include the ground truth output examples for the context inputs in the expanded training data set; [0080] at each improve step, the system trains the neural network using the expanded data set generated by performing the grow step to update the values of the parameters of the neural network).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-3, 9-10, and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Gulcehre in view of Mager et al. ("GPT-too: A Language-Model-First Approach for AMR-to-Text Generation" hereinafter Mager).
As to dependent Claim 2, Gulcehre teaches all the limitations of claim 1. Gulcehre further teaches wherein a weighting component that weights respective parses of the one or more parses as functions of a score that indicates respective levels of coherence of the respective parses to the input text ([0032] an output example generated by the generative neural network for the context input to generate as output a reward score that measures the quality of the output example relative to the context input; [0034] the reward function can be a textual coherence measure (i.e., score that indicates level of coherence); [0074] the system generates a respective reward score for each training example using a reward function; [0090] the system filter the expanded training data set to remove the training examples that have a reward score that is below the respective threshold value - thus, weighing the parses based on reward score/ coherence measure).
However, Gulcehre does not expressly teach wherein weighing parses as functions of a large language model (LLM)-produced cycle consistency score.
In the same field of endeavor, Mager teaches wherein weighing parses as functions of a large language model (LLM)-produced cycle consistency score (page 2, section 3 - cycle consistency is to assess the quality of a system’s output based on how well an external ‘reverse’ system can reconstruct the input from it; page 2, section 3 - the use of a cycle consistency measure to re-score the system outputs (i.e., weighing parses); page 5, section 5 - cycle consistency-based re-scoring using AMR parser and the Smatch metric).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated wherein weighing parses as functions of a large language model (LLM)-produced cycle consistency score, as suggested in Kim into Sherman. Doing so would be desirable because it would improve performance of the language model (Mager, page 2, section 1), thereby enhancing user experience.
As to dependent Claim 3, Gulcehre and Mager teach all the limitations of claim 1. Mager further teaches wherein weighting the respective parses of the one or more parses as functions of an LLM-produced cycle consistency score results in a model that produces parses that are coherent with respect to the input text (page 2, section 3 - cycle consistency is to assess the quality of a system’s output based on how well an external ‘reverse’ system can reconstruct the input from it; page 2, section 1 - re-scoring technique based on cycle-consistency that further improves performance of language model).
Claims 9-10 are method claims corresponding to the system claims 2-3 above and therefore, rejected for the same reasons.
Claims 16-17 are product claims corresponding to the system claims 2-3 above and therefore, rejected for the same reasons.
Claims 4-5, 11-12, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Gulcehre in view of Mager, further in view of Xiao et al. (US 2018/0349767 A1 hereinafter Xiao).
As to dependent Claim 4, Gulcehre and Mager teach all the limitations of claim 2. However, Gulcehre and Mager do not expressly teach wherein the weighting component further weights the respective parses as functions of a count-based prior probability that assigns scores above a defined threshold to parses that are syntactically valid and share a common substructure with the one or more parses.
In the same field of endeavor, Xiao teaches wherein the weighting component further weights the respective parses as functions of a count-based prior probability ([0066] combine WCFG and WFSA(s) to form the background priors guiding the RNN; [0040] the background b is an arbitrary non-negative function used to incorporate prior knowledge about the generative process; [0069] semantic parser taking into account prior knowledge about LF well-formedness and about the likelihood of certain entities being present (i.e., prior probability) based on the input; [0071] the well-formedness of the logical forms is modeled by a weighted context-free grammar; the likelihood that certain entities present in the input utterance are also present in the logical form is modeled by weighted finite-state automata; the grammar and automata are combined together through an efficient intersection algorithm to form a soft guide to the RNN; [0051] γx denotes the unigram probability of the output symbol x used to express certain forms of regularities on expected logical forms (i.e., count of symbol occurrences); observations about certain patterns that are likely or unlikely to occur in the logical forms) that assigns scores above a defined threshold to parses that are syntactically valid and share a common substructure with the one or more parses ([0042]the background b takes a value in {0,1}, with b(xt+1|x1, . . . , xt)=1 indicating that x1, . . . , xt, xt+1 is a valid DS prefix relative to G; the BRNN cannot produce “ungrammatical” (in the sense of being valid according to the grammar) prefixes - thus, binary {0,1} on grammatical validity is equivalent to scores above threshold for valid parses; [0028] the weight of a specific parse tree in a WCFG is the product or sum of all rule weights in the tree; included as often as the rule is used in the tree; [0051] automata express certain forms of regularities on expected logical forms, such as, like here, unigram probabilities of output - thus, the parses sharing a grammar rule share that weight contribution (i.e., common substructure with the parses)).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have incorporated wherein the weighting component further weights the respective parses as functions of a count-based prior probability that assigns scores above a defined threshold to parses that are syntactically valid and share a common substructure with the one or more parses, as suggested in Xiao into Gulcehre and Mager. Doing so would be desirable because incorporating simple input-dependent prior knowledge via WFSAs, the model improves over its RNN baseline(Xiao [0031]).
As to dependent Claim 5, Gulcehre, Mager, and Xiao teach all the limitations of claim 4. Xiao further teaches wherein weighting the respective parses as functions of the count- based prior probability results in a model that produces structurally regular parses ([0047] enumerates exactly the set of all valid derivation sequences relative to the original G; [0042] guarantees that the evolving DS prefix always remains valid relative to G; the BRNN cannot produce “ungrammatical” prefixes; [0051] automata on the output used to express certain forms of regularities on expected logical forms; observations about certain patterns that are likely or unlikely to occur in the logical forms - thus, weighing each parse by WCFG prior constrains the trained model to output well-formed grammar-valid logical forms composed of recurring substructures).
Claims 11-12 are method claims corresponding to the system claims 4-5 above and therefore, rejected for the same reasons.
Claims 18-19 are product claims corresponding to the system claims 4-5 above and therefore, rejected for the same reasons.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Applicant is required under 37 CFR § 1.111(c) to consider these references fully when responding to this action.
Li et al. (US 2015/0310864 A1) teaches: convert the first speech signals to obtain at least two first texts and sends the at least two first texts to the parsing unit; and the parsing unit is specifically configured to score semantics of each first text of the at least two first texts according to a predetermined scoring rule and according to naturalness and coherence of the semantics of the at least two first texts, where a higher score represents better naturalness and coherence of the semantics, and acquire, from the semantics of the at least two first texts, semantics with a highest score and of the first text as the first target semantics (see [0020]).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to REJI KARTHOLY whose telephone number is (571)272-3432. The examiner can normally be reached on Monday - Thursday from 7:30 am to 3:30 pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch, can be reached at telephone number 571-272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from Patent Center. Status information for published applications may be obtained from Patent Center. Status information for unpublished applications is available through Patent Center for authorized users only. Should you have questions about access to Patent Center, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated- interview-request-air-form.
/REJI KARTHOLY/Primary Examiner, Art Unit 2143