Prosecution Insights
Last updated: August 17, 2026
Application No. 18/304,341

RUNTIME ALIGNMENT OF LANGUAGE MODELS IN CONVERSATIONAL AI SYSTEMS AND APPLICATIONS

Final Rejection §103
Filed
Apr 20, 2023
Examiner
WITHEY, THEODORE JOHN
Art Unit
2655
Tech Center
2600 — Communications
Assignee
NVIDIA Corporation
OA Round
4 (Final)
41%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 41% of resolved cases
41%
Career Allowance Rate
11 granted / 27 resolved
-21.3% vs TC avg
Strong +47% interview lift
Without
With
+46.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
23 currently pending
Career history
66
Total Applications
across all art units

Statute-Specific Performance

§101
19.3%
-20.7% vs TC avg
§103
54.8%
+14.8% vs TC avg
§102
15.8%
-24.2% vs TC avg
§112
9.5%
-30.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 27 resolved cases

Office Action

§103
DETAILED ACTION This office action is in response to Applicant’s Amendment/Request for Reconsideration, received on 05/11/2026. Claims 1, 2, 9, 11, 12, 16, 17, 19 have been amended. Claims 1-20 are pending and have been considered. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 05/11/2026, see pgs. 11-12, with respect to “Claim Rejections under 35 U.S.C. 103” for independent claims 1, 11, and 19 (Lam in view of Moon) have been fully considered but they are not persuasive. Applicant’s representative asserts, “Independent claim 1 as amended states that a dialog flow is ‘specified in configuration information maintained separately from the language model.’ Claim 1 has also been amended to recite that ‘the dialog flow being specified in configuration information maintained separately from the language model and defining a sequence of past and future dialog between a user and outputs of the language model, the dialog flow including one or more associated next operations to be performed.’ Independent claims 11 and 19 have been similarly amended with respect to the one or more example dialog flows and the LLM. These amendments are supported by the Specification at least at [0035], [0039], [0092]-[0093] and FIGS. 2-3. FIG. 2 illustrates the language model separately from the canonical form input definitions, dialog flow definitions, and canonical form output definitions. FIG. 3 illustrates example configuration information including ‘define user,’ ‘define bot,’ ‘define flow,’ and ‘define subflow’ definitions. The cited portions of Lam do not disclose the amended limitation that the dialog flow is ‘specified in configuration information maintained separately from the language model.’ The Office Action relies primarily on Lam's Table 3 and associated DST, ACD, DAG, and RG stages as allegedly corresponding to the claimed dialog flow (Office Action, pp. 4-7, citing Lam, Fig. 14, Appendix, Table 3, and Section 5.2). However, the cited portions of Lam do not identify configuration information maintained separately from the language model that specifies the alleged dialog flow. Nor do the cited portions disclose a dialog flow, specified in such separately maintained configuration information, that comprises associated next steps specifying next operation to be performed as now recited in claim 1. Lam's API/database discussion also does not cure this deficiency. The Office Action cites Lam's API-call-related disclosures as part of the alleged operations used to execute a dialog flow (Office Action, pp. 5-6, citing Lam, Appendix, Table 3, Turn 1, ACD/DST/DAG/RG entries and Section 9). The cited portions do not disclose that any API/database action is specified by a dialog flow that is itself specified in configuration information maintained separately from the language model. Thus, even assuming Lam discloses an API/database action, the cited portions do not disclose the claimed arrangement in which the dialog flow comprises associated next operations that are executed to generate an output using the language model at model runtime. Moon does not remedy this deficiency. The Office Action relies on Moon for ANN and runtime-related teachings (Office Action, pp. 6-7, citing Moon, col. 23, II. 50-60; col. 24, II. 54-60; col. 25, II. 65-67-col. 26, II. 1-5; col. 28, 11. 45-50). The cited portions of Moon do not disclose a dialog flow that is specified in configuration information maintained separately from the language model. The cited portions also do not disclose associated next steps of such a separately maintained dialog flow being executed to generate an output using the language model. Accordingly, even if Lam and Moon are combined, the cited portions do not disclose or suggest the amended claim requirement that the dialog flow is ‘specified in configuration information maintained separately from the language model,’ where the dialog flow ‘being specified in configuration information maintained separately from the language model and defining a sequence of past and future dialog between a user and outputs of the language model, the dialog flow including one or more associated next operations to be performed.’” In response, the examiner agrees with Applicant’s arguments with respect to Lam, but respectfully asserts that the combination of Lam in view of Moon, wherein the structure of Moon will be discussed with newly cited portions, discloses the entered amendment. Specifically, Fig. 1 of Moon discloses a network environment 100 which contains a client system 130 containing assistant application 136 which is connected via network 110 to an assistant system 140 and a third-party system 170. [Col. 6, Lines 55-61] discloses “The assistant application 136 may communicate the user input to the assistant system 140. Based on the user input, the assistant system 140 may generate responses. The assistant system 140 may send the generated responses to the assistant application 136. The assistant application 136 may then present the responses to the user at the client system 130”. The examiner asserts that the assistant application tracks to configuration information maintained separately from the language model, i.e. the assistant system, wherein Moon described the assistant system to be containing a language model: [Col. 6, Lines 20-25] “In particular embodiments, the NLG 271 may use different language models and/or language templates to generate natural language outputs”, wherein the NLG 271 is part of the assistant system as described in Fig. 2. The operations of Lam will now be applied to this structure of Moon. See updated rejections below. Specification The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: “Runtime Alignment of Language Models in Conversation AI Systems and Applications”. The examiner would like to note that this was the title for all actions prior to the amendment entered on 05/11/2026. It is unclear to the examiner why the title has been changed to be less descriptive. Regardless of whether or not Applicant opts to select the examiner’s selected title or would like to provide their own, the specification needs to be updated to reflect this change in title (see top of pg. 1 of specification of instant application). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 5, 7, 9-13, 16-17, 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lam et al. (“Zero and Few-Shot Localization of Task-Oriented Dialogue Agents with a Distilled Representation”), hereinafter Lam, in view of Moon et al. (US-11442992-B1), hereinafter Moon. Regarding claim 1, Lam discloses: a method comprising: generating, based at least on a user input ([pg. 14, Appendix, Table 3, Turn 1, Input] I’d like hotel recommendations), a canonical form representing a constrained semantic representation of the user input ([pg. 14, Appendix, Table 3, Turn 1, DST Target] ( hotels search ) [Concatenating the request to provide a recommendation to “hotel search” is a constrained semantic representation of input, i.e. a canonical form]); determining, based at least on the canonical form, a dialog flow that controls output of a language model ([pg. 14, Appendix, Table 3, Turn 1, DAG Target] ( hotels search ) request rating , request stars [Requesting a rating or stars given to a hotel from a user (In view of response generation “Do you have any requirements for the hotel’s rating or the number of stars of the hotel?”) before making a final recommendation is indicative of a dialog flow that controls the output of a language model, i.e. generated response, based on the canonical form “hotels search”. Further consider Section 5.2 where the used model is claimed to be mBART, a well-known language model]). Lam does not disclose: a language model implemented as an artificial neural network; and, the dialog flow being specified in configuration information maintained separately from the language model. Moon discloses: a language model implemented as an artificial neural network ([Col. 23, Lines 50-60] The assistant system 140 may then select, by a conversational reasoning model, one or more candidate nodes from the knowledge graph corresponding to one or more candidate entities, respectively. Each candidate node may be selected based on the nodes corresponding to the initial entities, one or more dialog states associated with the query, and a context associated with the query, [Col. 24, Lines 54-66] FIG. 10 illustrates an example artificial neural network (“ANN”) 1000. In particular embodiments, an ANN may refer to a computational model comprising one or more nodes. Example ANN 1000 may comprise an input layer 1010, hidden layers 1020, 1030, 1040, and an output layer 1050. Each layer of the ANN 1000 may comprise one or more nodes, such as a node 1005 or a node 1015. In particular embodiments, each node of an ANN may be connected to another node of the ANN. As an example and not by way of limitation, each node of the input layer 1010 may be connected to one of more nodes of the hidden layer 1020. In particular embodiments, one or more nodes may be a bias node, [A conversational reasoning, i.e. language, model for selecting nodes, wherein an ANN is defined as the model performing the operations, indicates the conversational reasoning model to be implemented using the ANN]); and, the dialog flow being specified in configuration information maintained separately from the language model ([Fig. 1, Assistant System 140 connected to Assistant Application 136 via Network 110], [Col. 6, Lines 55-61] The assistant application 136 may communicate the user input to the assistant system 140. Based on the user input, the assistant system 140 may generate responses. The assistant system 140 may send the generated responses to the assistant application 136. The assistant application 136 may then present the responses to the user at the client system 130, [The examiner asserts that the assistant application tracks to a separate entity from the assistant system 140 containing natural language generator (NLG) 271 which is disclosed to be containing a language model ([Col. 22, Lines 20-25]). There is motivation to take the user input of Moon and transform it into the dialog flow format of Lam as Lam discloses generating dialog flows for current turns based on previous turns and based on user utterances ([Section 3.2, “3. Dialogue Act Generation”]), all disclosed functionalities of the assistant system 140 of Moon ([Col. 17, Lines 1-55]). The user input of Moon would be transformed using the dialogue act generation of Lam before generating responses, resulting in configuration information to be sent to the assistant system. Further, there is motivation to separate the functionalities of each step of Lam into distinct components as seen in Moon because splitting the stages of Lam into multiple networks/models reduces the computational load on each network/model (Lam discloses cost of virtual assistants in the introduction)]). Lam and Moon are considered analogous art within conversational reasoning within knowledge bases. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam to incorporate the teachings of Moon, because of the novel way to associate walk paths of a knowledge graph with input contexts including dialog state, sentence, and initial entities mentioned in the conversation for ranking candidate entities using a zero-shot relevance learning model which results in more accurate and relevant entities in generated responses within multi-turn dialogs (Moon, [Col. 3, Lines 15-40]). Lam further discloses: the dialog flow defining a sequence of past and future dialog between a user and outputs of the language model ([pg. 14, Appendix, Table 3, Turn 1, RG Target, Turn 2, DST Input Agent Acts and User response], [In view of the above dialog flow “(hotels search) request rating, request stars”, wherein each request defines a sequence of dialog to be performed between a user and outputs of language model, i.e. the requests, in later turns such as rating and stars in turn 2, wherein at the DST stage, the “(hotels search)” is defining a past dialog request to be clarified with the future “request rating” and “request stars” operations]), the dialog flow including one or more associated next operations to be performed ([As previously disclosed, additional questions to be asked based on stars and/or hotel rating are one or more associated next operations to be performed based on the original request. Future dialog tracks to next operations to be performed/answered/responded to]); and, performing the one or more associated next operations specified by the dialog flow to generate an output using the language model at runtime ([pg. 15, Appendix, Table 3, Turn 1, RG Prediction] Do you have a preference on how many stars and what rating the hotel should have? [Sending a response to a user is an operation to execute the steps of asking for stars and rating, i.e. dialog flow, based on the previously determined dialog flow DAG Prediction to generate an output, i.e. the response. Further, see section 9, “Ethical Considerations”, which disclose runs indicating the operation to be performed at runtime]), the dialog flow configuring the language model to generate the output according to constraints defined in the dialog flow based at least on (1) a match between the canonical form and a user input defined in the dialog flow ([In view of the dialog flow being generated based upon the canonical form, which itself is based upon user input, it is unclear to the examiner how there would not inherently always be a match between the canonical form and user input defined in the dialog flow for generating output as the dialog flow is generated based on the canonical form, and, therefore, the user input]) and (2) a corresponding canonical form of a language model output defined in the dialog flow ([pg. 14, Appendix, Table 3, Turn 2, DAG prediction “request location/price_level”], [“request location” is the canonical form of the generated language model output “And what about location?” defined in the multi-turn dialog flow]). Moon further discloses: the constraints being applied at runtime to the language model ([Col. 25, Lines 65-67]-[Col. 26, Lines 1-5] (2) a zero-shot learning model that leverages previous sentence, dialog, and KG contexts to re-rank candidates from pruned decoder graph output based on their relevance and path scores, which allows for generalizable and robust classification with a large number of candidate classes, [A zero-shot learning model indicates the constraints of a previous sentence/dialog/context are used for generating/ranking candidates without prior training]), and training data sets used to train the language model excluding feedback related to the constraints such that the language model is not pre-trained with the constraints ([Col. 28, Lines 45-50] The embodiments disclosed herein compute zero-shot relevance score in the KG embeddings space, thus allowing for robust prediction for KG entities and domains unseen during training as well, [An entity/domain unseen during training indicates there is no pre-training operation associated with the entity/domain, necessarily having constraints related to the entity/domain]). Regarding claim 2, Lam in view of Moon discloses: the method of claim 1. Lam further discloses: wherein the performing the one or more associated next operations comprises using at least the language model to generate the output ([pg. 5, Section 5.2, Par. 1] All models use a standard Seq2Seq architecture with a bidirectional encoder and left-to-right autoregressive decoder. mBART is pre-trained to denoise text in 50 languages, while mT5 is trained on 101 languages [mBART can reasonably be classified as a language model]). Regarding claim 3, Lam in view of Moon discloses: the method of claim 1. Lam further discloses: generating a second canonical form based at least on the output ([pg. 14, Appendix, Table 3, Turn 2, DST Prediction] ( hotels search ) rating equal_to " don’t care " , stars at_least " 5 " [In view of the previously disclosed ( hotels search ) canonical form of Turn 1, it can be seen that the addition of rating equal_to and stars at_least is an second, updated canonical form based on the output from the first turn, i.e. asking for rating and stars]); determining a second dialog flow based at least on the second canonical form ([pg. 14, Appendix, Table 3, Turn 2, DAG Prediction] ( hotels search ) request location , request price_level [In view of turn 1, it can be seen that a second dialog flow, i.e. determining location and price_level requests, in view of the second canonical form (see above element) indicating a hotel search with rating and stars already decided (tracking to output from a first turn), indicating the flow should ask other questions based on the elements of a second canonical form, which is based on output from a first turn]); and, performing one or more second operations to execute the second dialog flow to generate a second output ([pg. 14, Appendix, Table 3, Turn 2, RG Prediction] And what about location? Do you have a price range for the hotel? [Asking a user for location a pricing is executing the steps of the dialog flow determined in the DAG prediction to generate a second output, i.e. the response, in view of the first output RG prediction of turn 1]). Regarding claim 5, Lam in view of Moon discloses: the method of claim 1. Lam further discloses: wherein the generating the canonical form comprises processing the user input using a trained machine learning model ([pg. 5, Section 5.2, Par. 1] We use mbart-large-50 as the neural model for our agent in all our experiments. All models use a standard Seq2Seq architecture with a bidirectial encoder and left-to-right autoregressive decoder. mBART is pre-trained to denoise text in 50 languages [In view of the previously generated canonical forms of Lam, disclosed in claim 1 rejection]). Regarding claim 7, Lam in view of Moon discloses: the method of claim 1. Lam further discloses: wherein the determining the dialog flow comprises generating the dialog flow based at least on the canonical form ([pg. 14, Table 3, Turn 1, DST Prediction] ( hotels search ) [Based on an input “I’d like hotel recommendations”, indicating “hotels search” is a canonical form], [pg. 14, Table 3, Turn 1, DAG Prediction]) ( hotels search ) request rating , request stars [Determining to ask for a rating and/or stars for the hotel indicates a generated dialog flow based on the canonical form “hotels search”]). Regarding claim 9, Lam in view of Moon discloses: the method of claim 1. Lam further discloses: wherein the performing the one or more associated next operations: generating an embedding of a second canonical form associated with the dialog flow in a semantic or latent space ([pg. 14, Table 3, Turns 1, 2 DST Prediction] [In view of the sentence embedding of Lam ([pg. 5, section 4.2, Par. 1]), indicating the canonical forms can also be embedded as they are in some form of sentence, further in view of the addition of rating equal_to “don’t care” and stars at_least “5” to the turn 2 DST prediction indicating a second canonical form associated with the dialogue flow in view of the canonical form “(hotels search)” of turn 1]); determining one or more canonical forms based at least on the embedding of the second canonical form and one or more embeddings of one or more predefined canonical forms in the semantic or latent space ([Fig. 1, History], [pg. 14, Table 3, Turn 2, ACD, DAG], [Determining to add request location and request price_level canonical forms to the canonical form DAG prediction which comes after the rating and star determinations of the earlier DST section of turn 2 indicates determination of the canonical form “(hotels search) request location request price_level” is based on the second canonical form, i.e. [pg. 14, Appendix, Table 3, Turn 2, DST Prediction] ( hotels search ) rating equal_to " don’t care " , stars at_least " 5 ", e.g. not needing to include these pieces of information again, and a predefined canonical form in a semantic or latent space in view of the previous dialogue acts and retrieved results of Fig. 1 of Lam indicating predefined, i.e. historical canonical forms, further in view of the API and ACTS calls of the turns indicating embedding to transmit information and perform those calls in a semantic or latent space]); generating a prompt that includes the one or more canonical forms ([pgs. 14-15, Turns 1-3, DAG Predictions]), one or more example outputs associated with the one or more canonical forms ([pgs. 14-15, Turns 1-2, RG Predictions]), and at least a portion of a current conversation ([pgs. 14-15, Turns 1-3, DST Inputs]); and, processing the prompt using the language model to generate the output ([pgs. 14-15, Turn 3, RG Prediction] “There are 4 available hotels. I recommend Royal Plaza Hotel. Its rating is 9.” [In view of the mBART, i.e. language model, used to perform the operations of Lam as disclosed in Section 5.2]). Regarding claim 10, Lam in view of Moon discloses: the method of claim 1. Lam further discloses: wherein the canonical form and the dialog flow are specified in a formal modeling language ([pg. 15, Turn 3, ACD Input], [In view of [0035] of the instant application which defines modeling language as a programming language that “requires a particular syntax defining combinations of symbols that are considered to be correctly structured statements” indicating that the labels <state>…<endofstate> and <history>…<endofhistory>, which respectively correspond to dialogue flow and canonical forms, place dialog flow and canonical form information in a formal modeling language format]). Regarding claim 11, Lam discloses: a processor comprising: one or more processing units to perform operations ([Section 5.2, Par. 1] We also use the Dialogues4 library for data preprocessing and evaluation [Preprocessing data indicates a processor to perform that action]) comprising: generating, based at least on a user input ([pg. 14, Appendix, Table 3, Turn 1, Input] I’d like hotel recommendations), a canonical form representing a constrained semantic representation of the user input ([pg. 14, Appendix, Table 3, Turn 1, DST Target] ( hotels search ) [Concatenating the request to provide a recommendation to “hotel search” is a constrained semantic representation of input, i.e. a canonical form]); determining, based at least on the canonical form, a dialog flow that controls output of a language model ([pg. 14, Appendix, Table 3, Turn 1, DAG Target] ( hotels search ) request rating , request stars [Requesting a rating or stars given to a hotel from a user (In view of response generation “Do you have any requirements for the hotel’s rating or the number of stars of the hotel?”) before making a final recommendation is indicative of a dialog flow that controls the output of a language model, i.e. generated response, based on the canonical form “hotels search”. Further consider Section 5.2 where the used model is claimed to be mBART, a well-known language model]). Lam does not disclose: a language model implemented as an artificial neural network; and, the dialog flow being specified in configuration information maintained separately from the language model. Moon discloses: a language model implemented as an artificial neural network ([Col. 23, Lines 50-60] The assistant system 140 may then select, by a conversational reasoning model, one or more candidate nodes from the knowledge graph corresponding to one or more candidate entities, respectively. Each candidate node may be selected based on the nodes corresponding to the initial entities, one or more dialog states associated with the query, and a context associated with the query, [Col. 24, Lines 54-66] FIG. 10 illustrates an example artificial neural network (“ANN”) 1000. In particular embodiments, an ANN may refer to a computational model comprising one or more nodes. Example ANN 1000 may comprise an input layer 1010, hidden layers 1020, 1030, 1040, and an output layer 1050. Each layer of the ANN 1000 may comprise one or more nodes, such as a node 1005 or a node 1015. In particular embodiments, each node of an ANN may be connected to another node of the ANN. As an example and not by way of limitation, each node of the input layer 1010 may be connected to one of more nodes of the hidden layer 1020. In particular embodiments, one or more nodes may be a bias node, [A conversational reasoning, i.e. language, model for selecting nodes, wherein an ANN is defined as the model performing the operations, indicates the conversational reasoning model to be implemented using the ANN]); and, the dialog flow being specified in configuration information maintained separately from the language model ([Fig. 1, Assistant System 140 connected to Assistant Application 136 via Network 110], [Col. 6, Lines 55-61] The assistant application 136 may communicate the user input to the assistant system 140. Based on the user input, the assistant system 140 may generate responses. The assistant system 140 may send the generated responses to the assistant application 136. The assistant application 136 may then present the responses to the user at the client system 130, [The examiner asserts that the assistant application tracks to a separate entity from the assistant system 140 containing natural language generator (NLG) 271 which is disclosed to be containing a language model ([Col. 22, Lines 20-25]). There is motivation to take the user input of Moon and transform it into the dialog flow format of Lam as Lam discloses generating dialog flows for current turns based on previous turns and based on user utterances ([Section 3.2, “3. Dialogue Act Generation”]), all disclosed functionalities of the assistant system 140 of Moon ([Col. 17, Lines 1-55]). The user input of Moon would be transformed using the dialogue act generation of Lam before generating responses, resulting in configuration information to be sent to the assistant system. Further, there is motivation to separate the functionalities of each step of Lam into distinct components as seen in Moon because splitting the stages of Lam into multiple networks/models reduces the computational load on each network/model (Lam discloses cost of virtual assistants in the introduction)]). Lam and Moon are considered analogous art within conversational reasoning within knowledge bases. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam to incorporate the teachings of Moon, because of the novel way to associate walk paths of a knowledge graph with input contexts including dialog state, sentence, and initial entities mentioned in the conversation for ranking candidate entities using a zero-shot relevance learning model which results in more accurate and relevant entities in generated responses within multi-turn dialogs (Moon, [Col. 3, Lines 15-40]). Lam further discloses: the dialog flow defining a sequence of past and future dialog between a user and outputs of the language model ([pg. 14, Appendix, Table 3, Turn 1, RG Target, Turn 2, DST Input Agent Acts and User response], [In view of the above dialog flow “(hotels search) request rating, request stars”, wherein each request defines a sequence of dialog to be performed between a user and outputs of language model, i.e. the requests, in later turns such as rating and stars in turn 2, wherein at the DST stage, the “(hotels search)” is defining a past dialog request to be clarified with the future “request rating” and “request stars” operations]), the dialog flow including one or more associated next operations to be performed ([As previously disclosed, additional questions to be asked based on stars and/or hotel rating are one or more associated next operations to be performed based on the original request. Future dialog tracks to next operations to be performed/answered/responded to]); and, performing the one or more associated next operations specified by the dialog flow to generate an output using the language model at runtime ([pg. 15, Appendix, Table 3, Turn 1, RG Prediction] Do you have a preference on how many stars and what rating the hotel should have? [Sending a response to a user is an operation to execute the steps of asking for stars and rating, i.e. dialog flow, based on the previously determined dialog flow DAG Prediction to generate an output, i.e. the response. Further, see section 9, “Ethical Considerations”, which disclose runs indicating the operation to be performed at runtime]), the dialog flow configuring the language model to generate the output according to constraints defined in the dialog flow based at least on (1) a match between the canonical form and a user input defined in the dialog flow ([In view of the dialog flow being generated based upon the canonical form, which itself is based upon user input, it is unclear to the examiner how there would not inherently always be a match between the canonical form and user input defined in the dialog flow for generating output as the dialog flow is generated based on the canonical form, and, therefore, the user input]) and (2) a corresponding canonical form of a language model output defined in the dialog flow ([pg. 14, Appendix, Table 3, Turn 2, DAG prediction “request location/price_level”], [“request location” is the canonical form of the generated language model output “And what about location?” defined in the multi-turn dialog flow]). Moon further discloses: the constraints being applied at runtime to the language model ([Col. 25, Lines 65-67]-[Col. 26, Lines 1-5] (2) a zero-shot learning model that leverages previous sentence, dialog, and KG contexts to re-rank candidates from pruned decoder graph output based on their relevance and path scores, which allows for generalizable and robust classification with a large number of candidate classes, [A zero-shot learning model indicates the constraints of a previous sentence/dialog/context are used for generating/ranking candidates without prior training]), and training data sets used to train the language model excluding feedback related to the constraints such that the language model is not pre-trained with the constraints ([Col. 28, Lines 45-50] The embodiments disclosed herein compute zero-shot relevance score in the KG embeddings space, thus allowing for robust prediction for KG entities and domains unseen during training as well, [An entity/domain unseen during training indicates there is no pre-training operation associated with the entity/domain, necessarily having constraints related to the entity/domain]). Regarding claim 12, Lam in view of Moon discloses: the processor of claim 11. Lam further discloses: wherein the performing the one or more associated next operations comprises using at least the language model to generate the output ([pg. 5, Section 5.2, Par. 1] All models use a standard Seq2Seq architecture with a bidirectional encoder and left-to-right autoregressive decoder. mBART is pre-trained to denoise text in 50 languages, while mT5 is trained on 101 languages [mBART can reasonably be classified as a language model]). Regarding claim 13, Lam in view of Moon discloses: the processor of claim 11. Lam further discloses: wherein the one or more processing units further perform operations comprising: generating a second canonical form based at least on the output ([pg. 14, Appendix, Table 3, Turn 2, DST Prediction] ( hotels search ) rating equal_to " don’t care " , stars at_least " 5 " [In view of the previously disclosed ( hotels search ) canonical form of Turn 1, it can be seen that the addition of rating equal_to and stars at_least is an second, updated canonical form based on the output from the first turn, i.e. asking for rating and stars]); determining a second dialog flow based at least on the second canonical form ([pg. 14, Appendix, Table 3, Turn 2, DAG Prediction] ( hotels search ) request location , request price_level [In view of turn 1, it can be seen that a second dialog flow, i.e. determining location and price_level requests, in view of the second canonical form (see above element) indicating a hotel search with rating and stars already decided (tracking to output from a first turn), indicating the flow should ask other questions based on the elements of a second canonical form, which is based on output from a first turn]); and, performing one or more second operations to execute the second dialog flow to generate a second output ([pg. 14, Appendix, Table 3, Turn 2, RG Prediction] And what about location? Do you have a price range for the hotel? [Asking a user for location a pricing is executing the steps of the dialog flow determined in the DAG prediction to generate a second output, i.e. the response, in view of the first output RG prediction of turn 1]). Regarding claim 16, Lam in view of Moon discloses: the processor of claim 11. Lam further discloses: wherein the performing the one or more associated next operations comprises: generating an embedding of a second canonical form associated with the dialog flow in a semantic or latent space ([pg. 14, Table 3, Turns 1, 2 DST Prediction] [In view of the sentence embedding of Lam ([pg. 5, section 4.2, Par. 1]), indicating the canonical forms can also be embedded as they are in some form of sentence, further in view of the addition of rating equal_to “don’t care” and stars at_least “5” to the turn 2 DST prediction indicating a second canonical form associated with the dialogue flow in view of the canonical form “(hotels search)” of turn 1]); determining one or more canonical forms based at least on the embedding of the second canonical form and one or more embeddings of one or more predefined canonical forms in the semantic or latent space ([Fig. 1, History], [pg. 14, Table 3, Turn 2, ACD, DAG], [Determining to add request location and request price_level canonical forms to the canonical form DAG prediction which comes after the rating and star determinations of the earlier DST section of turn 2 indicates determination of the canonical form “(hotels search) request location request price_level” is based on the second canonical form, i.e. [pg. 14, Appendix, Table 3, Turn 2, DST Prediction] ( hotels search ) rating equal_to " don’t care " , stars at_least " 5 ", e.g. not needing to include these pieces of information again, and a predefined canonical form in a semantic or latent space in view of the previous dialogue acts and retrieved results of Fig. 1 of Lam indicating predefined, i.e. historical canonical forms, further in view of the API and ACTS calls of the turns indicating embedding to transmit information and perform those calls in a semantic or latent space]); generating a prompt that includes the one or more canonical forms ([pgs. 14-15, Turns 1-3, DAG Predictions]), one or more example outputs associated with the canonical forms ([pgs. 14-15, Turns 1-2, RG Predictions]), and at least a portion of a current conversation ([pgs. 14-15, Turns 1-3, DST Inputs]); and, processing the prompt using the language model to generate the output ([pgs. 14-15, Turn 3, RG Prediction] “There are 4 available hotels. I recommend Royal Plaza Hotel. Its rating is 9.” [In view of the mBART, i.e. language model, used to perform the operations of Lam as disclosed in Section 5.2]). Regarding claim 17, Lam in view of Moon discloses: the processor of claim 16. Lam further discloses: wherein the performing the one or more associated next operations further comprises: accessing at least one of a knowledge base, a computational knowledge engine, a search engines, or an automation service to generate a second output ([Fig. 1, “Knowledge Base”], [Following the flow of Fig. 1, using output from a knowledge base to generate dialogue acts (see dialogue act generation DAG) to be then used for response generation (RG) indicates accessing the knowledge base to generate a second output, i.e. response, in view of the plurality of responses of Table 3]), wherein the prompt is further generated to include at least a portion of the second output ([pgs. 14-15, Table 3, Turns 1-3], [The hotel is recommended based on price, rating, stars, etc. all of which could reasonably be considered to be portions of a second output, i.e. any step of the hotel recommendation]). Regarding claim 19, Lam discloses: a system comprising: one or more processors to: execute a dialog engine to manage an interplay between a large language model (LLM) and one or more user inputs ([pgs. 14-15, Table 3, Turns 1-3], [representing a multi-step interplay between a large language model (mBART) and user inputs, i.e. DST inputs]), the dialog engine dynamically generating a prompt for the LLM including one or more example dialog flows associated with one or more predefined user inputs that are within a threshold similarity to the one or more user inputs ([pg. 5, Section 4.2, Par. 2] In creating the RG training set, we first translate the source agent utterances to the target language and use LaBSE to remove pairs whose similarity score is below a threshold. We found a threshold of 0.8 to work best empirically. Higher thresholds would inadvertently filter correctly translated utterances. We construct the final training data by pairing aligned translated utterances that pass the filter with their corresponding translated agent dialogue acts [Pairing aligned utterances with agent dialogue acts indicates a prompt for action response for the LLM including a dialog flow, in view of the dialog flow of Table 3, wherein the predefined inputs, i.e. source utterances, are within a threshold similarity to one or more user inputs, i.e. translated versions. The examiner would like to note that though related to translations, the concepts of comparing words between languages to determine accuracies of translations is equivalent to the word comparison for task determination between tasks in the same language, i.e. both word embedding comparisons]), the one or more example dialog flows comprising one or more sequences of past and future dialog between a user and the LLM and configuring the LLM at runtime ([pg. 10, Ethical Considerations, Par. 2] …performing one run instead of averaging multiple runs…”, [Multiple and/or one run indicates a runtime operation. Further, in view of the previously cited dialog flow “(hotels search) request rating, request stars”, wherein each request defines a sequence of dialog to be performed between a user and outputs of language model, i.e. the requests, in later turns such as rating and stars in turn 2, wherein at the DST stage, the “(hotels search)” is defining a past dialog request to be clarified with the future “request location” and “request price level” operations]) to generate the output according to constraints defined in the one or more example dialog flows based on (1) a match between a canonical form and a user input defined in at least one dialog flow ([In view of the dialog flow being generated based upon the canonical form, which itself is based upon user input, it is unclear to the examiner how there would not inherently always be a match between the canonical form and user input defined in the dialog flow for generating output as the dialog flow is generated based on the canonical form, and, therefore, the user input]) and (2) a corresponding canonical form of an LLM output defined in the at least one dialog flow ([pg. 14, Appendix, Table 3, Turn 2, DAG prediction “request location/price_level”], [“request location” is the canonical form of the generated language model output “And what about location?” defined in the multi-turn dialog flow.]). Lam does not disclose: the language model being implemented as an artificial neural network; the one or more example dialog flows being specified in configuration information maintained separately from the language model; and, the constraints being applied at runtime to the language model, and training data sets used to train the language model excluding feedback related to the constraints such that the language model is not pre-trained with the constraints. Moon discloses: the language model being implemented as an artificial neural network ([Col. 23, Lines 50-60] The assistant system 140 may then select, by a conversational reasoning model, one or more candidate nodes from the knowledge graph corresponding to one or more candidate entities, respectively. Each candidate node may be selected based on the nodes corresponding to the initial entities, one or more dialog states associated with the query, and a context associated with the query, [Col. 24, Lines 54-66] FIG. 10 illustrates an example artificial neural network (“ANN”) 1000. In particular embodiments, an ANN may refer to a computational model comprising one or more nodes. Example ANN 1000 may comprise an input layer 1010, hidden layers 1020, 1030, 1040, and an output layer 1050. Each layer of the ANN 1000 may comprise one or more nodes, such as a node 1005 or a node 1015. In particular embodiments, each node of an ANN may be connected to another node of the ANN. As an example and not by way of limitation, each node of the input layer 1010 may be connected to one of more nodes of the hidden layer 1020. In particular embodiments, one or more nodes may be a bias node, [A conversational reasoning, i.e. language, model for selecting nodes, wherein an ANN is defined as the model performing the operations, indicates the conversational reasoning model to be implemented using the ANN]); the one or more example dialog flows being specified in configuration information maintained separately from the language model ([Fig. 1, Assistant System 140 connected to Assistant Application 136 via Network 110], [Col. 6, Lines 55-61] The assistant application 136 may communicate the user input to the assistant system 140. Based on the user input, the assistant system 140 may generate responses. The assistant system 140 may send the generated responses to the assistant application 136. The assistant application 136 may then present the responses to the user at the client system 130, [The examiner asserts that the assistant application tracks to a separate entity from the assistant system 140 containing natural language generator (NLG) 271 which is disclosed to be containing a language model ([Col. 22, Lines 20-25]). There is motivation to take the user input of Moon and transform it into the dialog flow format of Lam as Lam discloses generating dialog flows for current turns based on previous turns and based on user utterances ([Section 3.2, “3. Dialogue Act Generation”]), all disclosed functionalities of the assistant system 140 of Moon ([Col. 17, Lines 1-55]). The user input of Moon would be transformed using the dialogue act generation of Lam before generating responses, resulting in configuration information to be sent to the assistant system. Further, there is motivation to separate the functionalities of each step of Lam into distinct components as seen in Moon because splitting the stages of Lam into multiple networks/models reduces the computational load on each network/model (Lam discloses cost of virtual assistants in the introduction)]); and, the constraints being applied at runtime to the language model ([Col. 25, Lines 65-67]-[Col. 26, Lines 1-5] (2) a zero-shot learning model that leverages previous sentence, dialog, and KG contexts to re-rank candidates from pruned decoder graph output based on their relevance and path scores, which allows for generalizable and robust classification with a large number of candidate classes, [A zero-shot learning model indicates the constraints of a previous sentence/dialog/context are used for generating/ranking candidates without prior training]), and training data sets used to train the language model excluding feedback related to the constraints such that the language model is not pre-trained with the constraints ([Col. 28, Lines 45-50] The embodiments disclosed herein compute zero-shot relevance score in the KG embeddings space, thus allowing for robust prediction for KG entities and domains unseen during training as well, [An entity/domain unseen during training indicates there is no pre-training operation associated with the entity/domain, necessarily having constraints related to the entity/domain]). Lam and Moon are considered analogous art within conversational reasoning within knowledge bases. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam to incorporate the teachings of Moon, because of the novel way to associate walk paths of a knowledge graph with input contexts including dialog state, sentence, and initial entities mentioned in the conversation for ranking candidate entities using a zero-shot relevance learning model which results in more accurate and relevant entities in generated responses within multi-turn dialogs (Moon, [Col. 3, Lines 15-40]). Claim(s) 4, 6, 8, 14, 15, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lam in view of Moon, further in view of Hall et al. (US-20210406718-A1), hereinafter Hall. Regarding claim 4, Lam in view of Moon discloses: the method of claim 1. Lam further discloses: generating an embedding of the user input in a semantic or latent space ([pg. 5, section 4.2, Par. 1] To score a pair of sentences, the model first calculates an embedding for each sentence and computes the cosine distance between those vectors [Computing a distance between embedding vectors, wherein the vectors are encoded using BERT (indicating semantic analysis), indicates the vectors are embeddings in a semantic or latent space]). Lam in view of Moon does not disclose: wherein the generating the canonical form comprises: determining one or more example user inputs that are associated with one or more predefined canonical forms based at least on the embedding of the user input and one or more embeddings of the one or more example user inputs in the semantic or latent space; and, generating a prompt that includes the one or more example user inputs, the one or more predefined canonical forms, and at least a portion of a current conversation. Hall discloses: wherein the generating the canonical form comprises: determining one or more example user inputs that are associated with one or more predefined canonical forms based at least on the embedding of the user input and one or more embeddings of the one or more example user inputs in the semantic or latent space ([0022] As shown in FIG. 1C, in a new conversational event 105C, the user 110 may request that the conversational computing interface 100 “order another pepperoni pizza.” Accordingly, the conversational computing interface 100 may be configured to repeat the previously-executed action 106B in a replicated planned action 107B, e.g., so as to order another pepperoni pizza , and, [0017] the conversational computing interface may be trained to generate an action that is similar to an action that was taken for some other event in an annotated dialogue [In view of the embedding similarity comparison of Lam, Hall’s system determines an example input “Order a pepperoni pizza” (See Fig. 1B) with associated canonical form “order_pizza” (See Fig. 1B, 107B) based on similarity of requests of past orders (cheese pizza of training Fig. 1A and first pepperoni pizza of Fig. 1B]); generating a prompt that includes the one or more example user inputs, the one or more predefined canonical forms, and at least a portion of a current conversation ([Fig. 3A-3B, Annotated Dialogue History 300A, Updated Dialogue Plan 300B, traced/executable steps], [Updating a dialogue plan, i.e. prompt(s), based on a dialogue history, i.e. example user input(s), in view of the API invocations of the dialogue history “Create User Account”, “Get Menu”, “Add [type-of] pizza”, etc., i.e. predefined canonical forms, to represent a current conversation, i.e. cheese pizza instead of pepperoni, tracks to generating a new prompt featuring the example inputs and predefined canonical forms of the dialogue history with the current conversation, substituting pepperoni for cheese]); and, processing the prompt using the language model to generate the canonical form ([Fig. 3B, Conversation Event 312B], [Processing the prompts of “Which would you like?” and “Cheese!” in reference to kinds of pizza resulting in an API invocation “Add Cheese Pizza” indicates a generated canonical form, i.e. “Add Cheese Pizza”, based on processing of the prompt “Cheese!”, in view of the language model of Lam]). Lam are considered analogous art within task-oriented dialogue. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam in view of Moon to incorporate the teachings of Hall, because of the novel way to adjust previous dialogue conversations to a current situation allowing for the previous actions to be directly executable in a new context, reducing required training samples (Hall, [0001]). Regarding claim 6, Lam in view of Moon discloses: the method of claim 1. Lam in view of Moon does not disclose: wherein the determining the dialog flow comprises matching the canonical form to a predefined canonical form associated with the dialog flow. Hall discloses: wherein the determining the dialog flow comprises matching the canonical form to a predefined canonical form associated with the dialog flow ([Figs. 3A-3B], [0047] Updated dialogue plan 300B includes an updated conversational event 312B in which a user asks for a cheese pizza instead of a pepperoni pizza, and subsequent updated executable steps including updated executable step 314B asking for user confirmation and updated executable step 318B in which the conversational computing interface orders the cheese pizza instead of the pepperoni pizza…The updated dialogue plan 300B may be derived by automatically and/or manually editing the annotated dialogue history 300A [In view of Figs. 3A-3B, performing an updated dialogue plan that contains all the same steps as the dialogue history event indicates that based on the canonical form of “Order a pizza!”, in view of the canonical form generation for tasks of Lam, a dialogue flow is determined by matching canonical forms to predefined, i.e. historical, canonical forms in view of the embedding similarity matching of Lam]). Lam are considered analogous art within task-oriented dialogue. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam in view of Moon to incorporate the teachings of Hall, because of the novel way to adjust previous dialogue conversations to a current situation allowing for the previous actions to be directly executable in a new context, reducing required training samples (Hall, [0001]). Regarding claim 8, Lam in view of Moon discloses: the method of claim 1. Lam further discloses: generating an embedding of the canonical form in a semantic or latent space ([pg. 5, section 4.2, Par. 1] To score a pair of sentences, the model first calculates an embedding for each sentence and computes the cosine distance between those vectors [Computing a distance between embedding vectors, wherein the vectors are encoded using BERT (indicating semantic analysis), indicates the vectors are embeddings in a semantic or latent space]). Lam in view of Moon does not disclose: wherein the determining the dialog flow comprises: determining one or more canonical forms that are associated with one or more predefined dialog flows based at least on the embedding of the canonical form and one or more embeddings of the one or more canonical forms in the semantic or latent space; generating a prompt that includes the one or more canonical forms, the one or more predefined dialog flows, and at least a portion of a current conversation; and processing the prompt using the language model to generate the dialog flow. Hall discloses: wherein the determining the dialog flow comprises: determining one or more canonical forms that are associated with one or more predefined dialog flows based at least on the embedding of the canonical form and one or more embeddings of the one or more canonical forms in the semantic or latent space ([Figs. 3A-3B, Conversation Events 302A, 302B], [In view of the canonical form generation based on input of Lam, determining that an updated dialogue plan should have the same conversation pattern as a dialogue from history based on the original event “Order a pizza!” indicates a determination that the canonical forms of “Order a pizza!” or other conversational events, as would be generated in Lam, between the historical and current dialogues are associated based on canonical form embeddings, in view of the embeddings of Lam, associated with predefined dialogue flows, i.e. for ordering a pizza]); generating a prompt that includes the one or more canonical forms, the one or more predefined dialog flows, and at least a portion of a current conversation ([Fig. 3A-3B, Annotated Dialogue History 300A, Updated Dialogue Plan 300B, traced/executable steps], [Updating a dialogue plan, i.e. prompt(s), based on a dialogue history, i.e. predefined dialogue flow, in view of the API invocations of the dialogue history “Create User Account”, “Get Menu”, “Add [type-of] pizza”, etc., i.e. canonical forms, to represent a current conversation, i.e. cheese pizza instead of pepperoni, tracks to generating a new prompt featuring one or more canonical forms (as would be determined in Lam) and predefined dialogue flows of the dialogue history with the current conversation, i.e. for ordering a pizza with different toppings]); and processing the prompt using the language model to generate the dialog flow ([Fig. 3B, Conversation Event 312B], [Processing the prompts of “Cheese!” and “The total for your order is $8…” in reference to kinds of pizza resulting in an API invocation “Add Cheese Pizza” indicates a generated dialogue flow, i.e. “Add Cheese Pizza”, based on processing of the prompt “Cheese!”, in view of the language model of Lam and the previously defined prompts of cheese pizza and associated price]). Lam are considered analogous art within task-oriented dialogue. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam in view of Moon to incorporate the teachings of Hall, because of the novel way to adjust previous dialogue conversations to a current situation allowing for the previous actions to be directly executable in a new context, reducing required training samples (Hall, [0001]). Regarding claim 14, Lam in view of Moon discloses: the processor of claim 11. Lam further discloses: generating an embedding of the user input in a semantic or latent space ([pg. 5, section 4.2, Par. 1] To score a pair of sentences, the model first calculates an embedding for each sentence and computes the cosine distance between those vectors [Computing a distance between embedding vectors, wherein the vectors are encoded using BERT (indicating semantic analysis), indicates the vectors are embeddings in a semantic or latent space]). Lam in view of Moon does not disclose: wherein the generating the canonical form comprises: determining one or more example user inputs that are associated with one or more predefined canonical forms based at least on the embedding of the user input and one or more embeddings of the one or more example user inputs in the semantic or latent space; and, generating a prompt that includes the one or more example user inputs, the one or more predefined canonical forms, and at least a portion of a current conversation. Hall discloses: wherein the generating the canonical form comprises: determining one or more example user inputs that are associated with one or more predefined canonical forms based at least on the embedding of the user input and one or more embeddings of the one or more example user inputs in the semantic or latent space ([0022] As shown in FIG. 1C, in a new conversational event 105C, the user 110 may request that the conversational computing interface 100 “order another pepperoni pizza.” Accordingly, the conversational computing interface 100 may be configured to repeat the previously-executed action 106B in a replicated planned action 107B, e.g., so as to order another pepperoni pizza , and, [0017] the conversational computing interface may be trained to generate an action that is similar to an action that was taken for some other event in an annotated dialogue [In view of the embedding similarity comparison of Lam, Hall’s system determines an example input “Order a pepperoni pizza” (See Fig. 1B) with associated canonical form “order_pizza” (See Fig. 1B, 107B) based on similarity of requests of past orders (cheese pizza of training Fig. 1A and first pepperoni pizza of Fig. 1B]); generating a prompt that includes the one or more example user inputs, the one or more predefined canonical forms, and at least a portion of a current conversation ([Fig. 3A-3B, Annotated Dialogue History 300A, Updated Dialogue Plan 300B, traced/executable steps], [Updating a dialogue plan, i.e. prompt(s), based on a dialogue history, i.e. example user input(s), in view of the API invocations of the dialogue history “Create User Account”, “Get Menu”, “Add [type-of] pizza”, etc., i.e. predefined canonical forms, to represent a current conversation, i.e. cheese pizza instead of pepperoni, tracks to generating a new prompt featuring the example inputs and predefined canonical forms of the dialogue history with the current conversation, substituting pepperoni for cheese]); and, processing the prompt using the language model to generate the canonical form ([Fig. 3B, Conversation Event 312B], [Processing the prompts of “Which would you like?” and “Cheese!” in reference to kinds of pizza resulting in an API invocation “Add Cheese Pizza” indicates a generated canonical form, i.e. “Add Cheese Pizza”, based on processing of the prompt “Cheese!”, in view of the language model of Lam]). Lam are considered analogous art within task-oriented dialogue. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam in view of Moon to incorporate the teachings of Hall, because of the novel way to adjust previous dialogue (Hall, [0001]). Regarding claim 15, Lam in view of Moon discloses: the processor of claim 11. Lam further discloses: generating an embedding of the canonical form in a semantic or latent space ([pg. 5, section 4.2, Par. 1] To score a pair of sentences, the model first calculates an embedding for each sentence and computes the cosine distance between those vectors [Computing a distance between embedding vectors, wherein the vectors are encoded using BERT (indicating semantic analysis), indicates the vectors are embeddings in a semantic or latent space]). Lam in view of Moon does not disclose: wherein the determining the dialog flow comprises: determining one or more canonical forms that are associated with one or more predefined dialog flows based at least on the embedding of the canonical form and one or more embeddings of the one or more canonical forms in the semantic or latent space; generating a prompt that includes the one or more canonical forms, the one or more predefined dialog flows, and at least a portion of a current conversation; and processing the prompt using the language model to generate the dialog flow. Hall discloses: wherein the determining the dialog flow comprises: determining one or more canonical forms that are associated with one or more predefined dialog flows based at least on the embedding of the canonical form and one or more embeddings of the one or more canonical forms in the semantic or latent space ([Figs. 3A-3B, Conversation Events 302A, 302B], [In view of the canonical form generation based on input of Lam, determining that an updated dialogue plan should have the same conversation pattern as a dialogue from history based on the original event “Order a pizza!” indicates a determination that the canonical forms of “Order a pizza!” or other conversational events, as would be generated in Lam, between the historical and current dialogues are associated based on canonical form embeddings, in view of the embeddings of Lam, associated with predefined dialogue flows, i.e. for ordering a pizza]); generating a prompt that includes the one or more canonical forms, the one or more predefined dialog flows, and at least a portion of a current conversation ([Fig. 3A-3B, Annotated Dialogue History 300A, Updated Dialogue Plan 300B, traced/executable steps], [Updating a dialogue plan, i.e. prompt(s), based on a dialogue history, i.e. predefined dialogue flow, in view of the API invocations of the dialogue history “Create User Account”, “Get Menu”, “Add [type-of] pizza”, etc., i.e. canonical forms, to represent a current conversation, i.e. cheese pizza instead of pepperoni, tracks to generating a new prompt featuring one or more canonical forms (as would be determined in Lam) and predefined dialogue flows of the dialogue history with the current conversation, i.e. for ordering a pizza with different toppings]); and processing the prompt using the language model to generate the dialog flow ([Fig. 3B, Conversation Event 312B], [Processing the prompts of “Cheese!” and “The total for your order is $8…” in reference to kinds of pizza resulting in an API invocation “Add Cheese Pizza” indicates a generated dialogue flow, i.e. “Add Cheese Pizza”, based on processing of the prompt “Cheese!”, in view of the language model of Lam and the previously defined prompts of cheese pizza and associated price]). Lam are considered analogous art within task-oriented dialogue. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam in view of Moon to incorporate the teachings of Hall, because of the novel way to adjust previous dialogue conversations to a current situation allowing for the previous actions to be directly executable in a new context, reducing required training samples (Hall, [0001]). Regarding claim 18, Lam in view of Moon discloses: the processor of claim 11. Lam further discloses: wherein the processor is comprised in at least one of: a system incorporating one or more virtual machines (VMs) ([Section 5.2, Par. 4] Our models were trained on virtual machines with a single NVIDIA V100 (16GB memory) GPU on the AWS platform); and, a system implementing one or more large language models (LLMs) ([Section 5.2, Par. 1] We use mbart-large-50 as the neural model for our agent in all our experiments. [BART is a well-known language model, adapted to “large”]). Lam in view of Moon does not disclose: wherein the processor is comprised in at least one of: an infotainment system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for generating or presenting virtual reality, augmented reality, or mixed reality content; a system for performing conversational AI operations; a system for generating synthetic data; a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources. Hall discloses: a system for performing deep learning operations ([0084] Non-limiting examples of training procedures for adjusting trainable parameters include… reinforcement learning (e.g., deep Q learning based on feedback)); and, a system for generating or presenting virtual reality, augmented reality, or mixed reality content ([0079] In some implementations, display subsystem may include one or more virtual-, augmented-, or mixed reality displays [The examiner would like to note that due to the current disjunctive nature of the claims, not all elements require a mapping]). Lam are considered analogous art within task-oriented dialogue. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam in view of Moon to incorporate the teachings of Hall, because of the novel way to adjust previous dialogue conversations to a current situation allowing for the previous actions to be directly executable in a new context, reducing required training samples (Hall, [0001]). Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lam in view of Moon, further in view of Andreas et al. (US-11410643-B2), hereinafter Andreas. Regarding claim 20, Lam in view of Moon discloses: the system of claim 19. Lam in view of Moon does not disclose: wherein the prompt is dynamically generated based at least on the one or more user inputs being dissimilar from the one or more predefined user inputs by more than a threshold amount. Andreas discloses: wherein the prompt is dynamically generated based at least on the one or more user inputs being dissimilar from the one or more predefined user inputs by more than a threshold amount ([Col. 4, Lines 45-50] For example, the conversational computing interface may be configured to generate computer-executable plans for events that did not occur in the training data, so as to generalize the training from the exemplary annotated dialogues provided in training data to other, similar and/or dissimilar situations [In view of the similarity calculations of Lam, indicating a threshold similarity, in view of the similarities of Andreas, in which new prompts are dynamically generated based on dissimilar user inputs indicating a threshold in order to be considered “dissimilar”]). Lam are considered analogous art within task. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Lam in view of Moon to incorporate the teachings of Andreas, because of the novel way to adapt and ability to respond to diverse situations (Andreas, [Col. 4, Lines 30-50]). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Coope et al. (US-11204964-A1) discloses “A system comprising: an input configured to receive input speech data originating from a user; an output configured to output speech or text information; and a processor configured to: provide first input data to a character sequence determination module to determine a character sequence from the first input data, wherein determining a character sequence comprises: obtaining a first list of one or more candidate character sequences from the first input data; selecting a first candidate character sequence from the first list; generating a first confirm request to confirm the selected first candidate character sequence, wherein the first confirm request is outputted by way of the output; if second input data indicating that the first candidate character sequence is not confirmed is received, selecting a second candidate character sequence and generating a second confirm request to confirm the selected second candidate if the second candidate character sequence is different from the first candidate character sequence, wherein the second confirm request is outputted by way of the output; and if second input data indicating that the first candidate character sequence is confirmed is received, the one or more processors are further configured to: provide third input data to a dialogue module, wherein the dialogue module is configured to: determine, based on the third input data, a dialogue act that specifies speech or text information; and output, by way of the output, the speech or text information specified by the determined dialogue act.” (abstract). See entire document. Kawano et al. (“Neural Conversation Model Controllable by Given Dialogue Act Based on Adversarial Learning and Label-aware Objective”) discloses “Building a controllable neural conversation model (NCM) is an important task. In this paper, we focus on controlling the responses of NCMs by using dialogue act labels of responses as conditions. We introduce an adversarial learning framework for the task of generating conditional responses with a new objective to a discriminator, which explicitly distinguishes sentences by using labels. This change strongly encourages the generation of label-conditioned sentences” (abstract). Wang et al. (“Task-Oriented Dialogue System as Natural Language Generation”) discloses “we propose to formulate the task-oriented dialogue system as the purely natural language generation task, so as to fully leverage the large-scale pre-trained models like GPT-2 and simplify complicated delexicalization prepossessing. However, directly applying this method heavily suffers from the dialogue entity inconsistency caused by the removal of delexicalized tokens, as well as the catastrophic forgetting problem of the pre-trained model during fine-tuning, leading to unsatisfactory performance. To alleviate these problems, we design a novel GPT-Adapter-CopyNet network, which incorporates the lightweight adapter and Copy-Net modules into GPT-2 to achieve better performance on transfer learning and dialogue entity generation” (abstract). Any inquiry concerning this communication or earlier communications from the examiner should be directed to THEODORE JOHN WITHEY whose telephone number is (703)756-1754. The examiner can normally be reached Monday - Friday, 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached on (571) 272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /THEODORE WITHEY/Examiner, Art Unit 2655 /ANDREW C FLANDERS/Supervisory Patent Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Show 9 earlier events
Dec 09, 2025
Examiner Interview Summary
Dec 30, 2025
Request for Continued Examination
Jan 17, 2026
Response after Non-Final Action
Feb 13, 2026
Non-Final Rejection mailed — §103
May 06, 2026
Applicant Interview (Telephonic)
May 06, 2026
Examiner Interview Summary
May 11, 2026
Response Filed
Jul 21, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12657389
TECHNOLOGIES FOR ERROR REDUCTION IN INTENT CLASSIFICATION
3y 5m to grant Granted Jun 16, 2026
Patent 12646499
METHOD, DEVICE, AND COMPUTER PROGRAM PRODUCT FOR PROCESSING INFORMATION
3y 3m to grant Granted Jun 02, 2026
Patent 12632670
Natural Language Processing for Identifying Bias in a Span of Text
3y 2m to grant Granted May 19, 2026
Patent 12591744
METHOD FOR TRAINING SEMANTIC REPRESENTATION MODEL, DEVICE AND STORAGE MEDIUM
4y 0m to grant Granted Mar 31, 2026
Patent 12536994
APPARATUS FOR CLASSIFYING SOUNDS BASED ON NEURAL CODE IN SPIKING NEURAL NETWORK AND METHOD THEREOF
2y 9m to grant Granted Jan 27, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
41%
Grant Probability
87%
With Interview (+46.7%)
2y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 27 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month