Prosecution Insights
Last updated: October 02, 2026
Application No. 18/204,247

DIRECTED MANAGEMENT OF INTERACTIVE ELEMENTS IN AN INTERACTIVE ENVIRONMENT UTILIZING MACHINE LEARNING

Final Rejection §101§103§112
Filed
May 31, 2023
Priority
Feb 28, 2023 — provisional 63/448,950
Examiner
LU, HWEI-MIN
Art Unit
2142
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
63%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 63% of resolved cases
63%
Career Allowance Rate
152 granted / 240 resolved
+8.3% vs TC avg
Strong +40% interview lift
Without
With
+40.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
26 currently pending
Career history
264
Total Applications
across all art units

Statute-Specific Performance

§101
9.6%
-30.4% vs TC avg
§103
50.4%
+10.4% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
28.9%
-11.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 240 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment This office action is in response to the amendment filed on 06/11/2026. Claims 1-20 remain pending in the application. Claims 1, 9, and 19 are independent. Drawings Applicant's amendment to specification corrects some of previous objections; therefore, some of previous objections are withdrawn. Applicant's amendment to specification also raises new issues; therefore, the remaining objections are shown below. The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: 821 in FIG. 8 since 821 has been removed from ¶ [0114] of amended specification dated on 06/11/2026. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “730” has been used to designate both "PERIPHERAL DEVICE PORT" in FIG. 7 and "on-board camera" in ¶ [0110] of amended specification dated on 06/11/2026. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Specification Applicant's amendment to specification corrects previous objections; therefore, the previous objections are withdrawn. Claim Objections Applicant's amendment to claims corrects previous objections; therefore, the previous objections are withdrawn. Claim Rejections - 35 USC § 112 Applicant's amendment to claims corrects some of previous rejections; therefore, some of previous rejections are withdrawn. The remaining rejections are shown below. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2, 5, and 10 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 2 and 10 recites the limitation "… analyze/analyzing (…) the video game environment for a specific context based on the subsequent input … determine/determining a second intent objective based on one or more of the subsequent input, the specific context and …" in lines and lines 4-9 respectively, which rendering these claims indefinite because ". Claim 5 recites the limitation "… identify" in lines 6, which rendering these claims indefinite because ". Claim Rejections - 35 USC § 101 Applicant's amendment to claims corrects previous rejections; therefore, previous rejections are withdrawn. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-3, 7-11, and 15-20 are rejected under 35 U.S.C. 103 as being unpatentable over GOSLIN et al. (US 2021/0081498 A1, pub. date: 03/18/2021), hereinafter GOSLIN in view of O’Malia et al. (US 2022/0036153 A1, pub. date: 02/03/2022), hereinafter O’Malia. Independent Claims 1, 9, and 19 GOSLIN discloses a system (GOSLIN, ¶ [0034] with 115 in FIG. 3: AI System 115) comprising: at least one processor (GOSLIN, ¶ [0034] with 310 in FIG. 3: Processor 310); and memory storing instructions (GOSLIN, ¶ [0034] with 315 in FIG. 3: programming instructions stored in Memory 315) that, when executed by the at least one processor, cause the system to perform a set of operations (GOSLIN, ¶ [0034]: Processor 310 retrieves and executes programming instructions stored in Memory 315 as well as stores and retrieves application data residing in Storage 320; ¶ [0062]: the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions/acts specified in the block(s) of the flowchart illustrations or block diagrams), the set of operations comprising: receive an input, by a director service, to modify an interactive element of a video game environment rendered at one or more display devices; analyze, by the director service, the video game environment for a specific context based on the input (GOSLIN, ¶¶ [0014]-[0015] and [0018]-[0019]: the AI systems that utilize machine learning to interact with users in a dynamic and immersive manner; the AI system utilizes a collection of machine learning (ML) models, each trained on specified context or scope; the AI system determines the context of a given input; dynamically select ML models as the context of the interaction shifts, in order to continue to provide deep conversation; the AI system acts as an intelligent character in a role-playing game; the AI system infers the context of the conversation based on input as the user interacts with the character; the AI system uses natural language processing (NLP) and/or natural language understanding (NLU) to attempt to identify role-playing scenario the user is partaking in; the AI system repeatedly determines the context for each input (which may include analyzing prior input), such that the AI system can respond to shifting contexts (e.g., if the user switches from a stealth methodology to a brute-force methodology); the input comprises natural language, and may include text and/or audio input; ¶¶ [0023]-[0025] and [0027] with FIG. 1: a User 105 provides Input 110 to the AI System 115; the Input 110 include natural language, and may include text (e.g., typed by the User 105) or audio speech data (e.g., recorded by a microphone); if the Input 110 is audio, the AI System 115 utilizes speech-to-text techniques, and processes the resulting text; the AI System 115 determines the context of the Input 110 based at least in part on prior user selection; the AI System 115 uses NLP and/or NLU (e.g., performed on audio, text, and/or combinations thereof) to determine some or all of the context of the Input 110; e.g., the user may select an objective, and the AI System 115 uses NLP to identify the means the user is pursuing, and/or to infer the character or role the user is playing; the User 105 may provide additional Input 110; the User 105 and AI System 115 can interact during the role-playing scenario until the User 105 quits, or until predefined criteria are met; ¶¶ [0028]-[0030] with FIG. 2: the user interacts with a scenario creator to select one or more aspects of the role-playing scenario they wish to use; the options available for a given selection can depend on one or more other selections; i.e., the selections for each of the Objectives 210, Roles 220, and/or Means 230 may have predefined relationships defining combinations that can be selected; ¶ [0042] with FIG. 3: the Context Component 340 may determine the context based at least in part on the original input provided by the user (e.g., the explicit selection(s) the user made in initiating the scenario); ¶ [0049] with 405 and 410 in FIG. 4 and FIG. 3: at block 405, where an Interactivity Application 330 receives user input; this input may include textual input, audio input, and the like; at block 410, the Interactivity Application 330 evaluates the input to determine the current context; ¶ [0057] with 605 and 610 in FIG. 6 and FIG. 3: at block 605, the Interactivity Application 330 receives a first input to an artificial intelligence (AI) system; at block 610, the Interactivity Application 330 determines a first context of the first input, wherein the first context indicates a first role-playing scenario); receive one or more environment guidelines, wherein the one or more environment guidelines include conditional prompt generation logic; associate, by the director service, the input with the one or more environment guidelines (GOSLIN, ¶¶ [0015]-[0018]: a roleplaying scenario is an interactive session that includes a role for the user to play, an objective for the user to pursue, and/or a means that the user should use to achieve the objective; the user plays a more specific role, such as a particular character from a movie or television show; the user is expected to understand the personality of the selected role, and must "stay in character" (e.g., by using an appropriate means) to achieve the objective; each role of the user is related or correlated to an expected or anticipated means or approach; e.g., an "intellect" means may correspond to a "wizard" character, while an "intimidation" means corresponds to a "knight" character; the AI system responds in part based on whether the determined means the user is relying on aligns with the expected means for the determined role the user is playing (i.e., conditional prompt generation logic); in order to respond based on the alignment between the expected means and the actual means the user is utilizing, the AI system uses one or more fitness functions to ensure the output makes logical sense; the AI system similarly uses a fitness function to ensure that the output is authentic to a particular character (e.g., the character that the AI system is playing); this system can also be used to evaluate input from the user; e.g., the fitness function can evaluate each line corresponding to the character from the movie, show, or other predefined script, and then evaluate the user input to determine a mathematical distance between the user's word choice and phrasing, as compared to the "canonical" phrases; based on this distance, the AI system can determine whether the user is accurately playing the role (i.e., determining whether the objective has been achieved); the AI system may search for keywords in user input that have predefined associations with a role, means, and/or objective; some or all of the context may be known or provided to the AI system; e.g., the user can select one or more aspects of the scenario (e.g., the objective, the role, and/or the means), and the AI system can infer the remaining aspects; the AI system repeatedly determines the context for each input (which may include analyzing prior input), such that the AI system can respond to shifting contexts; ¶¶ [0002], [0023]-[0024] and [0035]-[0038] with FIG. 1: conversational agents, such as chat bots, can provide rudimentary responses to users, but typically rely on limited decision-trees with restricted scripts; further, existing chat agents have limited scope and topics of discussion; e.g., a help desk bot may respond well to IT requests, but is useless for questions about commercial products; similarly, a bot cannot be used to respond to generic requests without sacrificing the specificity required to engage in useful discussions that are more narrowly-focused; moreover, existing chat agents fail to provide immersive interactivity, as their limited responses and narrow scope of useful discussion are severely limiting; the AI System 115 includes a set of ML Models 120A-N; each ML Models 120 corresponds to a finite state machine, a decision tree, and the like; each ML Model 120 corresponds to a machine learning model that was trained to receive textual input and generate or select a corresponding output; this output includes dynamically generated text or audio; each ML Model 120A-N was trained based on training data corresponding to a given context (e.g., a specific combination of role, means, and objective); the training data includes samples of input text and a corresponding output text for each sample of input text; each ML Model 120 includes a set of learned weights or other parameters that were learned during a training phase, based on labeled training data; e.g., during training, each ML Model 120 can be refined by adjusting weighted paths in a decision tree, tweaking a genetic algorithm, modifying a rules-based system, and the like; the training data is collected from roleplaying sessions between users (e.g., as part of a development team), and/or from early test users; the input from each user can be used as either exemplar input or target output, depending on the particular model being trained ( e.g., depending on which role the AI system will be playing, and which role the user will be playing); a separate set of training data is used for each ML Model 120, where each set of training data is associated with a respective context (e.g., a role-playing scenario); each ML Model 120 is labeled or otherwise associated with the context that corresponds to the underlying training data; the AI System 115 can selectively use each ML Model 120 based on their specific contexts (e.g., based on matching the context of the current input with the ML Models 120) to generate deeper and more specific responses within each context); determine, by the director service, an intent objective based on one or more of the input, the specific context and the one or more environment guidelines; generate, by the director service, a prompt using a generative machine learning model based on the intent objective, the input, and the conditional prompt generation logic (GOSLIN, ¶¶ [0015]-[0019] and [0022]: the AI system can utilize techniques including keyword identification, sentiment analysis, intent evaluation, parsing to determine meaning, and the like; the AI system responds in part based on whether the determined means the user is relying on aligns with the expected means for the determined role the user is playing (i.e., conditional prompt generation logic); each objective corresponds to a suggested or best means or role; in order to respond based on the alignment between the expected means and the actual means the user is utilizing, the AI system uses one or more fitness functions to ensure the output makes logical sense; the AI system similarly uses a fitness function to ensure that the output is authentic to a particular character (e.g., the character that the AI system is playing); the AI system utilizes NLP and/or NLU to determine the objective, role, and/or means the user is using for the role-playing scenario; once the context of a given input is determined, the AI system selects a corresponding ML model to process the input; the AI system includes a respective ML model trained for each roleplaying scenario (e.g., trained for each combination of objective, role, and means); once the scenario is identified, the AI system can select the corresponding ML model for evaluating the input; this can allow each model to be specialized with a constrained context, while allowing the AI system to dynamically shift between contexts by selecting other models; the AI system maintains context-specific weights for each ML model; given an input context, the AI system can probabilistically select a ML model to use based on the context-specific weight associated with each; the AI system can modify these context-specific weights based on user feedback, in order to better select ML models for future interactions; ¶¶ [0023]-[0025] and [0035]-[0037] with FIG. 1: the AI System 115 includes a set of ML Models 120A-N; each ML Models 120 corresponds to a finite state machine, a decision tree, and the like; each ML Model 120 corresponds to a machine learning model that was trained to receive textual input and generate or select a corresponding output; this output includes dynamically generated text or audio; each ML Model 120A-N was trained based on training data corresponding to a given context (e.g., a specific combination of role, means, and objective); the AI System 115 determines the context of the Input 110, in order to select an appropriate ML Model 120A-N; the training data includes samples of input text and a corresponding output text for each sample of input text; the AI System 115 uses NLP and/or NLU (e.g., performed on audio, text, and/or combinations thereof) to determine some or all of the context of the Input 110; e.g., the user may select an objective, and the AI System 115 uses NLP to identify the means the user is pursuing, and/or to infer the character or role the user is playing; once the context is determined, the AI System 115 selects the corresponding ML Model 120; the AI System 115 may rely on weights associated with each ML Model 120 in determining which model to select; the AI System 115 may select one aspect of the context to change, and select the corresponding ML Model 120 for this changed aspect; each ML Model 120 includes a set of learned weights or other parameters that were learned during a training phase, based on labeled training data; e.g., during training, each ML Model 120 can be refined by adjusting weighted paths in a decision tree, tweaking a genetic algorithm, modifying a rules-based system, and the like; ¶¶ [0040]-[0046] with FIG. 3: the NLP Component 335 performs a variety of NLP processing such as sentiment analysis, keyword detection, parsing to determine intent or meaning, and the like; once the NLP Component 335 has determined the intent and/or sentiment, identified keywords, or performed any other NLP processing, some or all of the resulting data is passed as input to the ML Models 120; the Context Component 340 determines the current context of the user input in order to facilitate selection of an appropriate ML Model 120; the Context Component 340 infers the user's objective, role, and/or means based on identified keywords, or based on other results from the NLP Component 335; e.g., the Context Component 340 can determine that the objective is to retrieve an item, based on determining that predefined keywords relating to the item were included in one or more inputs received from the user; similarly, the Context Component 340 may determine that the user is roleplaying as a particular character or is using a particular means, based on keywords, and/or based on the sentiment or intent of the input(s), as identified by the NLP Component 335; the Context Component 340 can determine whether the user has shifted the role-playing scenario (such as by deciding to attempt a different means to achieve the objective) (i.e., conditional prompt generation logic); the Context Component 340 may identify this transition and select a different ML Model 120 (e.g., based on keywords in the current input, based on sentiment analysis on the input, based on intent analysis, and the like); in one embodiment, the Context Component 340 can allow this change and the role-playing scenario will be shifted accordingly (e.g., by selecting other ML Models 120); in another embodiment, the Context Component 340 can continue to output the previous context; e.g., the Context Component 340 can determine that the user is switching roles, but may nevertheless indicate, to the ML Component 345, that the role remains the same (e.g., because predefined rules indicate that the role cannot be changed mid-scenario); once the Context Component 340 has determined the current context (e.g., the objective, role, and means), the Context Component 340 provides an indication to the ML Component 345 reflecting these determinations; the ML Component 345 receives the current context from the Context Component 340, and selects one or more of the ML Models 120 based on this context; the ML Component 345 identifies and selects the ML Model 120 that is associated with a matching context; the ML Component 345 may select one or more other ML Models 120 (e.g., periodically, or in a probabilistic manner); e.g., the ML Models 120 are associated with context-specific weights, indicating a likelihood that each will be selected given a particular context; given a first context C, a first ML Model 120 (e.g., one trained on the same context C) may be associated with a relatively high weight, such that it will be selected frequently; similarly, a second ML Model 120 (trained on a different context) can be associated with a relatively lower weight, such that it is selected less frequently than the first; the weight of each ML Model 120 is determined based in part on the vector distance between the current context C and the respective context C' of each respective ML Model 120; the ML Component 345 can dynamically modify the context-specific weights of each ML Model 120 during use; ¶¶ [0050]-[0051] with 415 in FIG. 4 and FIG. 3: block 415, where the Interactivity Application 330 identifies and selects one or more ML Models 120 based on the determined context; the Interactivity Application 330 selects the ML Model 120 with a matching context; the Interactivity Application 330 probabilistically selects a model based on the determined context, the confidence in this determination, and the context-specific weights associated with each ML Model 120; the Interactivity Application 330 further refines the context-specific weights of each ML Model 120, and/or the internal weights of the previously-selected model, based on the current input; the Interactivity Application 330 performs one or more NLP operations on the input (e.g., keyword identification, sentiment analysis, intent determination, and the like), and processes the result with the ML Model 120; ¶ [0057] with 615 in FIG.6 and FIG. 3: block 615, where the Interactivity Application 330 selects a first ML model of the plurality of ML models based on the determined first context, wherein the first ML model was trained based at least in part on the first role-playing scenario); execute the generative machine learning model (GOSLIN, ¶ [0014]: the AI system generates responses using one or more ML models that correspond to that context; ¶¶ [0023] and [0026] with FIG. 1: each ML Model 120 corresponds to a machine learning model that was trained to receive textual input and generate or select a corresponding output; this output includes dynamically generated text or audio; the output corresponds to text or audio that is selected from a predefined set of responses; once the AI System 115 has selected a ML Model 120, the Input 110 is provided as input in order to generate a corresponding Response 125; the selected ML Model 120 dynamically generates output text based on weights learned during the training phase; the ML Model 120 acts as a classifier to classify the input, and uses corresponding predefined text for that category as the output; ¶¶ [0038] and [0046]-[0047] with FIG. 3: use each ML Model 120 based on their specific contexts (e.g., based on matching the context of the current input with the ML Models 120) to generate deeper and more specific responses within each context; using a particular ML Model 120 to generate a response; in addition to receiving the current context from the Context Component 340, the ML Component 345 receives results of the NLP analysis from the NLP Component 335; once an ML Model 120 has been selected, the ML Component 345 provides this input to the selected model in order to generate an output; ¶ [0051] with 420 in FIG. 4 and FIG. 3: at block 420, the Interactivity Application 330 generates a response using the identified and selected ML Model 120; the ML Models 120 are trained to dynamically generate a response; ¶ [0057] with 620 in FIG.6 and FIG. 3: at block 620, the Interactivity Application 330 generates a first output by processing the first input using the first ML model.); at the director service, determine that the model output is responsive to the input and the one or more environment guidelines (GOSLIN, ¶ [0020]: the AI system may periodically or randomly use a different ML model for a given context, and evaluate the user's response; i.e., for a given a first context C, the AI system may determine to use an ML model trained on context C' to generate the response; the AI system can then analyze how the user responds, in order to determine whether the ML model corresponding to context C' should be used at least occasionally in the future, given context C; e.g. the system may determine whether the user liked the response, or appeared confused or frustrated; this evaluation can include receiving user feedback (e.g., a thumbs up or thumbs down, a score, and the like) and/or using NLP and/or NLU to analyze the user's response (e.g., to determine the sentiment); ¶ [0025]: if user's rate the "flattery" response higher for entertainment value, the system can learn to use this model more often, given the context; ¶ [0033]: the system further utilizes techniques such as NLP to perform semantic analysis in order to determine whether the user enjoyed the scenario; ¶¶ [0040]-[0041] with FIG.3: the NLP Component 335 may perform sentiment analysis to determine whether the user is enjoying the ongoing scenario; the ML Models 120 can be refined based on this analysis; evaluating results output by the NLP Component 335, which may include data relating to the most recent input, as well as from one or more prior inputs; ¶¶ [0046]-[0048] with FIG. 3: after using a particular ML Model 120 to generate a response, the ML Component 345 evaluates the user's next input in order to refine the context-specific weight(s) associated with the ML Model 120; e.g., if the user-response is positive (e.g., with a positive sentiment evaluation from the NLP Component 335), the ML Component 345 may increase the weight of the previously-selected ML Model 120, in order to increase the probability that it will be selected in the future, given the same context; similarly, if the user's subsequent response is negative, the ML Component 345 may reduce the weight of the previously-selected ML Model 120; the ML Model 120 is trained as a classifier that receives input (e.g., text, the results of one or more NLP processes, and the like) and select from a large set of predefined responses; these responses may similarly include textual responses and/or pre-recorded audio; in order to determine the quality of a given output, the ML Component 345 can evaluate explicit ratings from the user, subsequent responses from the user, facial expressions or other non-verbal emotional cues (e.g., laughing) from the user, and the like; if the subsequent user input is positive, the ML Component 345 can refine the model to increase the probability that the previous output will be selected again, given the same input and/or context; similarly, if the subsequent input is negative, the ML Component 345 reduces the probability that the ML Model 120 will select the same response again, given the previous input/context; ¶ [0051]: the models are trained to classify the input in order to select from a predefined set of responses); in response to determining that the model output is responsive to the input and the one or more environment guidelines, modify the interactive element of the video game environment based on the model output; and render the modified interactive element within the video game environment at the one or more display devices (GOSLIN, ¶ [0021]: in order to improve responses and user-engagement, the learning system can provide differing responses to different users; the AI system can sometimes reach better solution by randomly (or pseudo-randomly) changing the order of things, prioritizing something that was previously lower priority, and the like; this randomness enables the AI system to continue searching for better solutions; e.g., the system may be at a local maxima for a solution, but there can be one or more better global maxima available which can be discovered through occasional random selections; ¶ [0027] with FIG. 1: these Responses 125 are provided to the User 105; ¶¶ [0047]-[0048] with FIG. 3: the ML Model 120 dynamically generates an output, which may be textual, audio, text converted to audio using text-to-speech models, and the like; this output is then provided to the user; based on the subsequent user input, the ML Component 345 can modify or refine the selected ML Model 120; the Interactivity Application 330 can provide immersive role-playing to the user; ¶¶ [0051]-[0052] with 425 in FIG.4 and FIG. 3: at block 425, the response is returned to the user (e.g., by displaying it on a screen, or by outputting audio); this process can then be repeated until the user exits the scenario, or the scenario otherwise terminates; each objective has one or more predefined termination points; these termination points are associated with particular responses that may be selected by the ML Model 120; e.g., if a particular response is generated by the model, the Interactivity Application 330 may determine that the scenario has ended in success or failure, and end the role-playing interaction after the response is output) (GOSLIN, ¶¶ [0031]-[0033] with FIG. 2: the GUI 200 further includes a Button 235 to generate suggested scenarios based on the user's previous interactions with the system; the suggested scenario is based in part on the number of times a given selection has been used by the user, and/or how recently the selection has occurred; the suggestion is based in part on the length of time the user spends in each scenario; as the user interacts with the system in the scenario, the system maintains data about the interaction, such as the length of time the scenario lasts (e.g., until the user succeeds, fails, or ends the scenario), the result of the scenario (e.g., whether the user succeeded, failed, or quit), and the like; ¶¶ [0051]-[0056] with FIGS. 3-5: the Interactivity Application 330 generates a response using the identified and selected ML Model 120; the response is returned to the user (e.g., by displaying it on a screen, or by outputting audio); this process can then be repeated until the user exits the scenario, or the scenario otherwise terminates; if a particular response is generated by the model, the Interactivity Application 330 may determine that the scenario has ended in success or failure, and end the role-playing interaction after the response is output; providing dynamic interactions using machine learning; at block 505, where an Interactivity Application 330 identifies the user for which the suggestion should be tailored (e.g., the user making the request); at block 510, the Interactivity Application 330 then retrieves a profile of the identified user; at block 515, the Interactivity Application 330 evaluates the data contained in the user profile to evaluate the roles the user previously played in prior interactions; at blocks 520 and 525, the Interactivity Application 330 performs similar evaluations of the user's prior objectives and means, respectively; the Interactivity Application 330 further determines, for each such prior scenario, a level of satisfaction the user experienced; this may be determined based on user selection (e.g., indicating a rating at the end of the interaction), and/or inferred based on sentiment analysis of the user input during and/or after the interaction; the Interactivity Application 330 evaluates other user profiles and/or recent interactions from other users in order to identify role(s), objective(s), and/or mean(s) have been recently used by others or that are popular among other users; at block 530, the Interactivity Application 330 generates and suggests one or more new scenarios to the user, based on the above analysis; the generated suggestions include scenarios that closely match with the user's preferences inferred by the Interactivity Application 330 based on the number of times a particular selection was used, the duration of interactivity with the selection, the average satisfaction when the selection was used, how recently the selection was used, and the like). GOSLIN further discloses a method comprising the set of operations described above, wherein the interactive environment is the gaming environment (GOSLIN, ¶ [0015]: the AI system acts as an intelligent character in a role-playing game; a roleplaying scenario is an interactive session that includes a role for the user to play, an objective for the user to pursue, and/or a means that the user should use to achieve the objective; the role can include any character, such as a child, a strong warrior, a wise wizard, a stealthy archer, and the like; similarly, objectives can include any goal, such as infiltrating a secure area, questioning or interrogating a character for information, retrieving items, rescuing characters, and the like; further, the means can similarly include any methodology of achieving the objective, such as stealth, force, intimidation, flattery, diversion, distraction, and the like) GOSLIN further discloses a computer storage media including instructions (GOSLIN, ¶¶ [0034]-[0035] with 315 and 320 in FIG. 3: programming instructions stored in Memory 315; application data residing in Storage 320; the Storage 320 includes one or more ML Models 120, as well as User Profiles 350), which when executed by a processor (GOSLIN, ¶ [0034] with 310 in FIG. 3: Processor 310 retrieves and executes programming instructions stored in Memory 315 as well as stores and retrieves application data residing in Storage 320), cause the processor to perform the set of operations described above (GOSLIN, ¶ [0063]: these computer program instructions may also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other device to function in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function/act specified in the block(s) of the flowchart illustrations or block diagrams). GOSLIN fails to explicitly disclose to generate a prompt for a generative machine learning model based on the intent objective, the input, and the conditional prompt generation logic; and execute the generative machine learning model with the prompt to produce a model output. O’Malia teaches a system and a method relating to artificial intelligence (AI) Agents (O’Malia, ¶ [0002]), wherein generate a prompt for a generative machine learning model based on the intent objective, the input, and the conditional prompt generation logic; and execute the generative machine learning model with the prompt to produce a model output (O’Malia, ABSTRACT: visual data and/or text data are received from the artificial intelligence agent and/or an environment of the artificial intelligence agent; a text prompt is generated based on the visual information and/or the text data; the text prompt is provided to an ultra-large language model; text output of the ultra large language model is received in response to the text prompt; the artificial intelligence agent is supplied with the text output of the ultra-large language model and/or the text output converted into an alternative format; the artificial intelligence agent is configured to select an action, a series of actions, and/or the policy based on the state of an environment of the artificial intelligence agent and on the text output of the ultra-large language model and/or the text output converted into the alternative format; ¶¶ [0003]-[0010]: the Agent typically evaluates each action and / or observation in the context of the variables the Agent may observe within the environment , and in the context of goals which the agent may perceive or have in the environment; the Agent per forms this evaluation in an attempt to learn associations between the actions, observations, and/or goals, to build knowledge about the environment, and to develop increasingly successful strategies, actions, or policies for a given environment state; the environment may include any component in which, or to which, the Al Agent may carry out actions and/or policies selected by the Al Agent; the environment may include, e.g., a video game, a robot, a drone, a vehicle, an aircraft, a watercraft, and any other apparatus and/or software component; when the Al Agent acts in a given environment, whether in a simulation, the real-world, or a game, the Agent may be at a disadvantage in comparison with human players because the Agent does not have the human ability to (1) reference facts about the observable objects in an environment, and (2) generalize from any of the following: (i) past experience, (ii) the context provided by the environment about the observable objects' likely characteristics, and (iii) how the observable objects may impact goals to be achieved and/or the means to achieve the goals; this human "commonsense reasoning” capability is different from factual knowledge in that this human capability is rooted in generalizable mental models adapted to context and the human actor's understanding of narrative and context, rather than based on static knowledge; humans gain such contextual knowledge continually via experience across a vast range of situations, may build abstractive mental models of the relevance of such past knowledge to novel situations, and may generalize across scenarios very fluently; e.g., in a murder mystery game, a bloody knife is probably a useful and desirable object — a clue — whereas in another game, it is more likely to injure or hurt the player and is best avoided; this is not "knowledge" but inference based on past experience and the context that is presented to the Agent; such associative capability enables humans to immediately identify likely threats and goals in the environment based on context, and make good choices and rapid progress (on balance) as a result; an Agent may use attention over a static, external semantic knowledge bases pertaining to the objects in its environment and their relationships in order to self-guide its action selection; it may be advantageous to the learning and progression of Al Agents over time in some environments to have the capability to gain access to information that humans use; ¶¶ [0016]-[0031]: enable information from the Al Agent to be converted into a format that enables the Al Agent Controller to use an Ultra-Large-Language-Model (ULLM) as an engine to process data from the Al Agent and/or the Al Agent’s environment, and convert outputs of the ULLM into a format usable by the Al Agent in order to inform the Al Agent's actions within the environment; surprisingly , this processing may include generalizing past scenarios to new contexts and environments, and/or attributing value to certain goals or actions, and providing guidance or signals or other forms of input to the Al Agent pertaining to the Agent's environment and the choices and actions which may be advantageous to the Agent in that environmental context; advances in ULLM began to not only model language on the word level, but successfully model and capture the structure and abstractive capability of human language on a higher level; this novel capability to replicate some of the abstractive capability of human writing enables the use of such models in combination with environment and goal observations to make suggestions which provide the same associative advantages that humans may use when they interact with such environments; the Al Agent Controller may use models in the BERT family of attention/transformer models or other natural language processing models which capture a reflection of human thought, knowledge, associations, and abstractions and generalizations to provide input to the Al Agent in a way that it may affect aspects of the Al Agent's behavior via recommendations regarding the action selection process, salient goals, features, or other attributes, features, or factors in an environment which may influence the behaviors of the Al Agent; enable the functionality of the novel methods and systems provided herein, which rely on the model's capability of combining certain mental models and thought templates commonly used by people, and successfully adapting them to novel scenarios and data sets; capable of producing extended passages of creative writing such as could be written by a human, based on a very short and simple prompt ("left-to-right" text generation task); through bi-directional information conversion between Al Agent environment and Al Agent Controller/ULLM, enable the Al Agent to use information which may be encoded in the ULLM; process information regarding the environment and the information in the context of past experience and mental models derived from, or encoded in, the vast volumes of text used to train the ULLMs; the Al Agent Controller may return to the Al Agent, via novel conversion or translation methods, information or guidance, or reward signal (s) regarding elements of the environment, components of the environment, goals, actions, or any combinations thereof which may be relevant or important for any positive or negative reasons; the Al Agent Controller may return to the Al Agent any other influence or guidance which enables the Al Agent to obtain similar performance benefits that a human may otherwise have had based on the human's use of past knowledge and its generalized application to a given environment / environment state and the goals, objects, relationships, actions and / or other factors which may exist within the environment of the Al Agent; the Al Agent accessing the Al Agent Controller for assistance with processing the Al Agent’s environment may help to shape the Al Agent's performance selection; the assistance provided by the Al Agent Controller may leverage the extensive training of ULLMs on human mental models, general and generalized knowledge, and associations between components or actions , and their potential application(s) to the Al Agent's environment state; the Al Agent Controller may even be trained or customized for specific domains, in order to enhance the Al Agent Controller predictive power and usefulness to the Al Agent; the Al Agent performs a series of actions , whether random or directed, and gathers feedback on the effectiveness of the actions in attaining a given reward or goal; ¶¶ [0033]-[0061] with FIG. 1: the Al Agent 112 may query the Al Agent Controller 102 which uses the ULLM 114 as an abstraction or generalization engine; specifically , the Al Agent 112 may use the ULLM 114 to generalize past experiences to the context of current and/or or past environment data of the Al Agent 112 or to provide additional context, information, or suggestions as to what action or policy might be most appropriate; the action or policy may be deemed most appropriate based on information or representations — visual or in text or other form — that the Al Agent 112 provides to the Al Agent Controller 102 about current and past observations and action space of the Al Agent 112, as well as perceived goals of the Al Agent 112; because most forms of information in environments of the Al Agent 12 may be visual, a method is implemented in which the environment information of the Al Agent 112 is converted into a format which may be processed by the Al Agent Controller 102, and via which the outputs of the Al Agent Controller 102 may be converted into a format which may be understood by the Al Agent; this may include methods to optimize the prompting, or structured querying, of the Al Agent Controller 102 to elicit certain types of responses which may provide particular value to the Al Agent 112; the Al Agent 112 may share a representation of the environment, or describe an element, or indicate a relationship between elements in an environment, or outline the components of its environment, and provide the information to the Al Agent Controller 102; the Al Agent Controller 102 may generalize the provided information to other, related scenarios or situations so as to propose actions, goals, and/or sub-goals based on the past experience of the Al Agent Controller 102; the Al Agent Controller 102 may use multiple aspects of the inputs or prompts to interpret the context of the environment state of the Al Agent 112 such that the Al Agent Controller 102 may generate a relevant output to provide to the Al Agent 112; by incorporating contextual information, the Al Agent Controller 102 may increase the likelihood of the Al Agent 112 selecting an appropriate or useful action and/or recognizing which elements of the environment of the Al Agent 112 may be important to consider when making a decision; the Priming Module 110 may be configured to convert visual data such as a digital image to a natural language description of the image, objects in the image, and/or action (s) occurring in the image or series of images; the Visual/Natural Language Mapping or captioning Module 108 may use methods with explicit relational and geometric reasoning components such as Image Captioning: Transforming Objects into Words, 2020; the Priming Module 110 is configured to use the ULLM 114; ULLMs are generally "prompted" or "primed" with text (such as with "masked" and "left-to-right" text generation tasks ), and the ULLM is then asked to produce text compatible with the prompt; the Priming Module 110 is configured to generate a text prompt for the ULLM 114 and to provide the text prompt to the ULLM 114; these generated prompts tend to benefit from specific and structured requests, in the sense that such specific requests tend to generate more reliably structured, relevant, and interpretable outputs; the Priming Module 110 is configured to receive a text output from the ULLM 114 in response to the supplied text prompt; the Priming Module 110 may be optimized to evaluate the relative value and success of the outputs from the ULLM 114 in terms of enhancing the performance of the Al Agent 112 in order to build a repository of proven mental models or thought templates to which the ULLM 114 responds in a reliable , accurate , and structured way; the Priming Module 110 may include a discriminator 118 which may be included in a generative adversarial network (GAN) included in the Priming Module 110' the Priming Module 110 may include any other type of reinforcement learning structure and/or an imitation learning structure to evaluate the relative value and success of the outputs from the ULLM 114 in terms of enhancing the performance of the Al Agent 112; the discriminator 118 and/or other learning structure may include a neural network which evaluates whether the output of the ULLM 114 is sufficiently relevant to the inputs of the Al Agent 112 and the environment state of the Al Agent 112 that the output of the ULLM 114 might provide value to the Al Agent 112; the Priming Module 110 may store and utilize Shared Representations 120, where Shared Representations 120 are models and templates that, when combined with a given data type or input type from the environment and/or the Al Agent 112, may reliably trigger the application or use of a given thought template, mental model, value system , or similar high-level logical framework by the Al Agent Controller 102; one example class of Shared Representations 120 may be: list completion of things which belong together; such representations may be activated by the Priming Module 110 and passed as a text prompt to the ULLM 114 in certain situations; when the Al Agent 112 has or sees a variety of objects in an inventory or on the screen, which may be presented to the ULLM 114 as a list for completion or matching; in such a case, the Priming Module 110 may pass the list to the ULLM 114 for separation or sorting and add a predetermined prompt for "things which belong"; the ULLM 114 may determine the pattern or class that coherently represents some or all of the objects and suggest outputs which continue the pattern; one example is to prompt the ULLM 114 with the first three colors of the rainbow "Red , Orange , Yellow ..." , and the model of the ULLM 114 would typically pick up on the nature of the pattern as a shared representation of "things which belong" and map this to color order in a rainbow — although the ULLM 114 may also make other associations; the expected return from the ULLM 114 in this example may be "Green , Blue , Indigo , Violet"; such information may favorably inform the action selection of the Al Agent 112 in a game without the Al Agent 112 having such direct knowledge or models; in Montezuma's Revenge, a snapshot of the environment may be conveyed to the Priming Module 110 via the Visual/Natural Language Map ping Module 108 which may then list likely semantics of the environment and positions of objects on the screen, which when provided to the ULLM 114, may result in suggestions via an interface overlay such as a heatmap which indicates the importance of avoiding the skull and the benefit of acquiring the key; e.g., a Heat Map Generator 126 included in the conversion module 124 may generate the heatmap from text returned by the ULLM 114; the Al Agent 112 may include a neural-network, artificially-intelligent, or deep learning model(s) 130; the models 130 may process information received from the Al Agent Controller 102 regarding potential goals, sub-goals, action-selection prioritization, threats , and contextual and associative and other forms of information, whether visual, text-based, or in other formats; the model(s) may produce outputs which are interpretable by the Al Agent 112 and may impact on the Al Agent's action selection in the environment over any given time horizon; in a first stage, the Al Agent Controller 102 may be customized for domain-specific uses; a role of the ULLM 114 is to process language-token inputs from the Priming Module 110 in the context of its training data, and produce relevant outputs which may be transferred to the Al Agent 112 to influence its actions; with a prompt such as "I was happy when I saw that the weather was," a likely output of the ULLM 114 is "sunny" , or "beautiful"; however, the output of the ULLM 114 may be substantially longer, including long-form text, depending on the prompt and model of the ULLM 114; by interpreting the data provided by the Priming Module 110, the ULLM 114 produces outputs which are consistent with human associations latent in its model weights; in the video game Montezuma's Revenge, such associative outputs mean negative human associations with the environment item "skull” yield low probability of directing the Al Agent 112 to interact with such an element in the environment; this example shows how the Al Agent Controller 102 may generate human associations and map the associations to the Al Agent's environment; as a result , the Al Agent Controller 102 may be dynamic and adaptive, improving performance of the Al Agent 112 in the environment; in a second Stage, the Priming Module 110, may translate or convert information from the Al Agent's environment into text or text-token format for processing by the ULLM 114; in a third Stage, the Priming Module 110 may also incorporate optimization processes that condition, translate, or otherwise transform the language representation outputs produced by the Visual/Natural Language Mapping Module 108 before the language representation outputs are transferred to the ULLM 114; in a fourth stage , the Al Agent Controller 102 may convert text data from the ULLM 114, where such data is not able to be processed by the Al Agent 112, into a form or format in which the information may be used or processed by the Al Agent 112 such that it may influence or improve action selection by the Al Agent 112 in the environment; a simple example might be to provide directional indications , or suggested actions, action tokens, or action types for the Al Agent to follow; a less direct example of such a process using the previous Montezuma's Revenge example is a heatmap- or bounding-box-based output to the Al Agent 112, which uses colors or textures associated by the Al Agent 112 with negative or positive rewards; e.g., when the ULLM 114 indicates a negative association with the skull, the Heat Map Generator 126 may highlight the skull to the Al Agent 112 in the color associated with negative rewards; the Al Agent 112 may then process this reward - expectation indication or associate it with a given action or object or location or other element of the environment; the Al Agent Controller 102 uses the ULLM 114 as an abstractive and generalizing engine which encodes human mental models such that for a given input , the Al Agent Controller 102 may measure or score the applicability of a given mental model via the outputs of the ULLM 114 and provide such data to the Al Agent 112; the Al Agent to leverage human knowledge and mental model frameworks which may be encoded in the ULLM 114 of the Al Agent Controller 102, such that the Al Agent 112 may receive data or information via which enables the Al Agent 112 to use or act upon the basis of human knowledge or mental models pertaining to certain elements of the environment and / or the interactions of such elements of the environment; the Al Agent Controller 102 provides a system which encodes human knowledge and association frame works in a framework which enables the Al Agent 112 to leverage such knowledge and associations in its action or policy selection(s); the Al Agent Controller 102 may provide a novel means for the Al Agent 112 to access past human "experience”, encoded in the model via vast volumes of training data used to shape the weights of the network of the ULLM 114 , to leverage thought templates for typical human reasoning or thought patterns, and to combine them with new information and context to provide inference related to human thought models, and a mechanism through which such outputs of the ULLM 114 may influence or direct the actions of Al Agent 112 in an environment; ¶¶ [0062]-[0070] with FIG. 2: operations may begin with the Visual/Natural Language Mapping Module 108 captioning (202) visual data; in the example , the environment is a video game called Montezuma's Revenge, and the image is a screenshot from the game; operations may continue by the Priming Module 110 generating (204) a text prompt for the ULLM 114; the ULLM 114 may generate (206) output text from the text prompt; the Al Agent Controller 102 may provide (208) the output text to the Al Agent 112; the Al Agent 112 may map semantics provided in the output text to actions; alternatively or in addition, the Heat Map Generator 126 may convert (210) the output text from the ULLM 114 into a heat map as shown in FIG . 2; e.g., in the heat map, an area around the key may be green, and areas around the fire and the skull, respectively, may be red; the logic may include additional, different, or fewer operations than illustrated in FIG. 2). GOSLIN and O’Malia are analogous art because they are from the same field of endeavor, a system and a method relating to artificial intelligence (AI) Agents. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of O’Malia to GOSLIN. Motivation for doing so would enhance the performance of the Al Agent in a . Claims 2 and 10 GOSLIN in view of O’Malia discloses all the elements as stated in Claims 1 and 9 respectively and further discloses monitor the video game environment for a subsequent input based on the provided model output, in response to receiving the subsequent input, analyze, by the director service, the video game environment for a specific context based on the subsequent input; associate the subsequent input with the one or more environment guidelines that provide systemic context about the video game environment; determine a second intent objective based on one or more of the subsequent input, the specific context and the one or more environment guidelines; and generate a second prompt for the generative machine learning model based on the second intent objective (GOSLIN, ¶¶ [0014]-[0022]: the AI system determines the context of a given input, and generates responses using one or more ML models that correspond to that context; the AI system can dynamically select ML models as the context of the interaction shifts, in order to continue to provide deep conversation; the AI system responds in part based on whether the determined means the user is relying on aligns with the expected means for the determined role the user is playing; in order to respond based on the alignment between the expected means and the actual means the user is utilizing, the AI system uses one or more fitness functions to ensure the output makes logical sense; the AI system similarly uses a fitness function to ensure that the output is authentic to a particular character (e.g., the character that the AI system is playing); the fitness function can evaluate each line corresponding to the character from the movie, show, or other predefined script, and then evaluate the user input to determine a mathematical distance between the user's word choice and phrasing, as compared to the "canonical" phrases; based on this distance, the AI system can determine whether the user is accurately playing the role; the AI system utilizes NLP and/or NLU to determine the objective, role, and/or means the user is using for the role-playing scenario; e.g., the AI system may search for keywords in user input that have predefined associations with a role, means, and/or objective; some or all of the context may be known or provided to the AI system; e.g., the user can select one or more aspects of the scenario (e.g., the objective, the role, and/or the means), and the AI system can infer the remaining aspects; the AI system repeatedly determines the context for each input (which may include analyzing prior input), such that the AI system can respond to shifting contexts (e.g., if the user switches from a stealth methodology to a brute-force methodology); once the context of a given input is determined, the AI system selects a corresponding ML model to process the input; the AI system includes a respective ML model trained for each roleplaying scenario (e.g., trained for each combination of objective, role, and means); the AI system uses a constant mapping from the determined context to the corresponding ML model; the AI system maintains context-specific weights for each ML model; given an input context, the AI system can probabilistically select a ML model to use based on the context-specific weight associated with each; the AI system can modify these context-specific weights based on user feedback, in order to better select ML models for future interactions; ¶ [0027] with FIG. 1: these Responses 125 are provided to the User 105. In tum, the User 105 may provide an additional Input 110. In this way, the User 105 and AI System 115 can interact during the role-playing scenario until the User 105 quits, or until predefined criteria are met; ¶¶ [0046]-[0048]: after using a particular ML Model 120 to generate a response, the ML Component 345 evaluates the user's next input in order to refine the context-specific weight(s) associated with the ML Model 120; e.g., if the user-response is positive (e.g., with a positive sentiment evaluation from the NLP Component 335), the ML Component 345 may increase the weight of the previously-selected ML Model 120, in order to increase the probability that it will be selected in the future, given the same context; similarly, if the user's subsequent response is negative, the ML Component 345 may reduce the weight of the previously-selected ML Model 120; based on the subsequent user-input, the ML Component 345 can modify or refine the selected ML Model 120; in order to determine the quality of a given output, the ML Component 345 can evaluate explicit ratings from the user, subsequent responses from the user, facial expressions or other non-verbal emotional cues (e.g., laughing) from the user, and the like; if the subsequent user input is positive, the ML Component 345 can refine the model to increase the probability that the previous output will be selected again, given the same input and/or context. Similarly, if the subsequent input is negative, the ML Component 345 reduces the probability that the ML Model 120 will select the same response again, given the previous input/context; in this way, the Interactivity Application 330 can provide immersive role-playing to the user; ¶¶ [0049]-[0052] with FIGS. 3 and 4: at block 405, where an Interactivity Application 330 receives user input; at block 410, the Interactivity Application 330 evaluates the input to determine the current context; at block 415, where the Interactivity Application 330 identifies and selects one or more ML Models 120 based on the determined context; the Interactivity Application 330 further refines the context-specific weights of each ML Model 120, and/or the internal weights of the previously-selected model, based on the current input; at block 420, the Interactivity Application 330 generates a response using the identified and selected ML Model 120; at block 425, the response is returned to the user (e.g., by displaying it on a screen, or by outputting audio); this process can then be repeated until the user exits the scenario, or the scenario otherwise terminates) (O'Malia, ¶¶ [0033]-[0061] with FIG. 1: the Al Agent 112 may query the Al Agent Controller 102 which uses the ULLM 114 as an abstraction or generalization engine; specifically , the Al Agent 112 may use the ULLM 114 to generalize past experiences to the context of current and/or or past environment data of the Al Agent 112 or to provide additional context, information, or suggestions as to what action or policy might be most appropriate; the action or policy may be deemed most appropriate based on information or representations — visual or in text or other form — that the Al Agent 112 provides to the Al Agent Controller 102 about current and past observations and action space of the Al Agent 112, as well as perceived goals of the Al Agent 112; because most forms of information in environments of the Al Agent 12 may be visual, a method is implemented in which the environment information of the Al Agent 112 is converted into a format which may be processed by the Al Agent Controller 102, and via which the outputs of the Al Agent Controller 102 may be converted into a format which may be understood by the Al Agent; this may include methods to optimize the prompting, or structured querying, of the Al Agent Controller 102 to elicit certain types of responses which may provide particular value to the Al Agent 112; the Al Agent 112 may share a representation of the environment, or describe an element, or indicate a relationship between elements in an environment, or outline the components of its environment, and provide the information to the Al Agent Controller 102; the Al Agent Controller 102 may generalize the provided information to other, related scenarios or situations so as to propose actions, goals, and/or sub-goals based on the past experience of the Al Agent Controller 102; the Al Agent Controller 102 may use multiple aspects of the inputs or prompts to interpret the context of the environment state of the Al Agent 112 such that the Al Agent Controller 102 may generate a relevant output to provide to the Al Agent 112; by incorporating contextual information, the Al Agent Controller 102 may increase the likelihood of the Al Agent 112 selecting an appropriate or useful action and/or recognizing which elements of the environment of the Al Agent 112 may be important to consider when making a decision; the Priming Module 110 may be configured to convert visual data such as a digital image to a natural language description of the image, objects in the image, and/or action (s) occurring in the image or series of images; the Visual/Natural Language Mapping or captioning Module 108 may use methods with explicit relational and geometric reasoning components such as Image Captioning: Transforming Objects into Words, 2020; the Priming Module 110 is configured to use the ULLM 114; ULLMs are generally "prompted" or "primed" with text (such as with "masked" and "left-to-right" text generation tasks ), and the ULLM is then asked to produce text compatible with the prompt; the Priming Module 110 is configured to generate a text prompt for the ULLM 114 and to provide the text prompt to the ULLM 114; these generated prompts tend to benefit from specific and structured requests, in the sense that such specific requests tend to generate more reliably structured, relevant, and interpretable outputs; the Priming Module 110 is configured to receive a text output from the ULLM 114 in response to the supplied text prompt; the Priming Module 110 may be optimized to evaluate the relative value and success of the outputs from the ULLM 114 in terms of enhancing the performance of the Al Agent 112 in order to build a repository of proven mental models or thought templates to which the ULLM 114 responds in a reliable , accurate , and structured way; the Priming Module 110 may include a discriminator 118 which may be included in a generative adversarial network (GAN) included in the Priming Module 110' the Priming Module 110 may include any other type of reinforcement learning structure and/or an imitation learning structure to evaluate the relative value and success of the outputs from the ULLM 114 in terms of enhancing the performance of the Al Agent 112; the discriminator 118 and/or other learning structure may include a neural network which evaluates whether the output of the ULLM 114 is sufficiently relevant to the inputs of the Al Agent 112 and the environment state of the Al Agent 112 that the output of the ULLM 114 might provide value to the Al Agent 112; the Priming Module 110 may store and utilize Shared Representations 120, where Shared Representations 120 are models and templates that, when combined with a given data type or input type from the environment and/or the Al Agent 112, may reliably trigger the application or use of a given thought template, mental model, value system , or similar high-level logical framework by the Al Agent Controller 102; one example class of Shared Representations 120 may be: list completion of things which belong together; such representations may be activated by the Priming Module 110 and passed as a text prompt to the ULLM 114 in certain situations; when the Al Agent 112 has or sees a variety of objects in an inventory or on the screen, which may be presented to the ULLM 114 as a list for completion or matching; in such a case, the Priming Module 110 may pass the list to the ULLM 114 for separation or sorting and add a predetermined prompt for "things which belong"; the ULLM 114 may determine the pattern or class that coherently represents some or all of the objects and suggest outputs which continue the pattern; one example is to prompt the ULLM 114 with the first three colors of the rainbow "Red , Orange , Yellow ..." , and the model of the ULLM 114 would typically pick up on the nature of the pattern as a shared representation of "things which belong" and map this to color order in a rainbow — although the ULLM 114 may also make other associations; the expected return from the ULLM 114 in this example may be "Green , Blue , Indigo , Violet"; such information may favorably inform the action selection of the Al Agent 112 in a game without the Al Agent 112 having such direct knowledge or models; in Montezuma's Revenge, a snapshot of the environment may be conveyed to the Priming Module 110 via the Visual/Natural Language Map ping Module 108 which may then list likely semantics of the environment and positions of objects on the screen, which when provided to the ULLM 114, may result in suggestions via an interface overlay such as a heatmap which indicates the importance of avoiding the skull and the benefit of acquiring the key; e.g., a Heat Map Generator 126 included in the conversion module 124 may generate the heatmap from text returned by the ULLM 114; the Al Agent 112 may include a neural-network, artificially-intelligent, or deep learning model(s) 130; the models 130 may process information received from the Al Agent Controller 102 regarding potential goals, sub-goals, action-selection prioritization, threats , and contextual and associative and other forms of information, whether visual, text-based, or in other formats; the model(s) may produce outputs which are interpretable by the Al Agent 112 and may impact on the Al Agent's action selection in the environment over any given time horizon; in a first stage, the Al Agent Controller 102 may be customized for domain-specific uses; a role of the ULLM 114 is to process language-token inputs from the Priming Module 110 in the context of its training data, and produce relevant outputs which may be transferred to the Al Agent 112 to influence its actions; with a prompt such as "I was happy when I saw that the weather was," a likely output of the ULLM 114 is "sunny" , or "beautiful"; however, the output of the ULLM 114 may be substantially longer, including long-form text, depending on the prompt and model of the ULLM 114; by interpreting the data provided by the Priming Module 110, the ULLM 114 produces outputs which are consistent with human associations latent in its model weights; in the video game Montezuma's Revenge, such associative outputs mean negative human associations with the environment item "skull” yield low probability of directing the Al Agent 112 to interact with such an element in the environment; this example shows how the Al Agent Controller 102 may generate human associations and map the associations to the Al Agent's environment; as a result , the Al Agent Controller 102 may be dynamic and adaptive, improving performance of the Al Agent 112 in the environment; in a second Stage, the Priming Module 110, may translate or convert information from the Al Agent's environment into text or text-token format for processing by the ULLM 114; in a third Stage, the Priming Module 110 may also incorporate optimization processes that condition, translate, or otherwise transform the language representation outputs produced by the Visual/Natural Language Mapping Module 108 before the language representation outputs are transferred to the ULLM 114; in a fourth stage , the Al Agent Controller 102 may convert text data from the ULLM 114, where such data is not able to be processed by the Al Agent 112, into a form or format in which the information may be used or processed by the Al Agent 112 such that it may influence or improve action selection by the Al Agent 112 in the environment; a simple example might be to provide directional indications , or suggested actions, action tokens, or action types for the Al Agent to follow; a less direct example of such a process using the previous Montezuma's Revenge example is a heatmap- or bounding-box-based output to the Al Agent 112, which uses colors or textures associated by the Al Agent 112 with negative or positive rewards; e.g., when the ULLM 114 indicates a negative association with the skull, the Heat Map Generator 126 may highlight the skull to the Al Agent 112 in the color associated with negative rewards; the Al Agent 112 may then process this reward - expectation indication or associate it with a given action or object or location or other element of the environment; the Al Agent Controller 102 uses the ULLM 114 as an abstractive and generalizing engine which encodes human mental models such that for a given input , the Al Agent Controller 102 may measure or score the applicability of a given mental model via the outputs of the ULLM 114 and provide such data to the Al Agent 112; the Al Agent to leverage human knowledge and mental model frameworks which may be encoded in the ULLM 114 of the Al Agent Controller 102, such that the Al Agent 112 may receive data or information via which enables the Al Agent 112 to use or act upon the basis of human knowledge or mental models pertaining to certain elements of the environment and / or the interactions of such elements of the environment; the Al Agent Controller 102 provides a system which encodes human knowledge and association frame works in a framework which enables the Al Agent 112 to leverage such knowledge and associations in its action or policy selection(s); the Al Agent Controller 102 may provide a novel means for the Al Agent 112 to access past human "experience”, encoded in the model via vast volumes of training data used to shape the weights of the network of the ULLM 114 , to leverage thought templates for typical human reasoning or thought patterns, and to combine them with new information and context to provide inference related to human thought models, and a mechanism through which such outputs of the ULLM 114 may influence or direct the actions of Al Agent 112 in an environment; ¶¶ [0062]-[0070] with FIG. 2: operations may begin with the Visual/Natural Language Mapping Module 108 captioning (202) visual data; in the example , the environment is a video game called Montezuma's Revenge, and the image is a screenshot from the game; operations may continue by the Priming Module 110 generating (204) a text prompt for the ULLM 114; the ULLM 114 may generate (206) output text from the text prompt; the Al Agent Controller 102 may provide (208) the output text to the Al Agent 112; the Al Agent 112 may map semantics provided in the output text to actions; alternatively or in addition, the Heat Map Generator 126 may convert (210) the output text from the ULLM 114 into a heat map as shown in FIG . 2; e.g., in the heat map, an area around the key may be green, and areas around the fire and the skull, respectively, may be red; the logic may include additional, different, or fewer operations than illustrated in FIG. 2). Claims 3 and 11 GOSLIN in view of O’Malia discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein generating a prompt further comprises: associating one or more prompt templates with the intent objective; and combining the one or more prompt templates into the prompt (GOSLIN, ¶¶ [0015]-[0018]: the AI system responds in part based on whether the determined means the user is relying on aligns with the expected means for the determined role the user is playing; each objective corresponds to a suggested or best means or role; in order to respond based on the alignment between the expected means and the actual means the user is utilizing, the AI system uses one or more fitness functions to ensure the output makes logical sense; the AI system similarly uses a fitness function to ensure that the output is authentic to a particular character (e.g., the character that the AI system is playing); this system can also be used to evaluate input from the user; e.g., suppose the user is roleplaying as a particular character from a move or series; the fitness function can evaluate each line corresponding to the character from the movie, show, or other predefined script, and then evaluate the user input to determine a mathematical distance between the user's word choice and phrasing, as compared to the "canonical" phrases; based on this distance, the AI system can determine whether the user is accurately playing the role; the AI system may search for keywords in user input that have predefined associations with a role, means, and/or objective; ¶¶ [0023] and [0026] with FIG. 1: each ML Model 120 corresponds to a machine learning model that was trained to receive textual input and generate or select a corresponding output; the output corresponds to text or audio that is selected from a predefined set of responses; the ML Model 120 acts as a classifier to classify the input, and uses corresponding predefined text for that category as the output; using a classifier as the ML Model 120, the AI System 115 may output pre-recorded phrases as the Response 125; ¶¶ [0028]-[0030] with FIG. 2: the Objective 210 is selected from a predefined set; the Roles 220 and Means 230 are similarly selected from predefined sets; the options available for a given selection can depend on one or more other selections; i.e., the selections for each of the Objectives 210, Roles 220, and/or Means 230 may have predefined relationships defining combinations that can be selected; ¶¶ [0024] and [0036]-[0037]: the training data includes sample input phrases used as input, as well as corresponding target output phrases to train each model; the training data includes samples of input text and a corresponding output text for each sample of input text; e.g., the training data is collected from roleplaying sessions between users (e.g., as part of a development team), and/or from early test users; the input from each user can be used as either exemplar input or target output, depending on the particular model being trained e.g., depending on which role the AI system will be playing, and which role the user will be playing); in one embodiment, the same set of training data can be used to train two ML Models 120, one for each side of the conversation; in this way, the AI System 115 can be trained to play two different roles based on a single set of training data; in some embodiments, a separate set of training data is used for each ML Model 120, where each set of training data is associated with a respective context (e.g., a role-playing scenario); the context of a set of training data is the role-playing scenario that the training data corresponds to; e.g., a first piece of training data may include one or more textual inputs, along with corresponding responses, that were recorded during a role-playing experience (e.g., between users, writers, actors, or other humans); i.e., the training data is collected by recording textual interactions between people (e.g., two writers role-playing as part of a scenario); the context of this first piece of training data can include identifiers for the objective, role, and/or means that the users were engaging in when the text was recorded; in this way, each ML Model 120 can be trained based on the particular underlying scenario, which can improve responses for the given scenario; each ML Model 120 is labeled or otherwise associated with the context that corresponds to the underlying training data; in this way, the AI System 115 can selectively use each ML Model 120 based on their specific contexts (e.g., based on matching the context of the current input with the ML Models 120) to generate deeper and more specific responses within each context; ¶¶ [0041] and [0047] with FIG. 3: the Context Component 340 infers the user's objective, role, and/or means based on identified keywords, or based on other results from the NLP Component 335; e.g., the Context Component 340 can determine that the objective is to retrieve an item, based on determining that predefined keywords relating to the item were included in one or more inputs received from the user; similarly, the Context Component 340 may determine that the user is roleplaying as a particular character or is using a particular means, based on keywords, and/or based on the sentiment or intent of the input(s), as identified by the NLP Component 335; the ML Model 120 is trained as a classifier that receives input ( e.g., text, the results of one or more NLP processes, and the like) and select from a large set of predefined responses; these responses may similarly include textual responses and/or pre-recorded audio; ¶ [0051]: the models are trained to classify the input in order to select from a predefined set of responses). Claims 7 and 15 GOSLIN in view of O’Malia discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein the set of operations further comprises: store one or more of the input, the intent objective, the prompt, and the model output (GOSLIN, ¶¶ [0031]-[0033]: each user has a corresponding user profile that maintains interaction history for the user, such as the scenarios they have previously engaged in (e.g., the objective, role, and/or means they used), the length of time they have spent interacting during each scenario, and the like; as the user interacts with the system in the scenario, the system maintains data about the interaction, such as the length of time the scenario lasts (e.g., until the user succeeds, fails, or ends the scenario), the result of the scenario (e.g., whether the user succeeded, failed, or quit), and the like; the system further utilizes techniques such as NLP to perform semantic analysis in order to determine whether the user enjoyed the scenario; this data is maintained in the profile of the user; ¶¶ [0036]-[0038] with FIG. 3: the training data is collected from roleplaying sessions between users (e.g., as part of a development team), and/or from early test users; the input from each user can be used as either exemplar input or target output, depending on the particular model being trained (e.g., depending on which role the AI system will be playing, and which role the user will be playing); the context of a set of training data is the role-playing scenario that the training data corresponds to; e.g., a first piece of training data may include one or more textual inputs, along with corresponding responses, that were recorded during a role-playing experience (e.g., between users, writers, actors, or other humans); i.e., the training data is collected by recording textual interactions between people (e.g., two writers role-playing as part of a scenario); the context of this first piece of training data can include identifiers for the objective, role, and/or means that the users were engaging in when the text was recorded; in this way, each ML Model 120 can be trained based on the particular underlying scenario, which can improve responses for the given scenario; each ML Model 120 is labeled or otherwise associated with the context that corresponds to the underlying training data; in this way, the AI System 115 can selectively use each ML Model 120 based on their specific contexts (e.g., based on matching the context of the current input with the ML Models 120) to generate deeper and more specific responses within each context; ¶¶ [0053]-[0056]: maintain user profiles for each user, where the user profile specifies information about previous interactions, such as the previous scenarios they have participated in, the length of time they interacted with each, and the like; determine, for each such prior scenario, a level of satisfaction the user experienced; this may be determined based on user selection (e.g., indicating a rating at the end of the interaction), and/or inferred based on sentiment analysis of the user input during and/or after the interaction; evaluate other user profiles and/or recent interactions from other users in order to identify role(s), objective(s), and/or mean(s) have been recently used by others or that are popular among other users; infers the user's preferences based on the number of times a particular selection was used, the duration of interactivity with the selection, the average satisfaction when the selection was used, how recently the selection was used, and the like; determine the user's preferred selections with respect to some factors of the scenario (e.g., the role, objective, and/or means), and selects one of these factors to change). Claims 8, 16, and 20 GOSLIN in view of O’Malia discloses all the elements as stated in Claims 1, 9, and 19 respectively and further discloses when, in response to determining that the model output is not responsive, the set of operations further comprises: generate a new prompt for the generative machine learning model based on the intent objective; and execute the generative machine learning model with the new prompt to produce a new model output (GOSLIN, ¶¶ [0014] and [0018]-[0022]: dynamically select ML models as the context of the interaction shifts, in order to continue to provide deep conversation; the AI system repeatedly determines the context for each input (which may include analyzing prior input), such that the AI system can respond to shifting contexts (e.g., if the user switches from a stealth methodology to a brute-force methodology); the AI system includes a respective ML model trained for each roleplaying scenario (e.g., trained for each combination of objective, role, and means); the AI system can select the corresponding ML model for evaluating the input; this can allow each model to be specialized with a constrained context, while allowing the AI system to dynamically shift between contexts by selecting other models; the AI system can select different ML models for the same context; e.g., the AI system may periodically or randomly use a different ML model for a given context, and evaluate the user's response; i.e., for a given a first context C, the AI system may determine to use an ML model trained on context C' to generate the response; the AI system can then analyze how the user responds, in order to determine whether the ML model corresponding to context C' should be used at least occasionally in the future, given context C; e.g., the system may determine whether the user liked the response, or appeared confused or frustrated; in order to improve responses and user-engagement, the learning system can provide differing responses to different users; the AI system can sometimes reach better solution by randomly (or pseudo-randomly) changing the order of things, prioritizing something that was previously lower priority, and the like; this randomness enables the AI system to continue searching for better solutions; e.g. the system may be at a local maxima for a solution, but there can be one or more better global maxima available which can be discovered through occasional random selections; the AI system maintains context-specific weights for each ML model; given an input context, the AI system can probabilistically select a ML model to use based on the context-specific weight associated with each; the AI system can modify these context-specific weights based on user feedback, in order to better select ML models for future interactions; ¶¶ [0023] and [0025]-[0026] with FIG. 1: each ML Model 120 corresponds to a machine learning model that was trained to receive textual input and generate or select a corresponding output; in one embodiment, this output includes dynamically generated text or audio; in another embodiment, the output corresponds to text or audio that is selected from a predefined set of responses; in one embodiment, the selected ML Model 120 dynamically generates output text based on weights learned during the training phase; in another embodiment, the ML Model 120 acts as a classifier to classify the input, and uses corresponding predefined text for that category as the output; i.e., when the output text or audio is not found in a predefined set of responses, the output text or audio can be dynamically generated; ¶¶ [0042]-[0048] with FIG. 3: the Context Component 340 continues to evaluate the context for each received input; the Context Component 340 can determine whether the user has shifted the role-playing scenario (such as by deciding to attempt a different means to achieve the objective); the Context Component 340 may identify this transition and select a different ML Model 120 (e.g., based on keywords in the current input, based on sentiment analysis on the input, based on intent analysis, and the like); in one embodiment, the Context Component 340 can allow this change and the role-playing scenario will be shifted accordingly (e.g., by selecting other ML Models 120); dynamically modify the context-specific weights of each ML Model 120 during use; after using a particular ML Model 120 to generate a response, the ML Component 345 evaluates the user's next input in order to refine the context-specific weight(s) associated with the ML Model 120; e.g., if the user-response is positive (e.g., with a positive sentiment evaluation from the NLP Component 335), the ML Component 345 may increase the weight of the previously-selected ML Model 120, in order to increase the probability that it will be selected in the future, given the same context. Similarly, if the user's subsequent response is negative, the ML Component 345 may reduce the weight of the previously-selected ML Model 120; the ML Model 120 dynamically generates an output, which may be textual, audio, text converted to audio using text-to-speech models, and the like; based on the subsequent user-input, the ML Component 345 can modify or refine the selected ML Model 120; in order to determine the quality of a given output, the ML Component 345 can evaluate explicit ratings from the user, subsequent responses from the user, facial expressions or other non-verbal emotional cues (e.g., laughing) from the user, and the like; if the subsequent user input is positive, the ML Component 345 can refine the model to increase the probability that the previous output will be selected again, given the same input and/or context; similarly, if the subsequent input is negative, the ML Component 345 reduces the probability that the ML Model 120 will select the same response again, given the previous input/context; ¶ [0051]: the ML Models 120 are trained to dynamically generate a response). Claim 17 GOSLIN in view of O’Malia discloses all the elements as stated in Claim 9 and further discloses wherein the interactive element comprises a non-player character (NPC), an animated infographic, a video, an image, a quiz, or a game object (GOSLIN, ¶ [0015]: the AI system acts as an intelligent character in a role-playing game; the AI system infers the context of the conversation based on input as the user interacts with the character; e.g., the AI system uses natural language processing (NLP) and/or natural language understanding (NLU) to attempt to identify role-playing scenario the user is partaking in; a roleplaying scenario is an interactive session that includes a role for the user to play, an objective for the user to pursue, and/or a means that the user should use to achieve the objective; the role can include any character, such as a child, a strong warrior, a wise wizard, a stealthy archer, and the like; similarly, objectives can include any goal, such as infiltrating a secure area, questioning or interrogating a character for information, retrieving items, rescuing characters, and the like; further, the means can similarly include any methodology of achieving the objective, such as stealth, force, intimidation, flattery, diversion, distraction, and the like; the user plays a more specific role, such as a particular character from a movie or television show; ¶¶ [0028]-[0033] with FIG. 2: the user interacts with a scenario creator to select one or more aspects of the role-playing scenario they wish to use; the user can select, in a first Box 205, an Objective 210A-N; the Objectives 210 can include, without limitation, an "Infiltrate" Objective 210A, a "Question" Objective 210B, and a "Retrieve" Objective 210N; the Objective 210 is selected from a predefined set; further, using the Box 215, the user can select a Role 220A-N; the user can select a Wizard 220A, a Warrior 220B, and a Ninja 220N; additionally, using the Box 225, the user can select a Means 230A-N, including Stealth 230A, Force 230B, and Flattery 230N; the Roles 220 and Means 230 are similarly selected from predefined sets; the user can select one or more of the options, and use the Button 245 to launch the selected scenario; the user may select only a subset of the aspects, and leave the AI System 115 to infer the remaining; the options available for a given selection can depend on one or more other selections; i.e., the selections for each of the Objectives 210, Roles 220, and/or Means 230 may have predefined relationships defining combinations that can be selected; e.g., the choice of available Roles 220 may depend on the Objective 210 selected; similarly, the available Means 230 may depend on the Objective 210 and/or the Role 220; the GUI 200 further includes a Button 235 to generate suggested scenarios, a Button 240 to generate random scenarios, and a Button 245 to launch the selected or suggested scenario; the Button 235 is used to generate suggested scenarios based on the user's previous interactions with the system; the Button 245 is used to begin the interactive scenario that has been selected and/or generated). Claim 18 GOSLIN in view of O’Malia discloses all the elements as stated in Claim 9 and further discloses wherein a gaming environment comprises a video game, an online game, a massively multiplayer online roleplaying game (MMORPG), or a virtual reality environment (GOSLIN, ¶ [0014]: provide AI systems that utilize machine learning to interact with users in a dynamic and immersive manner; ¶ [0015]: the AI system acts as an intelligent character in a role-playing game; ¶ [0016]: the user plays a more specific role, such as a particular character from a movie or television show; ¶ [0048]: the Interactivity Application 330 can provide immersive role-playing to the user). Claims 4-5 and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over GOSLIN in view of O’Malia as applied to Claims 3, 1, 11, and 9 respectively above, and further in view of Sharifi et al. (US 2022/0189474 A1, pub. date: 06/16/2022), hereinafter Sharifi. Claims 4 and 12 GOSLIN in view of O’Malia discloses all the elements as stated in Claims 3 and 11 respectively and further discloses wherein associating the one or more prompt templates further comprises: (GOSLIN, ¶¶ [0015]-[0018]: the AI system responds in part based on whether the determined means the user is relying on aligns with the expected means for the determined role the user is playing; each objective corresponds to a suggested or best means or role; in order to respond based on the alignment between the expected means and the actual means the user is utilizing, the AI system uses one or more fitness functions to ensure the output makes logical sense; the AI system similarly uses a fitness function to ensure that the output is authentic to a particular character (e.g., the character that the AI system is playing); this system can also be used to evaluate input from the user; e.g., suppose the user is roleplaying as a particular character from a move or series; the fitness function can evaluate each line corresponding to the character from the movie, show, or other predefined script, and then evaluate the user input to determine a mathematical distance between the user's word choice and phrasing, as compared to the "canonical" phrases; based on this distance, the AI system can determine whether the user is accurately playing the role; the AI system may search for keywords in user input that have predefined associations with a role, means, and/or objective; ¶¶ [0023] and [0026] with FIG. 1: each ML Model 120 corresponds to a machine learning model that was trained to receive textual input and generate or select a corresponding output; the output corresponds to text or audio that is selected from a predefined set of responses; the ML Model 120 acts as a classifier to classify the input, and uses corresponding predefined text for that category as the output; using a classifier as the ML Model 120, the AI System 115 may output pre-recorded phrases as the Response 125; ¶¶ [0028]-[0030] with FIG. 2: the Objective 210 is selected from a predefined set; the Roles 220 and Means 230 are similarly selected from predefined sets; the options available for a given selection can depend on one or more other selections; i.e., the selections for each of the Objectives 210, Roles 220, and/or Means 230 may have predefined relationships defining combinations that can be selected; ¶¶ [0024] and [0036]-[0037]: the training data includes sample input phrases used as input, as well as corresponding target output phrases to train each model; the training data includes samples of input text and a corresponding output text for each sample of input text; e.g., the training data is collected from roleplaying sessions between users (e.g., as part of a development team), and/or from early test users; the input from each user can be used as either exemplar input or target output, depending on the particular model being trained e.g., depending on which role the AI system will be playing, and which role the user will be playing); in one embodiment, the same set of training data can be used to train two ML Models 120, one for each side of the conversation; in this way, the AI System 115 can be trained to play two different roles based on a single set of training data; in some embodiments, a separate set of training data is used for each ML Model 120, where each set of training data is associated with a respective context (e.g., a role-playing scenario); the context of a set of training data is the role-playing scenario that the training data corresponds to; e.g., a first piece of training data may include one or more textual inputs, along with corresponding responses, that were recorded during a role-playing experience (e.g., between users, writers, actors, or other humans); i.e., the training data is collected by recording textual interactions between people (e.g., two writers role-playing as part of a scenario); the context of this first piece of training data can include identifiers for the objective, role, and/or means that the users were engaging in when the text was recorded; in this way, each ML Model 120 can be trained based on the particular underlying scenario, which can improve responses for the given scenario; each ML Model 120 is labeled or otherwise associated with the context that corresponds to the underlying training data; in this way, the AI System 115 can selectively use each ML Model 120 based on their specific contexts (e.g., based on matching the context of the current input with the ML Models 120) to generate deeper and more specific responses within each context; ¶¶ [0041] and [0047] with FIG. 3: the Context Component 340 infers the user's objective, role, and/or means based on identified keywords, or based on other results from the NLP Component 335; e.g., the Context Component 340 can determine that the objective is to retrieve an item, based on determining that predefined keywords relating to the item were included in one or more inputs received from the user; similarly, the Context Component 340 may determine that the user is roleplaying as a particular character or is using a particular means, based on keywords, and/or based on the sentiment or intent of the input(s), as identified by the NLP Component 335; the ML Model 120 is trained as a classifier that receives input ( e.g., text, the results of one or more NLP processes, and the like) and select from a large set of predefined responses; these responses may similarly include textual responses and/or pre-recorded audio; ¶ [0051]: the models are trained to classify the input in order to select from a predefined set of responses). GOSLIN in view of O’Malia fails to explicitly disclose generating an embedding for the intent objective; and identifying one or more prompt templates that are semantically associated with the intent objective based on the embedding. Sharifi teaches a system and a method relating to perform interactive actions using intelligent assistant (Sharifi, ¶¶ [0001]-[0002]), wherein generating an embedding for the intent objective; and identifying one or more prompt templates that are semantically associated with the intent objective based on the embedding (Sharifi, ¶ [0008]: if the NL only clarification prompt is "do you want news about the actor John Doe or the producer John Doe", it can be determined that "actor" and "producer" satisfy a semantic similarity threshold; e.g., embeddings can be generated for "actor" and "producer" using a trained encoder (e.g., a trained neural network model), a distance between the "actor" embedding and the "producer" embedding determined, and the distance determined to satisfy the semantic similarity threshold be determined to provide an enhanced clarification prompt instead, such as one that includes a first image of the actor and a second image of the producer; ¶¶ [0027]-[0029]: dialog manager 118 may be configured to map a representation of a user request to perform some action, e.g., using the annotations, to one or more "responsive actions" of a plurality of candidate responsive actions that are then performed by automated assistant 100; mappings may include mappings between entities and candidate responsive actions that are performable in association with those entities; dialog manager 118 may employ one or more trained machine learning models, alone or in combination with one or more grammars; these trained machine learning models may be trained to identify intents, e.g., by embedding data indicative of a user's utterance into a latent space, and then determining which other embeddings (and therefore, intents) are most proximate, e.g., using techniques such as Euclidean distance, cosine similarity, etc.; various contextual signals may be used to perform various aspects of the natural language processing and dialog managing features; entity or entity type recognition, entity or entity type ranking, identification of candidate responsive actions associated with entities or entity types, ranking of candidate responsive actions, and/or filtering of candidate responsive actions, may be performed based on contextual signals; ¶¶ [0037]-[0038]: the NL only clarification prompt can be generated based on an NL only clarification prompt template; the NL only clarification prompt template may be pre-generated, or may be generated responsive to identifying the two or more candidate responsive actions as corresponding to the user's spoken utterance; the NL only clarification prompts can include slots filled by the natural language characterizations of the candidate responsive actions to be rendered in the prompt; the system may generate such natural language characterizations of the candidate responsive actions based on data generated during the natural language processing; an NL only clarification prompt may be the default clarification prompt used when the automated assistant 100 determines it cannot select between two or more candidate responsive actions for user requests generally, or for certain types of user requests; however, the automated assistant 100 may instead select an enhanced clarification prompt based on a variety of factor (s)/condition(s); ¶ [0051]: the system can generate an embedding based on the identified semantic property (e.g., using a trained neural network encoder) and compare the embedding to a plurality of embeddings of respective semantic properties associated with the candidate responsive actions corresponding to the renderings of the clarification prompt; the plurality of embeddings of respective semantic properties associated with the candidate responsive actions may have been generated by the trained neural network encoder or another trained neural network encoder based on metadata that indicates semantic properties associated with the candidate responsive actions; the system can determine that the given semantic property matches, or most closely matches, a given embedding, of the plurality of embeddings of the respective semantic properties, based on the comparison; e.g., assume the embeddings are word2vec representations; in this example, a cosine distance between the word2vec representation of the semantic property and each of the word2vec representations of the respective semantic properties of the candidate responsive actions of the prompt can be determined, and a given semantic property of a candidate responsive action that is associated with a respective cosine distance that satisfies a distance threshold can be utilized to determine the semantic property of the spoken utterance matches, or most closely matches, the given semantic property that is associated with a given candidate responsive action ( e.g., an exact match or "fuzzy" match); as a result, the given candidate responsive action that is associated with the given semantic property may be selected for performance; ¶ [0059]: the NL only clarification prompt can be generated based on an NL only clarification prompt template; the NL only clarification prompt template may be pre-generated, or may be generated responsive to identifying the two or more candidate responsive actions as corresponding to the user's spoken utterance; in the case of pre-generated NL only clarification prompt templates, the NL only clarification prompt template can be selected from among various NL only clarification templates based on the identified two or more candidate responsive actions that correspond to the user's spoken utterance; e.g., there may be NL only clarification prompt templates for online shopping, viewing or retrieving media content, interactions with a restaurant reservation application, booking flights, etc.; there may also be NL only clarification prompt templates for the various combinations of such actions, e.g., a clarification prompt for selecting between an online shopping action and a flight booking action; the NL only clarification prompt template may be selected from among various NL only clarification templates at least in part based on the natural language characterizations of the candidate responsive actions to be rendered in the prompt; e.g., if the clarification prompt is to include natural language characterizations of candidate responsive actions that are detailed and/or long-winded, then an NL only clarification prompt template that includes long pauses before and/or after the characterizations or that provides a summary at the end may be selected. In implementations in which the NL only clarification prompt template is generated after receiving the spoken utterance, it may likewise be tailored to the candidate responsive actions and/or their characterizations that are to be rendered in the clarification prompt; ¶¶ [0074]-[0076]: the system determines a similarity measure that reflects a textual and/or semantic similarity between the first term(s) and the second term(s); the system can embed the first term(s) as a first embedding in an embedding space and can embed the second term(s) as a second embedding in the embedding space using a trained encoder (e.g., a trained neural network embedding model); the system can use the embeddings to generate the similarity measure; the system can thus determine to provide the enhanced clarification prompt rather than the NL only clarification prompt based on the comparison(s) of the embeddings of the first and second terms and/or based on. the similarity measure; determine to provide the enhanced clarification prompt rather than the NL only clarification prompt based on determining that the similarity measure and/or embeddings indicate threshold level(s) of similarity and/or dissimilarity; the NL only clarification prompt template(s) and/or natural language characterizations of the candidate responsive actions have previously been generated, for this user or for another user as indicated by the historical automated assistant interaction data). GOSLIN in view of O’Malia, and Sharif are analogous art because they are from the same field of endeavor, a system and a method relating to . It is also well known in the art that embeddings are commonly used in the language processing models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Sharifi to GOSLIN in view of O’Malia. Motivation for doing so would help to . Claims 5 and 13 GOSLIN in view of O’Malia discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein generating the intent objective further comprises: environment guidelines (GOSLIN, ¶¶ [0015]-[0018]: the AI system uses natural language processing (NLP) and/or natural language understanding (NLU) to attempt to identify role-playing scenario the user is partaking in; the AI system can utilize techniques including keyword identification, sentiment analysis, intent evaluation, parsing to determine meaning, and the like; a roleplaying scenario is an interactive session that includes a role for the user to play, an objective for the user to pursue, and/or a means that the user should use to achieve the objective; the user plays a more specific role, such as a particular character from a movie or television show; the user is expected to understand the personality of the selected role, and must "stay in character" (e.g., by using an appropriate means) to achieve the objective; each role of the user is related or correlated to an expected or anticipated means or approach; e.g., an "intellect" means may correspond to a "wizard" character, while an "intimidation" means corresponds to a "knight" character; the AI system responds in part based on whether the determined means the user is relying on aligns with the expected means for the determined role the user is playing; similarly, each objective corresponds to a suggested or best means or role; in order to respond based on the alignment between the expected means and the actual means the user is utilizing, the AI system uses one or more fitness functions to ensure the output makes logical sense; the AI system similarly uses a fitness function to ensure that the output is authentic to a particular character (e.g., the character that the AI system is playing); this system can also be used to evaluate input from the user; e.g., suppose the user is roleplaying as a particular character from a move or series; the fitness function can evaluate each line corresponding to the character from the movie, show, or other predefined script, and then evaluate the user input to determine a mathematical distance between the user's word choice and phrasing, as compared to the "canonical" phrases; based on this distance, the AI system can determine whether the user is accurately playing the role; the AI system utilizes NLP and/or NLU to determine the objective, role, and/or means the user is using for the role-playing scenario; e.g., the AI system may search for keywords in user input that have predefined associations with a role, means, and/or objective; some or all of the context may be known or provided to the AI system; e.g., the user can select one or more aspects of the scenario (e.g., the objective, the role, and/or the means), and the AI system can infer the remaining aspects; the AI system repeatedly determines the context for each input (which may include analyzing prior input), such that the AI system can respond to shifting contexts; ¶¶ [0040]-[0045] with FIG. 3: the NLP Component 335 performs a variety of NLP processing such as sentiment analysis, keyword detection, parsing to determine intent or meaning, and the like; once the NLP Component 335 has determined the intent and/or sentiment, identified keywords, or performed any other NLP processing, some or all of the resulting data is passed as input to the ML Models 120; the Context Component 340 determines the current context of the user input in order to facilitate selection of an appropriate ML Model 120; the Context Component 340 infers the user's objective, role, and/or means based on identified keywords, or based on other results from the NLP Component 335; e.g., the Context Component 340 can determine that the objective is to retrieve an item, based on determining that predefined keywords relating to the item were included in one or more inputs received from the user; similarly, the Context Component 340 may determine that the user is roleplaying as a particular character or is using a particular means, based on keywords, and/or based on the sentiment or intent of the input(s), as identified by the NLP Component 335; the Context Component 340 may determine the context based at least in part on the original input provided by the user (e.g., the explicit selection(s) the user made in initiating the scenario); the Context Component 340 can determine whether the user has shifted the role-playing scenario (such as by deciding to attempt a different means to achieve the objective); the Context Component 340 may identify this transition and select a different ML Model 120 (e.g., based on keywords in the current input, based on sentiment analysis on the input, based on intent analysis, and the like); in some embodiments, some or all of the selections are not changeable during the interaction; e.g., the Context Component 340 may infer that the user is attempting to play as a different role, or pursue a different objective; in one embodiment, the Context Component 340 can allow this change and the role-playing scenario will be shifted accordingly (e.g., by selecting other ML Models 120); in another embodiment, the Context Component 340 can continue to output the previous context; e.g., the Context Component 340 can determine that the user is switching roles, but may nevertheless indicate, to the ML Component 345, that the role remains the same (e.g., because predefined rules indicate that the role cannot be changed mid-scenario); once the Context Component 340 has determined the current context (e.g., the objective, role, and means), the Context Component 340 provides an indication to the ML Component 345 reflecting these determinations). GOSLIN in view of O’Malia fails to explicitly disclose generating an embedding for one or more of the input, specific context, and environment guidelines; and identifying an intent objective based on the embedding. Sharifi teaches a system and a method relating to perform interactive actions using intelligent assistant (Sharifi, ¶¶ [0001]-[0002]), wherein generating an embedding for one or more of the input, specific context, and environment guidelines; and identifying an intent objective based on the embedding (Sharifi, ¶ [0008]: if the NL only clarification prompt is "do you want news about the actor John Doe or the producer John Doe", it can be determined that "actor" and "producer" satisfy a semantic similarity threshold; e.g., embeddings can be generated for "actor" and "producer" using a trained encoder (e.g., a trained neural network model), a distance between the "actor" embedding and the "producer" embedding determined, and the distance determined to satisfy the semantic similarity threshold be determined to provide an enhanced clarification prompt instead, such as one that includes a first image of the actor and a second image of the producer; ¶¶ [0027]-[0029]: dialog manager 118 may be configured to map a representation of a user request to perform some action, e.g., using the annotations, to one or more "responsive actions" of a plurality of candidate responsive actions that are then performed by automated assistant 100; mappings may include mappings between entities and candidate responsive actions that are performable in association with those entities; dialog manager 118 may employ one or more trained machine learning models, alone or in combination with one or more grammars; these trained machine learning models may be trained to identify intents, e.g., by embedding data indicative of a user's utterance into a latent space, and then determining which other embeddings (and therefore, intents) are most proximate, e.g., using techniques such as Euclidean distance, cosine similarity, etc.; various contextual signals may be used to perform various aspects of the natural language processing and dialog managing features; entity or entity type recognition, entity or entity type ranking, identification of candidate responsive actions associated with entities or entity types, ranking of candidate responsive actions, and/or filtering of candidate responsive actions, may be performed based on contextual signals; ¶¶ [0037]-[0038]: the NL only clarification prompt can be generated based on an NL only clarification prompt template; the NL only clarification prompt template may be pre-generated, or may be generated responsive to identifying the two or more candidate responsive actions as corresponding to the user's spoken utterance; the NL only clarification prompts can include slots filled by the natural language characterizations of the candidate responsive actions to be rendered in the prompt; the system may generate such natural language characterizations of the candidate responsive actions based on data generated during the natural language processing; an NL only clarification prompt may be the default clarification prompt used when the automated assistant 100 determines it cannot select between two or more candidate responsive actions for user requests generally, or for certain types of user requests; however, the automated assistant 100 may instead select an enhanced clarification prompt based on a variety of factor (s)/condition(s); ¶ [0051]: the system can generate an embedding based on the identified semantic property (e.g., using a trained neural network encoder) and compare the embedding to a plurality of embeddings of respective semantic properties associated with the candidate responsive actions corresponding to the renderings of the clarification prompt; the plurality of embeddings of respective semantic properties associated with the candidate responsive actions may have been generated by the trained neural network encoder or another trained neural network encoder based on metadata that indicates semantic properties associated with the candidate responsive actions; the system can determine that the given semantic property matches, or most closely matches, a given embedding, of the plurality of embeddings of the respective semantic properties, based on the comparison; e.g., assume the embeddings are word2vec representations; in this example, a cosine distance between the word2vec representation of the semantic property and each of the word2vec representations of the respective semantic properties of the candidate responsive actions of the prompt can be determined, and a given semantic property of a candidate responsive action that is associated with a respective cosine distance that satisfies a distance threshold can be utilized to determine the semantic property of the spoken utterance matches, or most closely matches, the given semantic property that is associated with a given candidate responsive action ( e.g., an exact match or "fuzzy" match); as a result, the given candidate responsive action that is associated with the given semantic property may be selected for performance; ¶ [0059]: the NL only clarification prompt can be generated based on an NL only clarification prompt template; the NL only clarification prompt template may be pre-generated, or may be generated responsive to identifying the two or more candidate responsive actions as corresponding to the user's spoken utterance; in the case of pre-generated NL only clarification prompt templates, the NL only clarification prompt template can be selected from among various NL only clarification templates based on the identified two or more candidate responsive actions that correspond to the user's spoken utterance; e.g., there may be NL only clarification prompt templates for online shopping, viewing or retrieving media content, interactions with a restaurant reservation application, booking flights, etc.; there may also be NL only clarification prompt templates for the various combinations of such actions, e.g., a clarification prompt for selecting between an online shopping action and a flight booking action; the NL only clarification prompt template may be selected from among various NL only clarification templates at least in part based on the natural language characterizations of the candidate responsive actions to be rendered in the prompt; e.g., if the clarification prompt is to include natural language characterizations of candidate responsive actions that are detailed and/or long-winded, then an NL only clarification prompt template that includes long pauses before and/or after the characterizations or that provides a summary at the end may be selected. In implementations in which the NL only clarification prompt template is generated after receiving the spoken utterance, it may likewise be tailored to the candidate responsive actions and/or their characterizations that are to be rendered in the clarification prompt; ¶¶ [0074]-[0076]: the system determines a similarity measure that reflects a textual and/or semantic similarity between the first term(s) and the second term(s); the system can embed the first term(s) as a first embedding in an embedding space and can embed the second term(s) as a second embedding in the embedding space using a trained encoder (e.g., a trained neural network embedding model); the system can use the embeddings to generate the similarity measure; the system can thus determine to provide the enhanced clarification prompt rather than the NL only clarification prompt based on the comparison(s) of the embeddings of the first and second terms and/or based on. the similarity measure; determine to provide the enhanced clarification prompt rather than the NL only clarification prompt based on determining that the similarity measure and/or embeddings indicate threshold level(s) of similarity and/or dissimilarity; the NL only clarification prompt template(s) and/or natural language characterizations of the candidate responsive actions have previously been generated, for this user or for another user as indicated by the historical automated assistant interaction data). GOSLIN in view of O’Malia, and Sharif are analogous art because they are from the same field of endeavor, a system and a method relating to . It is also well known in the art that embeddings are commonly used in the language processing models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Sharifi to GOSLIN in view of O’Malia. Motivation for doing so would help to . Claims 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over GOSLIN in view of O’Malia as applied to Claims 1 and 9 respectively above, and further in view of Murdock, IV et al. (US 2022/0309246 A1, pub. date: 09/29/2022), hereinafter of Murdock, IV. Claims 6 and 14 GOSLIN in view of O’Malia discloses all the elements as stated in Claims 1 and 9 respectively and further discloses wherein determining that the model output is responsive further comprises: responsiveness to the input and satisfaction of the one or more environment guidelines based on one or more metrics; (GOSLIN, ¶¶ [0015]-[0022]: the AI system responds in part based on whether the determined means the user is relying on aligns with the expected means for the determined role the user is playing; in order to respond based on the alignment between the expected means and the actual means the user is utilizing, the AI system uses one or more fitness functions to ensure the output makes logical sense; the AI system similarly uses a fitness function to ensure that the output is authentic to a particular character (e.g., the character that the AI system is playing); this system can also be used to evaluate input from the user; e.g., suppose the user is roleplaying as a particular character from a move or series; the fitness function can evaluate each line corresponding to the character from the movie, show, or other predefined script, and then evaluate the user input to determine a mathematical distance between the user's word choice and phrasing, as compared to the "canonical" phrases; based on this distance, the AI system can determine whether the user is accurately playing the role; the AI system repeatedly determines the context for each input (which may include analyzing prior input), such that the AI system can respond to shifting contexts; the AI system can select different ML models for the same context; e.g., the AI system may periodically or randomly use a different ML model for a given context, and evaluate the user's response; i.e., for a given a first context C, the AI system may determine to use an ML model trained on context C' to generate the response; the AI system can then analyze how the user responds, in order to determine whether the ML model corresponding to context C' should be used at least occasionally in the future, given context C; e.g., the system may determine whether the user liked the response, or appeared confused or frustrated; this evaluation can include receiving user feedback (e.g., a thumbs up or thumbs down, a score, and the like) and/or using NLP and/or NLU to analyze the user's response (e.g., to determine the sentiment); in order to improve responses and user-engagement, the learning system can provide differing responses to different users; this randomness enables the AI system to continue searching for better solutions; e.g., the system may be at a local maxima for a solution, but there can be one or more better global maxima available which can be discovered through occasional random selections; the AI system maintains context-specific weights for each ML model; given an input context, the AI system can probabilistically select a ML model to use based on the context-specific weight associated with each; further, the AI system can modify these context-specific weights based on user feedback, in order to better select ML models for future interactions; ¶ [0025] with FIG. 1: the AI System 115 periodically or probabilistically selects an ML Model 120, which may not be the model associated with a perfectly matching context; e.g., the AI System 115 may rely on weights associated with each ML Model 120 in determining which model to select; ¶¶ [0040]-[0048] with FIG. 3: the NLP Component 335 performs a variety of NLP processing such as sentiment analysis, keyword detection, parsing to determine intent or meaning, and the like; e.g., the NLP Component 335 may perform sentiment analysis to determine whether the user is enjoying the ongoing scenario. In some embodiments, the ML Models 120 can be refined based on this analysis; the Context Component 340 determines the current context of the user input, in order to facilitate selection of an appropriate ML Model 120; the Context Component 340 does so at least in part by evaluating results output by the NLP Component 335; this may include data relating to the most recent input, as well as from one or more prior inputs; the Context Component 340 infers the user's objective, role, and/or means based on identified keywords, or based on other results from the NLP Component 335; e.g., the Context Component 340 can determine that the objective is to retrieve an item, based on determining that predefined keywords relating to the item were included in one or more inputs received from the user; similarly, the Context Component 340 may determine that the user is roleplaying as a particular character or is using a particular means, based on keywords, and/or based on the sentiment or intent of the input(s), as identified by the NLP Component 335; the Context Component 340 may determine the context based at least in part on the original input provided by the user (e.g., the explicit selection(s) the user made in initiating the scenario); the Context Component 340 performs this context identification for each received input in the scenario; i.e., even if the Context Component 340 has already inferred the context with a high degree of confidence (or determined it conclusively based on explicit user-selection), the Context Component 340 continues to evaluate the context for each received input; in this way, the Context Component 340 can determine whether the user has shifted the role-playing scenario (such as by deciding to attempt a different means to achieve the objective); the Context Component 340 may identify this transition and select a different ML Model 120 (e.g., based on keywords in the current input, based on sentiment analysis on the input, based on intent analysis, and the like); the ML Component 345 identifies and selects the ML Model 120 that is associated with a matching context; the ML Component 345 may select one or more other ML Models 120 (e.g., periodically, or in a probabilistic manner); e.g., the ML Models 120 are associated with context-specific weights, indicating a likelihood that each will be selected given a particular context; given a first context C, a first ML Model 120 (e.g., one trained on the same context C) may be associated with a relatively high weight, such that it will be selected frequently; similarly, a second ML Model 120 (trained on a different context) can be associated with a relatively lower weight, such that it is selected less frequently than the first; the weight of each ML Model 120 is determined based in part on the vector distance between the current context C and the respective context C' of each respective ML Model 120; the ML Component 345 can dynamically modify the context-specific weights of each ML Model 120 during use; after using a particular ML Model 120 to generate a response, the ML Component 345 evaluates the user's next input in order to refine the context-specific weight(s) associated with the ML Model 120; e.g., if the user-response is positive (e.g., with a positive sentiment evaluation from the NLP Component 335), the ML Component 345 may increase the weight of the previously-selected ML Model 120, in order to increase the probability that it will be selected in the future, given the same context; similarly, if the user's subsequent response is negative, the ML Component 345 may reduce the weight of the previously-selected ML Model 120; i.e., each weight of different ML Models 120 indicates confidence scores of each ML Model 120; in addition to receiving the current context from the Context Component 340, the ML Component 345 receives results of the NLP analysis from the NLP Component 335; once an ML Model 120 has been selected, the ML Component 345 provides this input to the selected model in order to generate an output; based on the subsequent user-input, the ML Component 345 can modify or refine the selected ML Model 120; in order to determine the quality of a given output, the ML Component 345 can evaluate explicit ratings from the user, subsequent responses from the user, facial expressions or other non-verbal emotional cues (e.g., laughing) from the user, and the like; if the subsequent user input is positive, the ML Component 345 can refine the model to increase the probability that the previous output will be selected again, given the same input and/or context; similarly, if the subsequent input is negative, the ML Component 345 reduces the probability that the ML Model 120 will select the same response again, given the previous input/context; ¶¶ [0049]-[0051] with FIGS. 3-4: the Interactivity Application 330evaluates the input to determine the current context; in addition to determining a context, the Interactivity Application 330 generates a confidence in this determination; e.g., if the Interactivity Application 330 is inferring the context, the Interactivity Application 330 may further generate a corresponding confidence in order to aid selection of an appropriate ML Model 120; the Interactivity Application 330 identifies and selects one or more ML Models 120 based on the determined context; in one embodiment, the Interactivity Application 330 selects the ML Model 120 with a matching context; in another embodiment, the Interactivity Application 330 probabilistically selects a model based on the determined context, the confidence in this determination, and the context-specific weights associated with each ML Model 120; the Interactivity Application 330 further refines the context-specific weights of each ML Model 120, and/or the internal weights of the previously-selected model, based on the current input; the Interactivity Application 330 performs one or more NLP operations on the input (e.g., keyword identification, sentiment analysis, intent determination, and the like), and processes the result with the ML Model 120). GOSLIN in view of O’Malia fails to explicitly disclose receiving a confidence threshold value for evaluating the model output; and comparing the one or more confidence score for the one or more components of the model output against the confidence threshold value. Murdock, IV teaches a system and a method relating to artificial intelligence application (Murdock, IV, ¶ [0002]), wherein receiving a confidence threshold value for evaluating the model output; and comparing the one or more confidence score for the one or more components of the model output against the confidence threshold value (Murdock, IV, ¶¶ [0092]-[0096] and [0131]-[00137] with in FIG. 6: the score of a particular answer is a confidence score that indicates a relative measure of confidence that the answer is correct; e.g., the score may be a value between 0.0 and 1.0, with 0.0 representing a lowest confidence and 1.0 representing a highest confidence, with the measure of confidence increasing linearly between 0.0 and 1.0; at step 610, receive input defining parameter values (e.g., an initial value of a confidence threshold); at step 610, receive a question from a user device; at step 615, generate one or more answers to the question from step 610 using question answering system techniques; determine a confidence score for each of the one or more answers from step 615 using question answering system technique; determines which of the answers (from step 615) to return to the user that asked the question (at step 610) by comparing the respective confidence scores (determined at step 620) to a confidence threshold; compare the confidence score of each answer to the confidence threshold, returns (to the user device) answers whose confidence score is greater than the confidence threshold, and does not return answers whose confidence score is less than the confidence threshold; adjusts the confidence threshold based on comparing a highest one of the confidence scores (from step 620) to the confidence threshold; either: (i) increasing the confidence threshold for a next question in response to the highest one of the confidence scores being greater than the confidence threshold, or (ii) decreasing the confidence threshold for the next question in response to the highest one of the confidence scores being less than the confidence threshold.). GOSLIN in view of O’Malia, and Murdock, IV are analogous art because they are from the same field of endeavor, a system and a method relating to artificial intelligence application. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Murdock, IV to GOSLIN in view of O’Malia. Motivation for doing so would improve . Response to Arguments Applicant’s arguments filed 06/11/2026 with respect to Claims 1, 9, and 19 have been fully considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Roberts et al. ("Steps towards prompt-based creation of virtual worlds", arXiv:2211.05875v1, Nov. 10, 2022, pp. 1-15) discloses in Abstract of Page 1 that (1) large language models trained for code generation can be applied to speaking virtual worlds into existence (creating virtual worlds); (2) in this work we show that prompt-based methods can both accelerate in-VR level editing, as well as can become part of gameplay rather than just part of game development; (3) as an example, we present Codex VR Pong which shows non-deterministic game mechanics using generative processes to not only create static content but also non-trivial interactions between 3D objects; (4) this demonstration naturally leads to an integral discussion on how one would evaluate and benchmark experiences created by generative models – as there are no qualitative or quantitative metrics that apply in these scenarios; and (5) conclude by discussing impending challenges of AI-assisted co-creation in VR. Roberts further discloses in Section 1 of Pages 1-2 that (1) propose in this paper that these capabilities can be combined to allow "speaking the world into existence", or taking natural language descriptions and turning them into interactive visual scenes within a game engine; (2) in particular, this has the potential for allowing authoring Virtual Reality (VR) experiences from within the headset, as well as allow completely novel modes of gameplay; (3) integrating Codex with a game engine however should allow for a much easier real-time natural language interface for interactive scene creation; (4) this should in principle allow us to not only create game levels on-demand, but also allow non-coding users to speak training or educational scenarios with much less effort and time expenditure; (5) to show that prompt-based creation can become part of gameplay rather than just part of game development, we demonstrate the first VR game with non-deterministic game mechanics powered by OpenAI’s text generative models, by integrating them with the Unity game engine; (6) in reference to one of the first video games ever made, we built a surreal tennis game, in which the players can transform both the paddles and the ball into any 3d objects; and (7) these transformed objects then interact in semantically sensible ways that were not predetermined by the developer, for example a ball transformed into an egg colliding with a frying pan results in a fried egg. Roberts also discloses in Section 3 of Pages 3-5 that (1) virtual Reality and related modalities would benefit from prompt-based content generation, analogous to what Dall-E has done for images or GPT for text; (2) a prompt-based method that reduces the cost of generating VR content might not only spur the appearance of new content by reducing the time and cost of development, but might also make it more feasible to make already-existing content continue to be available; (3) a second motivation to explore prompt-based generation in VR is that spontaneous user-generated content in the course of a VR experience has been imagined as a core element of VR from its inception; (4) fictional depictions of VR, such as in Star Trek’s Holodeck, or The Matrix movies, tend to depict characters asking for the content and dynamics of the world to change in response to prompts; (5) furthermore, practical uses of VR would often be made economical by a prompting methodology; e.g., a physical therapist might prompt for a virtual exercise machine or environment that is tuned the needs of a patient with an unusual case; (6) at present, however, textual prompt-based methods have achieved at least initial utility in text, code, and image generation; (7) therefore it is natural to investigate spoken prompts as a method to "speak the world into existence", or change the contents and dynamics of a virtual world while one is experiencing it; (8) there must be a degree of semantic coherence to how objects can interact; e.g., (a) one should be able to pick up a virtual suitcase, but not the water in a pond, or at least not in the same way; (b) at the same time, interactions must not be too rigid; and (c) it should be possible for surprising interactions to emerge between objects or other elements; (9) coherent whole scenes should be prompt-able; (10) users should be able to refer to what has been created in order to modify it, using both words, descriptions, pointing, and other modalities; and (11) consultants in an ideation session could use voice prompts to generate multiple interactive 3D assets with the intent to inspire more prompts with a novel output. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for replying to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /HWEI-MIN LU/Primary Examiner, Art Unit 2142
Read full office action

Prosecution Timeline

May 31, 2023
Application Filed
Feb 11, 2026
Non-Final Rejection mailed — §101, §103, §112
Apr 30, 2026
Examiner Interview Summary
Apr 30, 2026
Applicant Interview (Telephonic)
Jun 11, 2026
Response Filed
Aug 17, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749017
MACHINE LEARNING EVALUATION FOR DETECTING FEATURE BIAS
3y 9m to grant Granted Sep 29, 2026
Patent 12737096
DRAWER PAGE OVERLAY FOR MULTITASKING
2y 3m to grant Granted Sep 15, 2026
Patent 12718096
ENHANCED DISCRIMINATE FEATURE LEARNING DEEP RESIDUAL CNN FOR MULTI-TASK ROTATING MACHINERY FAULT DIAGNOSIS WITH INFORMATION FUSION
3y 5m to grant Granted Aug 25, 2026
Patent 12705533
SYSTEMS AND METHODS FOR IMPROVING PREDICTION PROCESS USING AUTOMATED RULE LEARNING FRAMEWORK
3y 8m to grant Granted Aug 11, 2026
Patent 12700003
SYSTEMS AND METHODS FOR FREQUENT MACHINE LEARNING MODEL RETRAINING AND RULE OPTIMIZATION
4y 2m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
63%
Grant Probability
99%
With Interview (+40.2%)
2y 11m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 240 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month