Prosecution Insights
Last updated: October 02, 2026
Application No. 18/742,627

DYNAMIC EVALUATION SYSTEM FOR RESPONSIBLE AI IN LARGE LANGUAGE MODELS

Final Rejection §103
Filed
Jun 13, 2024
Examiner
MANOHARAN, SHASHIDHAR SHANKAR
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
80%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
4 granted / 5 resolved
+18.0% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
26 currently pending
Career history
33
Total Applications
across all art units

Statute-Specific Performance

§101
17.2%
-22.8% vs TC avg
§103
64.8%
+24.8% vs TC avg
§102
4.7%
-35.3% vs TC avg
§112
9.4%
-30.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 5 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendments filed 07/09/2026 have been accepted and considered in this office action. Claims 1, 6-7, 13, 17, and 18 have been amended. Claim 8 has been cancelled. Claims 1-7, 9-20 are pending. Response to Arguments Applicant’s arguments with respect to 35 U.S.C. 101 rejection of claims 1-7 and 9-20 have been considered and are persuasive after amendments. The 35 U.S.C. 101 rejection has been withdrawn. The limitations recite generating embeddings for simulated conversations, performing cluster analysis of the embeddings, generative diversity and coverage metrics based on the cluster analysis, and using the resulting feedback to adjust conversation and/or language model parameters for subsequent simulated conversations. When considered as a whole, these additional limitations meaningfully apply the recited RAI compliance evaluation in a specific computer-implemented feedback process rather than merely performing the abstract evaluation on a generic computer. Applicant’s arguments with respect to 35 U.S.C. 103 rejection of claims 1-7 and 9-20 have been considered but are moot in view of new grounds of rejection necessitated by the applicant’s amendments to the claims. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-7, 9-12, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Akkiraju et al. (hereinafter Akkiraju) (US 11190464 B2) (Paragraph numbers from attached copy) in view of Larson et al. (hereinafter Larson) (US 20200193331 A1), in further view of Gurgu et al. (hereinafter Gurgu) (US 20230297887 A1). Regarding claim 1, Akkiraju discloses: A system for testing responsible artificial intelligence (RAI) compliance of a language model (LM) based application, the system comprising (Akkiraju, P[3]: “The method may additionally include one or more processors determining an assessment of the performance of the customer service agent based on the interaction between the customer service agent and the chatbot.” (determining an assessment of the performance reads on testing compliance of the LM based application), P[22]: “Customer service agents may be trained to identify the fourth customer chatbot and to apply company anti-fraud policies.”, P[7]: “Many organizations spend tremendous resources to provide extensive training to their agents pertaining to products, services, and guidelines for dealing with concerns and requests from customers.” (anti-fraud policies and guidelines for dealing with concerns of customers reads on RAI compliance): at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: (Akkiraju, P[47]: "computer processor(s) 404, memory 406, persistent storage 408" and P[50]: "Integrated learning environment program 106 is stored in persistent storage 408 for execution by one or more of the respective computer processors 404," reads on a system with a processor and memory configured to perform program operations.) generating a first conversational-input prompt requesting a first conversational input, wherein the first conversational input prompt includes at least a portion of the initial conversation parameters (Akkiraju, P[39]: "determines a chatbot behavior and responses using a customer service interaction data and the persona" and P[40]: "initiates the training by sending the message 'My flight is Delayed!'," reads on generating a prompt or initial message for the simulation based on the received persona parameters.) providing the generated first conversational-input prompt as input to a language model (Akkiraju, P[42]: "encoder 604 is a neural network that maps a variable-length input sequence 602 to fixed-length vector 606," reads on providing an input sequence to a neural network/language model component of the simulation.) receiving, from the language model in response to the first conversational-input prompt, the requested first conversational input (Akkiraju, P[24]: "customer chatbots module 108 may generate multiple responses and select at least one response according to the target style, tone, or persona," reads on receiving the generated conversational output from the model based on the prompt.) transmitting the first conversational input to the LM-based application (Akkiraju, P[42]: "a simulated customer sends variable-length input sequence 602" and "to the customer service agent," reads on sending the generated input to the target application being tested/trained.) receiving a first conversational response from the LM-based application in response to the first conversational input (Akkiraju, P[40]: "Ben sends the message 'I am sorry to hear that. I am here to help. What's your flight number?' in reply," reads on receiving the response from the target application/agent.) storing the first conversational input and the first conversational response as a first simulated conversation (Akkiraju, P[14]: "Database 112 is a repository for data" and "Customer service interaction data may include text and/or speech conversations between customers and customer service agents," reads on recording the input/response exchange in a repository.) generating feedback based on the stored first simulated conversation (Akkiraju, P[25]: "Feedback generation module 110 provides feedback" and "by continuous assessment of the style and performance of the interaction," reads on generating feedback based on the analysis of the recorded conversation.) wherein generating the feedback comprises: generating an embedding for each of the first simulated conversation (Akkiraju, Abstract: "determining an assessment of the performance of the customer service agent based on the interaction between the customer service agent and the chatbot. The method additionally includes generating feedback for the customer service agent based on the assessment of the performance of the customer service agent.", teaches generating feedback based upon the simulated conversational interaction)]; adjusting one or more of the initial conversational parameters based on the generated feedback (Akkiraju, P[34]: "customer chatbots module 108 adjusts simulated customer styles and tasks based on the evaluation results" and "by feedback generation module 110," reads on modifying the conversation parameters/styles based on the feedback results.) generating a second conversational-input prompt requesting a second conversational input, wherein the second conversational input prompt includes at least a portion of the adjusted conversation parameters (Akkiraju, P[41]: "may continue processing at operation 260 to determine a chatbot behavior and responses" and "based on the interaction between the chatbot and the customer service agent," (reads on continuing the simulation for subsequent turns, which inherently includes the steps of generating a second prompt based on the adjusted behavior, providing it to the model, receiving/transmitting the subsequent turn's inputs and responses, and storing the resulting conversation similar to the first iteration), Akkiraju, P[39]: "determines a chatbot behavior and responses using a customer service interaction data and the persona" and P[40]: "initiates the training by sending the message 'My flight is Delayed!'," reads on generating a prompt or initial message for the simulation based on the received persona parameters.) providing the generated second conversational-input prompt as input to the language model (Akkiraju, P[42]: "encoder 604 is a neural network that maps a variable-length input sequence 602 to fixed-length vector 606," reads on providing an input sequence to a neural network/language model component of the simulation.) receiving, from the language model in response to the second conversational-input prompt, the requested second conversational input (Akkiraju, P[24]: "customer chatbots module 108 may generate multiple responses and select at least one response according to the target style, tone, or persona," reads on receiving the generated conversational output from the model based on the prompt.) transmitting the second conversational input to the LM-based application (Akkiraju, P[42]: "a simulated customer sends variable-length input sequence 602" and "to the customer service agent," reads on sending the generated input to the target application being tested/trained.) receiving a second conversational response from the LM-based application in response to the second conversational input (Akkiraju, P[40]: "Ben sends the message 'I am sorry to hear that. I am here to help. What's your flight number?' in reply," reads on receiving the response from the target application/agent.) storing the second conversational input and the second conversational response as a second simulated conversation (Akkiraju, P[14]: "Database 112 is a repository for data" and "Customer service interaction data may include text and/or speech conversations between customers and customer service agents," reads on recording the input/response exchange in a repository.) and evaluating an RAI compliance of the LM-based application based on at least whether the second simulated conversation violated the RAI guideline (Akkiraju, P[22]: "identify the fourth customer chatbot" and "to apply company anti-fraud policies,", Akkiraju, P[42]: “generates a score for Ben’s performance…generates a score of 0.8 for the reply provided by Ben” (reads on evaluating the performance/compliance of the application by determining if the interaction violated established policies or guidelines, anti-fraud reads on RAI.) Akkiraju does not explicitly disclose: receiving initial conversation parameters including at least one or more initial persona settings for a simulated user, an RAI guideline, and a configuration setting for the language model; generating feedback based on the stored first simulated conversation, wherein generating the feedback comprises: performing a cluster analysis with the embedding and other embeddings for simulated conversations by grouping the embeddings that are within a threshold distance from one another in an embedding space; and generating a diversity and coverage metric based on the cluster analysis; wherein the adjustment include adjusting the configuration setting for the language model, and wherein at least one conversation parameter is adjusted responsive to the diversity and coverage metric; However, Larson discloses: performing a cluster analysis with the embedding and other embeddings for simulated conversations by grouping the embeddings that are within a threshold distance from one another in an embedding space (Larson, P[0028]: "the density of the plurality of distinct instances relates to a cluster or a grouping of distinct instances of training data of the corpus of raw machine learning training data in which each distinct instance of training data is within a predetermined distance of another distinct instance of training data within the cluster or the grouping", Larson teaches forming/analyzing a cluster of vector-represented instances based on the instances being within a predetermined distance of one another); and generating a diversity and coverage metric based on the cluster analysis (Larson, P[0099]-P[0101], clustering, [P[0034]: "Calculating one or more efficacy metrics of the joint corpus of training data, wherein calculating the one or more efficacy metrics includes calculating one or more of a coverage metric value and a diversity metric value of the joint corpus of training data", when read with Larson's vector-space cluster analysis, Larson teaches calculating coverage and diversity metrics for the corpus characterized by that cluster analysis); and wherein at least one conversation parameter is adjusted responsive to the diversity and coverage metric (Larson, P[0034]: "sourcing additional re-seeding corpora of raw machine learning training data until a coverage metric threshold and/or a diversity metric threshold of a resulting joint corpus of raw machine learning training data is satisfied", Larson teaches modifying subsequent generation/sourcing responsive specifically to the calculated coverage and diversity metrics); It would have been prima facie obvious to one of ordinary art before the effective filing date of the claimed invention to have modified Akkiraju in view of Larson. Doing so would have provided Larson’s embedding-based clustering and diversity/coverage metrics to systematically evaluate simulated conversational data (Larson, Abstract, P[0034]) with Akkiraju’s simulated conversational interacts and feedback based adjustment (Akkiraju, Abstract) to systematically evaluate the simulated conversations, thus, improving the diversity and coverage of subsequent simulated interactions. The combination of Akkiraju and Larson does not explicitly disclose: However, Gurgu discloses: receiving initial conversation parameters including at least one or more initial persona settings for a simulated user, an RAI guideline, and a configuration setting for the language model (Gurgu, P[0111]: "the hyperparameters that may be controlled include, for example, temperature, top-k and top-p.", Gurgu teaches configuration settings controlling the generative language model); adjusting one or more of the initial conversational parameters based on the generated feedback, wherein the adjustment include adjusting the configuration setting for the language model (Gurgu, P[0097]: "A parameter of the large language model may be altered based on the one or more metrics. For example, a hyperparameter of the language model may be altered.", Gurgu teaches adjusting an LM configuration parameter responsive to generated feedback/metrics) It would have been prima facie obvious to one of ordinary art before the effective filing date of the claimed invention to have modified Akkiraju in view of Larson and in further view of Gurgu. Doing so would have provided the adjustable language-model configuration parameters of Gurgu (Gurgu, Abstract, P[0097], P[0111]) and Larson’s embedding-based clustering and diversity/coverage metrics to systematically evaluate simulated conversational data (Larson, Abstract, P[0034]) with Akkiraju’s simulated conversational interacts and feedback based adjustment (Akkiraju, Abstract). This would have enabled the generated metrics to guide adjustment of the language model for subsequent simulated conversations, improving the diversity and coverage of the generated interactions. Regarding claim 2, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 1. The combination further discloses: The initial conversation parameters further include a system description of the LM-based application; and the first conversational-input prompt further includes the system description (Akkiraju, P[14]: "Customer service interaction data may further include contextual information," and P[39]: "include models to directly consider the persona, task, and context as additional constraints," and P[40]: "the task scenario includes an airline customer in an international flight schedule where the first flight has been delayed," reads on providing a description of the system environment (the airline flight schedule and delay parameters) as part of the initial simulation parameters.) Regarding claim 3, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 1. The combination further discloses: The initial conversation parameters further include an adversarial seed including at least one of an example topic, example query, or example conversation; and the first conversational-input prompt further includes the adversarial seed (Akkiraju, P[14]: "Customer service interaction data may include text and/or speech conversations," and P[37]: "derived from the customer service interaction data by performing unsupervised machine learning methods to extract tone and persona styles," reads on using specific examples from prior conversations to seed the current simulation. Text and/or speech conversations reads on example conversation and tone and persona styles reads on example topic) Regarding claim 4, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 3. The combination further discloses: The adversarial seed includes at least a portion of a logged prior conversation with the LM-based application (Akkiraju, P[14]: "Customer service interaction data may include text and/or speech conversations between customers and customer service agents," and "collected from customer service call centers, social media, and/or any other source," reads on using logged historical interactions as the basis for the simulation seed.) Regarding claim 5, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 1. The combination further discloses: The first simulated conversation is stored in a first set of simulated conversations, and the second simulated conversation is stored in a second set of simulated conversations (Akkiraju, P[14]: "Database 112 is a repository for data," and P[35]: "store the style and performance of a customer service agent with respect to different types of customers and task scenarios during the training process," reads on categorizing and storing interactions into distinct data sets for analysis.) Regarding claim 6, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 1. The combination further discloses: Generating the feedback includes generating one or more metrics for the first set of simulated conversations (Akkiraju, P[41]: "determining the similarity... by automatic metrics," and P[42]: "generates a score for Ben's performance," reads on calculating quantitative metrics based on the stored conversation sets.) Regarding claim 7, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 6. The combination further discloses: wherein the metrics further include at least one of a relevance metric and an adversarial metric (Akkiraju, P[0032]: "similarity between the responses generated by the model and the responses provided by the agent by automatic metrics (e.g., BLEU, Tf-idf, word2vec)," reads on the relevance metric; Akkiraju, P[0022]: "identify the fourth customer chatbot" and "to apply company anti-fraud policies," reads on an adversarial metric used to detect deceptive or non-compliant behavior)). Regarding claim 9, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 5. The combination further discloses: Evaluating an RAI compliance includes evaluating whether each simulated conversation in the second set of conversations violated the RAI guideline (Akkiraju, P[31]: "evaluates the performance of an agent ranging from utterance level to task level," and P[34]: "adjusts simulated customer styles and tasks based on the evaluation results," reads on assessing a full set of interactions against the required guidelines.) Regarding claim 10, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 1. The combination further discloses: Generating an evaluation prompt including the RAI guideline, the second simulated conversation, and an instruction for the language model to determine if the second simulated conversation violated the RAI guidelines (Akkiraju, P[33]: "provide multiple responses" "shown as a multiple-choice question" and "ask the customer service agent to select the best response," reads on generating an evaluative prompt with instructions to judge the content against the defined scenario/guideline.) providing the evaluation prompt to the language model (Akkiraju, P[42]: "encoder 604 is a neural network that maps a variable-length input sequence 602" and "simulated customer sends variable-length input sequence 602," reads on providing the generated evaluative sequence to a language model/neural network component.) receiving, from the language model in response to the evaluation prompt, a response indicating whether the second simulated conversation violated the RAI guideline (Akkiraju, P[44]: "identifies the high score for the provided reply" and P[22]: "apply company anti-fraud policies," reads on receiving an output from the model indicating whether the interaction complied with or violated the established policies/guidelines, anti-fraud reads on RAI.) Regarding claim 11, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 1. The combination further discloses: The LM-based application is a chatbot (Akkiraju, P[1]: "customer care training using chatbots," and P[8]: "training may be performed by using a chatbot," reads on the application under test being a chatbot.) Regarding claim 12, the combination of Akkiraju, Larson, and Gurgu discloses the system described in claim 1. The combination further discloses: The persona settings include a setting for at least one of a conscientiousness trait, an openness trait, an extraversion trait, a neuroticism trait, or an agreeableness trait (Akkiraju, P[18]: "unsupervised machine learning methods to extract tone and persona styles," P: "enthusiastic when communicating," and P[28]: "attempt to avoid confrontation," reads on assigning multiple specific behavioral traits associated with an extraversion trait (being enthusiastic) and an agreeableness trait (avoiding confrontation) to the simulated persona.) Regarding claim 18, the combination of Akkiraju and Larson discloses the system described in claim 17. The combination further discloses: adjusting the conversation parameters includes adjusting at least one of the configuration settings for the language model (Akkiraju, P[34]: "customer chatbots module 108 adjusts simulated customer styles and tasks" and "increases the uncertainty in customer responses simulated by chatbots," reads on the functional act of modifying model generation behavior based on performance.). The combination does not explicitly disclose: the conversation parameters further include configuration settings for a language model used for generating the first set of simulated conversations, including at least one of a top- p value, a top-k value, or a temperature; However, Gurgu discloses: the conversation parameters further include configuration settings for a language model used for generating the first set of simulated conversations, including at least one of a top- p value, a top-k value, or a temperature (Gurgu, P[0128]: "the evaluation system 304 sets a test value for a parameter or hyperparameter (collectively referenced as parameter) of the large language model", P[0133]: "the evaluation system 304 selects the value for the parameter resulting in the highest metrics (e.g., highest aggregate of the UGR, MSAV, and MSAR metrics), as the final value for the corresponding parameter of the large language model.", Gurgu teaches changing/testing the claimed LM configuration settings and selecting an adjusted setting according to the resulting metrics); It would have been prima facie obvious to one of ordinary art before the effective filing date of the claimed invention to have modified Akkiraju in view of Larson and in further view of Gurgu. Doing so would have provided the adjustable language-model configuration parameters of Gurgu (Gurgu, Abstract, P[0097], P[0111]) and Larson’s embedding-based clustering and diversity/coverage metrics to systematically evaluate simulated conversational data (Larson, Abstract, P[0034]) with Akkiraju’s simulated conversational interacts and feedback based adjustment (Akkiraju, Abstract). This would have enabled the generated metrics to guide adjustment of the language model for subsequent simulated conversations, improving the diversity and coverage of the generated interactions. Claims 13-17 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Akkiraju et al. (hereinafter Akkiraju) (US 11190464 B2) (Paragraph numbers from attached copy) in view of Larson et al. (hereinafter Larson) (US 20200193331 A1). Regarding claim 13, Akkiraju discloses: A computer-implemented method for testing responsible artificial intelligence (RAI) compliance of a language model (LM) based application, the method comprising: (Akkiraju, Abstract) generating feedback for a first set of simulated conversations with the LM-based application (Akkiraju, P[0025]: "Feedback generation module 110 provides feedback" and "by continuous assessment of the style and performance of the interaction," reads on generating feedback based on a recorded interaction set.) based on the feedback, adjusting one or more conversation parameters, wherein the conversation parameters include one or more persona settings for a simulated user (Akkiraju, P[0034]: "customer chatbots module 108 adjusts simulated customer styles and tasks based on the evaluation results," reads on modifying the persona behavioral settings based on previous turns.) generating a second set of simulated conversations, wherein generating the second set of simulated conversations comprises (Akkiraju, P[0041]: "integrated learning environment program 106 may continue processing at operation 260 to determine a chatbot behavior and responses based on the interaction between the chatbot and the customer service agent," reads on generating a subsequent turn or set of interactions by returning to the generation process after an initial assessment.): generating a conversational-input prompt requesting a conversational input, wherein the conversational input prompt includes an RAI guideline and the persona settings (Akkiraju, P[0039]: "determines a chatbot behavior and responses using a customer service interaction data and the persona" and P[0040]: "initiates the training by sending the message," reads on generating a subsequent turn prompt based on the persona and task guidelines.) providing the generated conversational-input prompt as input to a language model; receiving, from the language model in response to the conversational-input prompt, the requested conversational input (Akkiraju, P[0042]: "encoder 604 is a neural network that maps a variable-length input sequence 602 to fixed-length vector 606," and P[0024]: "generate multiple responses and select at least one response," reads on utilizing an internal model to generate the simulated user's input.) transmitting the conversational input to the LM-based application (Akkiraju, P[0042]: "a simulated customer sends variable-length input sequence 602 (e.g., 'My flight is Delayed!') to the customer service agent," reads on the act of sending the model-generated input sequence to the target agent/application being tested.); receiving a conversational response from the LM-based application in response to the conversational input (Akkiraju, P[0042]: "a simulated customer sends variable-length input sequence 602" and "to the customer service agent," reads on sending the input to the target application and receiving the reply.) storing the conversational response as part of a simulated conversation of the second set of simulated conversations (Akkiraju, P[0014]: "Database 112 is a repository for data" and "include text and/or speech conversations," reads on recording the subsequent interaction in a repository.) and evaluating an RAI compliance of the second set of simulated conversations (Akkiraju, P[0022]: "identify the fourth customer chatbot" and "to apply company anti-fraud policies," reads on evaluating the performance of the interaction set against established guidelines.) Akkiraju does not explicitly disclose: wherein generating the feedback comprises: generating embeddings for the first set of simulated conversations; performing a cluster analysis of the generated embeddings; and generating a diversity and coverage metric based on the cluster analysis; However, Larson discloses: wherein generating the feedback comprises: generating embeddings for the first set of simulated conversations (Larson, P[0081]: "sentence embedding may include a set of techniques that map instances of sentences, words, and/or phrases identified within the corpus of training data to vectors of real numbers.", P[0082]: "generate a vector representation for each instance within a corpus of training data.", Larson teaches generating embeddings for the respective instances comprising the first set/corpus); performing a cluster analysis of the generated embeddings (Larson, P[0028]: "density of the plurality of distinct instances relates to a cluster or a grouping of distinct instances of training data of the corpus of raw machine learning training data in which each distinct instance of training data is within a predetermined distance of another distinct instance of training data within the cluster or the grouping", because the instance shave previously been mapped to vector representations, Larson teaches cluster analysis of the generated embeddings); and generating a diversity and coverage metric based on the cluster analysis (Larson, P[0034]: "calculating one or more efficacy metrics of the joint corpus of training data, wherein calculating the one or more efficacy metrics includes calculating one or more of a coverage metric value and a diversity metric value of the joint corpus of training data", Larson teaches calculating both claimed types of metrics for the corpus subjected to the preceding vector/cluster analysis); It would have been prima facie obvious to one of ordinary art before the effective filing date of the claimed invention to have modified Akkiraju in view of Larson. Doing so would have provided Larson’s embedding-based clustering and diversity/coverage metrics to systematically evaluate simulated conversational data (Larson, Abstract, P[0034]) with Akkiraju’s simulated conversational interacts and feedback based adjustment (Akkiraju, Abstract) to systematically evaluate the simulated conversations, thus, improving the diversity and coverage of subsequent simulated interactions. Regarding claim 14, the combination of Akkiraju and Larson discloses the computer-implemented method of claim 13. The combination further discloses: the persona settings include settings for at least two of a conscientiousness trait, an openness trait, an extraversion trait, a neuroticism trait, or an agreeableness trait (Akkiraju, P[18]: "unsupervised machine learning methods to extract tone and persona styles," P[28]: "enthusiastic when communicating," and P[19]: "attempt to avoid confrontation," reads on assigning multiple specific behavioral traits associated with an extraversion trait (being enthusiastic when communicating) and an agreeableness trait (avoiding confrontation) to the simulated persona.) Regarding claim 15, the combination of Akkiraju and Larson discloses the computer-implemented method of claim 13. The combination further discloses: Generating an evaluation prompt including the RAI guideline, the second simulated conversation, and an instruction for the language model to determine if the second simulated conversation violated the RAI guidelines (Akkiraju, P[33]: "provide multiple responses" "shown as a multiple-choice question" and "ask the customer service agent to select the best response," reads on generating an evaluative prompt with instructions to judge the content against the defined scenario/guideline.) providing the evaluation prompt to the language model (Akkiraju, P[42]: "encoder 604 is a neural network that maps a variable-length input sequence 602" and "simulated customer sends variable-length input sequence 602," reads on providing the generated evaluative sequence to a language model/neural network component.) receiving, from the language model in response to the evaluation prompt, a response indicating whether the second simulated conversation violated the RAI guideline (Akkiraju, P[44]: "identifies the high score for the provided reply" and P[22]: "apply company anti-fraud policies," reads on receiving an output from the model indicating whether the interaction complied with or violated the established policies/guidelines, anti-fraud reads on RAI.) Regarding claim 16, the combination of Akkiraju and Larson discloses the computer-implemented method of claim 15. The combination further discloses: Generating an RAI compliance score based on whether the conversational response violated the RAI guideline (Akkiraju, P[41]: "determining the similarity" "by automatic metrics," and P[42]: "generates a score for Ben's performance," reads on producing a numerical assessment of adherence to guidelines.) Regarding claim 17, Akkiraju discloses: a system for testing responsible artificial intelligence (RAI) compliance of a language model (LM) based application, the system comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the system to perform operations comprising: (Akkiraju, P[0047]: "computer processor(s) 404, memory 406, persistent storage 408" and P[0050]: "stored in persistent storage 408 for execution by one or more of the respective computer processors 404," reads on a system with a processor and memory configured to perform program operations.); receiving initial conversation parameters, wherein the initial conversation parameters comprise: an RAI guideline; one or more persona settings for a simulated user; and a system description of the LM-based application (Akkiraju, P[0037]: "integrated learning environment program 106 receives a customer persona" and P[0007]: "guidelines for dealing with concerns" and P[0014]: "Customer service interaction data may further include contextual information," reads on receiving initial parameters including personas, guidelines, and context or descriptions of the system environment.); generating a first set of simulated conversations, with the LM-based application, based on the initial conversation parameters (Akkiraju, P[0039]: "determines a chatbot behavior and responses using a customer service interaction data and the persona" and P[0040]: "initiates the training by sending the message," reads on generating an initial set of simulated exchanges based on the input persona and task.); generating feedback based on the first set of simulated conversations (Akkiraju, P[0025]: "Feedback generation module 110 provides feedback" and "by continuous assessment of the style and performance of the interaction," reads on generating feedback derived from the analysis of a recorded conversation set.); adjusting one or more of the conversation parameters based on the generated feedback (Akkiraju, P[0034]: "customer chatbots module 108 adjusts simulated customer styles and tasks based on the evaluation results," reads on modifying behavioral or task parameters based on the assessed results of previous interactions.); generating a second set of simulated conversations, with the LM-based application, based on the adjusted conversation parameters (Akkiraju, P[0041]: "may continue processing at operation 260 to determine a chatbot behavior and responses based on the interaction between the chatbot and the customer service agent," reads on generating a subsequent set of interactions after the parameters have been modified.); and evaluating an RAI compliance of the LM-based application based on whether the second set of simulated conversations violated the RAI guideline (Akkiraju, P[0022]: "identify the fourth customer chatbot" and "to apply company anti-fraud policies," reads on evaluating the compliance of the interaction set by determining if a specific policy or guideline violation occurred.). Akkiraju does not explicitly disclose: wherein generating the feedback comprises: generating embeddings for the first set of simulated conversations; performing a cluster analysis of the generated embeddings; and generating a diversity and coverage metric based on the cluster analysis; However, Larson discloses: wherein generating the feedback comprises: generating embeddings for the first set of simulated conversations (Larson, P[0081]: "sentence embedding may include a set of techniques that map instances of sentences, words, and/or phrases identified within the corpus of training data to vectors of real numbers.", P[0082]: "generate a vector representation for each instance within a corpus of training data.", Larson teaches generating embeddings for the respective instances comprising the first set/corpus); performing a cluster analysis of the generated embeddings (Larson, P[0028]: "density of the plurality of distinct instances relates to a cluster or a grouping of distinct instances of training data of the corpus of raw machine learning training data in which each distinct instance of training data is within a predetermined distance of another distinct instance of training data within the cluster or the grouping", because the instance shave previously been mapped to vector representations, Larson teaches cluster analysis of the generated embeddings); and generating a diversity and coverage metric based on the cluster analysis (Larson, P[0034]: "calculating one or more efficacy metrics of the joint corpus of training data, wherein calculating the one or more efficacy metrics includes calculating one or more of a coverage metric value and a diversity metric value of the joint corpus of training data", Larson teaches calculating both claimed types of metrics for the corpus subjected to the preceding vector/cluster analysis); It would have been prima facie obvious to one of ordinary art before the effective filing date of the claimed invention to have modified Akkiraju in view of Larson. Doing so would have provided Larson’s embedding-based clustering and diversity/coverage metrics to systematically evaluate simulated conversational data (Larson, Abstract, P[0034]) with Akkiraju’s simulated conversational interacts and feedback based adjustment (Akkiraju, Abstract) to systematically evaluate the simulated conversations, thus, improving the diversity and coverage of subsequent simulated interactions. Regarding claim 19, the combination of Akkiraju and Larson discloses the computer-implemented method of claim 17. The combination further discloses: Generating an evaluation prompt including the RAI guideline, the second simulated conversation, and an instruction for the language model to determine if the second simulated conversation violated the RAI guidelines (Akkiraju, P[33]: "provide multiple responses" "shown as a multiple-choice question" and "ask the customer service agent to select the best response," reads on generating an evaluative prompt with instructions to judge the content against the defined scenario/guideline.) providing the evaluation prompt to the language model (Akkiraju, P[42]: "encoder 604 is a neural network that maps a variable-length input sequence 602" and "simulated customer sends variable-length input sequence 602," reads on providing the generated evaluative sequence to a language model/neural network component.) receiving, from the language model in response to the evaluation prompt, a response indicating whether the second simulated conversation violated the RAI guideline (Akkiraju, P[44]: "identifies the high score for the provided reply" and P[22]: "apply company anti-fraud policies," reads on receiving an output from the model indicating whether the interaction complied with or violated the established policies/guidelines, anti-fraud reads on RAI.) Regarding claim 20, the combination of Akkiraju and Larson discloses the computer-implemented method of claim 17. The combination further discloses: The initial conversation parameters are received through a configuration interface (Akkiraju, P[16]: "User interface 116" "display text, documents, web browser windows, user options," and "control sequences the user employs to control the program," reads on receiving testing parameters through a GUI/interface.) Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHASHIDHAR S MANOHARAN whose telephone number is (571)272-6772. The examiner can normally be reached M-F 8:00-4:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SHASHIDHAR SHANKAR MANOHARAN/ Examiner, Art Unit 2655 /ANDREW C FLANDERS/ Supervisory Patent Examiner, Art Unit 2655
Read full office action

Prosecution Timeline

Jun 13, 2024
Application Filed
Feb 09, 2026
Non-Final Rejection mailed — §103
Jul 08, 2026
Examiner Interview Summary
Jul 08, 2026
Applicant Interview (Telephonic)
Jul 09, 2026
Response Filed
Sep 24, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737537
GOVERNANCE AND CONFIDENCE ASSESSMENT OF LLM
2y 6m to grant Granted Sep 15, 2026
Patent 12682890
MASK-CONFORMER AUGMENTING CONFORMER WITH MASK-PREDICT DECODER UNIFYING SPEECH RECOGNITION AND RESCORING
2y 4m to grant Granted Jul 14, 2026
Patent 12682173
MODULAR FRAMEWORK FOR EVALUATING LANGUAGE MODELS
2y 4m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 3 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+33.3%)
2y 2m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 5 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month