Prosecution Insights
Last updated: October 02, 2026
Application No. 19/014,909

TRAINING DIALOGUE SYSTEMS TO AVOID HARMFUL CONTENT BY USING LANGUAGE MODELS

Non-Final OA §103
Filed
Jan 09, 2025
Examiner
SUBRAMANI, NANDINI
Art Unit
2656
Tech Center
2600 — Communications
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
65%
Grant Probability
Moderate
1-2
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% of resolved cases
65%
Career Allowance Rate
64 granted / 99 resolved
+2.6% vs TC avg
Strong +48% interview lift
Without
With
+47.5%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
15 currently pending
Career history
114
Total Applications
across all art units

Statute-Specific Performance

§101
12.9%
-27.1% vs TC avg
§103
65.3%
+25.3% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
8.5%
-31.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 99 resolved cases

Office Action

§103
DETAILED ACTION Introduction Applicant's submission filed on 01/09/2025 has been entered. Claims 1-20 are pending in the application and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 8-14 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Dinan, Emily, et al. "Build it break it fix it for dialogue safety: Robustness from adversarial human attack." Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019 (cited in IDS) in view of Tiwari, US PgPub. 2025/0335817. Regarding claim 1, Dinan teaches a method comprising: inputting a test case, PNG media_image1.png 540 368 media_image1.png Greyscale comprising one or more natural language statements, into a dialogue management system (DMS), where the DMS is configured to maximize a probability that processing the test case will result in an output by an artificial intelligence language model (LM) that contains harmful content ( Dinan, Fig. 1 Break it ( Round 1), Dinan pg. 4539, 1. Build it: Build a model capable of detecting OFFENSIVE messages ( trained BERT-based model, A0) 2. Break it: Ask crowdworkers to try to “beat the system” by submitting messages that our system (A0) marks as SAFE but that the worker considers to be OFFENSIVE(harmful)); processing, by the LM, the test case to generate a responsive natural language output based on a dialogue plan of the DMS (see Dinan, Fig. 1 A0 , pg. 4540 Models to Break During round 1,workerstryto break the baseline model A0, trained on Wikipedia Toxic Comments); analyzing the responsive natural language output to determine if the responsive natural language output comprises harmful content (Dinan, pg. 4539, Build it: Build a model capable of detecting OFFENSIVE messages. This is our best performing BERT-based model; We establish baselines using two models. The first one is a binary classifier built on top of a large pre-trained transformer model.; Fig. 1, A0 prediction of Offensive or Safe ( analyzing response) ); and in response to the responsive natural language output being determined to comprise harmful content (see Dinan, Fig, 1 prediction of A0 is predicted as SAFE while actually considered OFFENSIVE): generating a conversation log for the test case based on a context of the test case, wherein the context of the test case comprises a series of natural language inputs of the test case and natural language responses generated by the LM during processing of the test case (see Dinan, sect 6 Multi turn tasks with models using context dialogue safety we posit it is important to move beyond classifying single utterances, as it may be the case that an utterance is entirely innocuous on its own but extremely offensive in the context of the previous dialogue history, sect 6.3.1-2, Table 8 lists the conversation(multiturn task) that has been used in Break It Phase); and outputting the conversation log to a language model development computing system for retraining the LM to avoid generation of responses comprising harmful content(see Dinan sect 6.3.2 multiturn dialog used for adversarial task retraining, results in Table 10; Dinan, sect. 4.2 During the “fix it” round, we update the models with the newly collected adversarial data from the “break it” round. The training data consists of all previous rounds of data, so that model Ai is trained on all rounds n for n <= I; Fig. 1 Fix it ( Round 1) for retraining A0 + A1 ). Dinan teaches in response to the responsive natural language output being determined to comprise harmful content: generating a conversation log for the test case based on a context of the test case, wherein the context of the test case comprises a series of natural language inputs of the test case and natural language responses generated by the LM during processing of the test case, however to further compact prosecution and teach generating a conversation log for the test case based on a context of the test case, Tiwari is used to further teach in response to the responsive natural language output being determined to comprise harmful content: generating a conversation log for the test case based on a context of the test case, wherein the context of the test case comprises a series of natural language inputs of the test case and natural language responses generated by the LM during processing of the test case ( see Tiwari, [0030-0031] The undesirability of first conversational outputs 362 may be assessed using undesirable expression assessment model 350 in the form of an ML model trained to detect undesirable expressions. The training of dialogue model 360 to avoid one or both of hallucinations and undesirable expressions, in action 243, may be performed by hardware processor 104 of system 100, using ML model training pipeline 130 Tiwari [0014] discusses the contexts around conversations); and outputting the conversation log to a language model development computing system for retraining the LM to avoid generation of responses comprising harmful content (see Tiwari, Fig. 2, step 243 training the Dialogue model 360 to reducing the generation of toxic or otherwise undesirable expressions as well as to reducing hallucinations by dialogue model using reinforcement learning to o provide guardrailed dialogue model 364. Further training can be included as indicated in Fig. 2 steps 248a/b ). Dinan and Tiwari are considered to be analogous to the claimed invention because both relate to training Dialogue Models based on detection of offensive language in the dialogue. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Dinan to retrain language models to detect offensive language across a variety of content classes with the training a large language model to project a consistent personality that is substantially guard railed against hallucination and the generation of toxic, offensive, or otherwise undesirable language teachings of Tiwari to provide an efficient and resource sparing solution for imbuing machine learning models (see Tiwari, [0002]). Regarding claim 2, Dinan in view of Tiwari teaches method of claim 1. Dinan further teaches wherein the harmful content is content that is one of offensive, inappropriate, or derogatory with regard to a user (see Dinan, pg. 4540 Crowderworker Task We ask crowdworkers to try to “beat the system” by submitting messages that our system marks as SAFE but that the worker considers to be OFFENSIVE). Tiwari also teaches wherein the harmful content is content that is one of offensive, inappropriate, or derogatory with regard to a user (see Tiwari, [0031] The undesirability of first conversational outputs 362 may be assessed using undesirable expression assessment model 350 in the form of an ML model trained to detect undesirable expressions at both the word level and a more abstract intent level, such as detecting sarcasm, bullying and the like, and to output undesirability score 352 corresponding to the undesirableness of each of first conversational outputs 362). The same motivation to combine as claim 1 applies here. Regarding claim 3, Dinan in view of Tiwari teaches method of claim 1. Dinan further teaches further comprising retraining the LM to avoid generation of responses comprising harmful content at least by generating training examples from the conversation log, wherein the training examples comprise ground truth data modified to not include harmful content (see Dinan, sect 7, The broken examples are used to train the models using the adversarial data which includes complex examples with less profanity, which existing classifiers can pick up on, and is instead offensive due to figurative language ,negation, offensive language in the context of the dialogue is nuanced ( table 8) and by requiring more world knowledge, which in turn trains the classifier to be more robust). Regarding claim 4, Dinan in view of Tiwari teaches method of claim 3. Dinan further teaches wherein retraining the LM further comprises populating user response and context information of one or more training examples from content of the conversation log(see Dinan sect 4.2 & Sect 1, to retrain/FIX IT In this work we instead fully automate such an approach using crowdworkers as the humans in-the-loop, and also apply a fixing stage where models are retrained to improve them. Finally, we repeat the whole build, break, and fix sequence over a number of iterations; Dinan sect 6.2-3 discusses model with contexts and break it phase with contexts,); submitting the one or more training examples to the LM to obtain a response; comparing the response to a ground truth response that the LM should generate that does not include harmful content, to determine a loss (see Dinan, sect 6.3.2 discusses Fix it Phases with context training and the results as indicated in table 10 with training tests included as in Table 9); and modifying at least one operational parameter of the LM to reduce the loss (see Dinan, sect 5.1.2 discusses weighting mult-tasking with mixing parameter which tuned based on the validation set. The Cross entropy loss is adjusted based on the final bias using the validation set and is optimized for the sensitive class (offensive class) metric on the standard and adversarial validation sets respectively; operational parameter of the LM/BERT based model ). Regarding claim 8, Dinan in view of Tiwari teaches method of claim 1. Dinan further teaches wherein analyzing the responsive natural language output to determine if the responsive natural language output comprises harmful content comprises: inputting the responsive natural language output to a trained machine learning classifier computer model that outputs a classification as to whether the responsive natural language output contains the harmful content or not (Dinan, pg. 4539 discusses models to include the binary classifier to the Bert, which is then used in the :”Build it “ model A0 (baseline classifier) classify if Offensive or Safe ( indicated in Fig. 1)); and in response to the responsive natural language output being classified as containing harmful content, generating the conversation log from content of a conversation of a dialogue system managed by the DMS. (see Dinan, sect. 4.2 During the “fix it” round, the model Ai is trained with previous rounds of data and the newly collected adversarial data from “break it” round and validated with robustness(do not result in harmful content due to user input) to new adversarial attacks as indicated as “Broken : add to new dataset” in Fig. 1 which is further extended for multi-turn dialogs( conversation); Dinan, Table 4). Regarding claim 9, Dinan in view of Tiwari teaches method of claim 1. Dinan further teaches wherein the retrained LM prioritizes paths through a dialogue plan that do not result in the harmful content being generated in response to a user input (see Dinan, sect. 4.2 During the “fix it” round, the model Ai is trained with previous rounds of data and the newly collected adversarial data from “break it” round and validated with robustness(do not result in harmful content due to user input) to new adversarial attacks ). Regarding claim 10, Dinan in view of Tiwari teaches method of claim 8. Tiwari further teaches wherein the dialogue system is one of a chat bot or automated assistant conversation system (see Tiwari, Fig. 3 describes the Dialogue Model and Guardrailed Dialogue model based on personas ( digital assistants etc. as described in Tiwari [0011] ). The same motivation to combine as claim 1 applies here. Regarding claim 11, is directed to a computer program product claim corresponding to the method claim presented in claim 1 and is rejected under the same grounds stated above regarding claim 1. Regarding claim 12, is directed to a computer program product claim corresponding to the method claim presented in claim 2 and is rejected under the same grounds stated above regarding claim 2. Regarding claim 13, is directed to a computer program product claim corresponding to the system claim presented in claim 3 and is rejected under the same grounds stated above regarding claim 3. Regarding claim 14, is directed to a computer program product claim corresponding to the system claim presented in claim 4 and is rejected under the same grounds stated above regarding claim 4. Regarding claim 18, is directed to a computer program product claim corresponding to the system claim presented in claim 8 and is rejected under the same grounds stated above regarding claim 8. Regarding claim 19, is directed to a computer program product claim corresponding to the system claim presented in claim 9 and is rejected under the same grounds stated above regarding claim 9. Regarding claim 20, is directed to a computer system claim corresponding to the method claim presented in claim 1 and is rejected under the same grounds stated above regarding claim 1. Claims 5 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Dinan, Emily, et al. "Build it break it fix it for dialogue safety: Robustness from adversarial human attack." Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019 (cited in IDS) in view of Tiwari, US PgPub. 2025/0335817 further in view of Coursey US Patent 12,210,849. Regarding claim 5, Dinan in view of Tiwari teaches method of claim 1. However, Dinan in view of Tiwari fail to teach further comprising determining the dialogue plan based on a state-space search algorithm that determines a set of actions to take to progress from a starting state to a goal state corresponding to the harmful content. However, Coursey teaches further comprising determining the dialogue plan based on a state-space search algorithm that determines a set of actions to take to progress from a starting state to a goal state corresponding to the harmful content (see Coursey, col 4 line 58-col 5 line 3 describes the MCTS (state space search) method to determine the search for an action; Coursey, Fig. 8 describes the sentiment analysis ( negative sentiment- harmful) and the goal oriented analysis to determine the goal oriented evaluation ( goal state corresponding to harmful content)). Dinan, Tiwari and Coursey are considered to be analogous to the claimed invention because both relate to training Dialogue Models based on detection of offensive language in the dialogue. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Dinan in view of Tiwari to retrain language models to detect offensive language across a variety of content classes with the AI language self-improvement agent using language modeling and tree search teachings of Coursey to virtual agent for language interactions which utilizes self-play learning to generate conversation logs from tree search processes in determining language utterances (see Coursey, col 1 lines 50-54). Regarding claim 15, is directed to a computer program product claim corresponding to the system claim presented in claim 5 and is rejected under the same grounds stated above regarding claim 5. Claims 6-7 and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Dinan, Emily, et al. "Build it break it fix it for dialogue safety: Robustness from adversarial human attack." Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019 (cited in IDS) in view of Tiwari, US PgPub. 2025/0335817 further in view of Coursey US Patent 12,210,849 further in view of Hansen, Eric A et. al. "LAO∗: A heuristic search algorithm that finds solutions with loops." Artificial Intelligence 129.1-2 (2001): 35-6. Regarding claim 6, Dinan in view of Tiwari further in view of Coursey teaches method of claim 5. However, Dinan in view of Tiwari further in view of Coursey fail to teach wherein the state-space search algorithm is a modified Loopy AO* algorithm in which the goal state is the harmful content. However, Hansen teaches wherein the state-space search algorithm is a modified Loopy AO* algorithm in which the goal state is the harmful content (see Hansen, pg. 46 describes LAO* with goal state ( goal state corresponding to harmful content)). Dinan in view of Tiwari further in view of Coursey teaches a Monte Carlo Tree Search state search space dialog system to create an AI language self-improvement virtual agent using language modeling and tree search techniques, however does not teach using Loopy AO* algorithm for the state space search algorithm. Hansen teaches using LAO* ( Loop AO*) algorithm, a heuristic search algorithm that can find optimal solutions for Markov decision processes (MDPs) without evaluating the entire state space. Using the known technique of LAO* as taught by Hansen, to provide the state-space search algorithm in the references Dinan in view of Tiwari further in view of Coursey to find heuristic search approach to find solutions with loops, such as detection of goal state which has the harmful content would have been obvious to one of ordinary skill in the art. Regarding claim 7, Dinan in view of Tiwari further in view of Coursey further in view of Hansen teaches method of claim 6. Hansen further teaches wherein the modified Loopy AO* algorithm implements an objective function F defined as follows: F(n) = 0 if n is a terminal node of the dialogue plan that has the harmful content (see Hansen, pg. 43 equation 10 (if i is a goal state : harmful content)); F(n) = 1 if n is a terminal node of the dialogue plan that does not have the harmful content(see Hansen, pg. 43 equation 10 (if i is a nonterminal tip state : does not have harmful content)); F(n) = min(F(sl), F(s2),...,F(sk)), where n is an OR node and where sl,...,sk are successor states(see Hansen, pg. 52, in Equation 11 the cost for an OR node For an OR node, only one child is selected. The cost is the minimum among all children which includes Cost from node nn to child and Evaluation cost of child ); F(n)=(F(sl) x p(sl))+...+(F(sk) x p(sk)), where n is an AND node (see Hansen, pg. 52, equation 13 For an AND node, all children must be solved. The total cost is the sum of all child costs. All successors contribute to the final cost and Every child must be included ); and p(si) is a probability of reaching a successor state si of state n(see Hansen, pg. 42 We let pij (a) denote the probability that taking action a in state i results in a transition to state j). The same motivation to combine as claim 6 applies here. Regarding claim 16, is directed to a computer program product claim corresponding to the system claim presented in claim 6 and is rejected under the same grounds stated above regarding claim 6. Regarding claim 17, is directed to a computer program product claim corresponding to the system claim presented in claim 7 and is rejected under the same grounds stated above regarding claim 7. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Misra, US Patent 12,210,849 teaches techniques for updating a LLM to cause generation of desired response which includes identifying a portion of the LLM, such as a layer of the LLM, that is responsible for (i.e., has the most impact in) generating an output by processing an input, where the output may be an undesired response for the input. An undesired response, as used herein, may refer to an incorrect response to an input, an inappropriate response (e.g., a biased, harmful, violent, etc. response), or otherwise a response that is not expected/desired (see Mishra, col 3 lines 11-30). Irving et. al. US PgPub. 2024/0104336 teaches a rule violation detection neural network can be used by the dialogue system to determine when a response violates one or more rules (see Irving, Fig. 5). Gupta et. al. US PgPub. 2024/0420453 teaches techniques for generating synthetic data for machine learning including undesired responses as target output (see Gupta, Fig. 4). Any inquiry concerning this communication or earlier communications from the examiner should be directed to NANDINI SUBRAMANI whose telephone number is (571)272-3916. The examiner can normally be reached Monday - Friday 12:00pm - 5:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh M Mehta can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NANDINI SUBRAMANI/ Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Jan 09, 2025
Application Filed
Jul 10, 2026
Non-Final Rejection mailed — §103
Sep 13, 2026
Interview Requested
Sep 22, 2026
Applicant Interview (Telephonic)
Sep 22, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749498
QUALITY ESTIMATION MODEL FOR PACKET LOSS CONCEALMENT
3y 9m to grant Granted Sep 29, 2026
Patent 12731575
Attention-Based Joint Acoustic and Text On-Device End-to-End Model
3y 7m to grant Granted Sep 08, 2026
Patent 12700416
Audio Transcoding Method and Apparatus, Audio Transcoder, Device, and Storage Medium
3y 9m to grant Granted Aug 04, 2026
Patent 12688863
EMOTIONALLY-AWARE VOICE RESPONSE GENERATION METHOD AND APPARATUS
4y 7m to grant Granted Jul 21, 2026
Patent 12670912
ATTENTIVE SCORING FUNCTION FOR SPEAKER IDENTIFICATION
2y 9m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+47.5%)
3y 0m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 99 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month