Prosecution Insights
Last updated: October 02, 2026
Application No. 18/745,550

GENERATIVE DIALOG MODEL TRAINING METHOD AND APPARATUS AS WELL AS GENERATIVE DIALOG IMPLEMENTING METHOD AND APPARATUS

Final Rejection §103
Filed
Jun 17, 2024
Priority
Jun 30, 2023 — CN 202310797318.6
Examiner
MEIS, JON CHRISTOPHER
Art Unit
2654
Tech Center
2600 — Communications
Assignee
Baidu Online Network Technology (Beijing) Co., Ltd.
OA Round
2 (Final)
33%
Grant Probability
At Risk
3-4
OA Rounds
7m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants only 33% of cases
33%
Career Allowance Rate
11 granted / 33 resolved
-28.7% vs TC avg
Strong +52% interview lift
Without
With
+52.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
17 currently pending
Career history
60
Total Applications
across all art units

Statute-Specific Performance

§101
21.7%
-18.3% vs TC avg
§103
55.9%
+15.9% vs TC avg
§102
12.5%
-27.5% vs TC avg
§112
9.2%
-30.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 33 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claims 1-2, 4-15, and 17-20 are pending. Claims 1, 13, and 14 are independent. This Application was published as US 20240338530. Apparent priority is 30 June 2023. The instant Application is directed to a method of training a dialog model to meet a safety specification. Applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection that, if presented, were necessitated by the amendments to the Claims. This action is Final. Response to Arguments 35 USC 102/3 Applicant’s arguments with respect to 35 USC 102 and 103 have been considered but not persuasive. In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., a greater proportion of inputs corresponding to the safety specification) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993). Applicant argues (pg. 11) that the instant invention deliberately skews data toward the desired class in order to reinforce learning of the new content. However, the claims only require that the proportion of data for the first-class dialog inputs is larger than that of second-class dialog inputs. The first-class inputs correspond to the update iteration, and the second-class inputs correspond to the pre-update iteration. While Beaver discloses (in one embodiment) adding samples in order to balance the distribution over classes, adding any samples results in an overall increase in the total number of samples. The claims as written limit the proportion of samples between updates, not the proportion of samples between classes. Therefore, the combination reads on the claim language, and the rejection is maintained. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-2, 4, 13-15, and 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chamarthy et al. (US 20220164606 A1) in view of Beaver (US 20200143794 A1). Regarding claim 1, Chamarthy discloses: 1. A generative dialog model training method, comprising: in response to determination of an update of a safety specification, taking an updated safety specification as a target safety specification, (Fig. 4 shows "COMPARE NEW FAIRNESS METRIC TO THE FAIRNESS THRESHOLD 420". The new fairness metric reads on an updated safety specification and the new metric is what is compared as the target. ) the update being performed on a previous safety specification when a generative dialog model after last optimization is determined not to meet a launch requirement, (Fig. 4 shows the steps continue to 420 after step 412 fails to meet the threshold (launch requirement).) wherein the safety specification comprises an evaluation specification of at least one evaluation dimension corresponding to different combinations respectively, ("[0028] Evaluating the fairness metric on the reference/monitored group as a whole provides an indication of whether erroneous inclinations exists or not based on the accepted ranges of fairness metrics indicating fairness or erroneous inclination…" – fairness and erroneous inclinations are both dimensions) any one combination consists of one content field and one application scenario, (see mapping of next two limitations) the content field is a safe content field involved by a generative dialog, and ("[0017] There are multiple techniques that exist to detect erroneous inclinations in machine learning models for both regression and classification type models. These techniques compute whether the machine learning model is exhibiting erroneous inclination against a monitored group as compared its counterpart a reference group. For example, a loan machine learning model may be erroneously inclined to give 90% of favorable outcomes to loan applicants within group B as compared to only 75% of favorable outcomes to group A applicants. If a fairness calculation is determined using a disparate impact ratio, then the metric would turn out to be 75/90=83%. If a machine learning model validator has a fairness threshold as 90% (meaning at least 90% of the monitored population should get 90% of favorable outcomes) then an inference may be made that the machine learning model is erroneous inclined against the group A." – each group reads on a content field) the application scenario is an application scenario of the generative dialog; ("[0044] The logic of the cognitive system implements the cognitive operation(s), examples of which include, but are not limited to, question answering, identification of related concepts within different portions of content in a corpus, intelligent search algorithms, such as Internet web page searches, for example, medical diagnostic and treatment recommendations, financial trend analysis, financial investment recommendations, credit scoring and credit/loan approval recommendations, and other types of recommendation generation, e.g., items of interest to a particular user, potential new contact recommendations, or the like..." – any of the applications listed reads on an application scenario) determining a dialog input corresponding to a current optimization according to the target safety specification, ("[0037]...The trained artificial intelligence based computing system or computer model may be implemented as part of or utilized by a cognitive computing system that employs the trained artificial intelligence (AI) based computing system or computer model to generate results upon which the cognitive computing system operates to generate cognitive computing responses to user requests, for example, e.g., answering natural language questions, performing image recognition, generating recommendations, decision support operations, or any other cognitive computing operation. The cognitive computing system may comprise any artificial intelligence based computing system that is trained through a machine learning process so as to generate results from given inputs, where the results have an acceptable level of error or loss after such training..." - The system optimizes for all inputs so any input corresponds to a current optimization.) comprising: obtaining a first dialog input set, and taking dialog inputs therein as the dialog inputs corresponding to the current optimization; (“[0083] Utilizing the identified set of user-defined monitored and reference groupings, the machine learning model data quality improvement detection engine looks at the monitored behavior data (step 406) and computes the percentage of favorable outcomes for the set of user-defined monitored and reference groupings (step 408)…) the first dialog input set at least comprises the dialog inputs corresponding to an updated combination, and (Fig. 4 shows step 422 updates the groups which would update the inputs corresponding to an updated combination) the first dialog input set meets the following predetermined condition: a number proportion of first-class dialog inputs is larger than that of second-class dialog inputs, the first-class dialog inputs are the dialog inputs corresponding to the updated combination, and the second-class dialog inputs are the dialog inputs corresponding to the combination without the update; and (not explicitly disclosed) optimizing the generative dialog model according to the dialog input and a principle that a reply generated by the generative dialog model conforms to the target safety specification, ("[0085] Having identified the new monitored group and new reference group, the machine learning model data quality improvement detection engine transmits the accurate identification of the monitored group and the reference group to an authorized computing system (step 426). In response to receiving the accurate identification of the monitored group and the reference group, the machine learning model data quality improvement detection engine reduces erroneous inclinations of the trained cognitive computing system and/or trained computer machine learning model to address the newly identified the accurate identification of the monitored group and the reference group (step 428). That is, utilizing the newly identified ranges for the monitored and reference groups, the machine learning model data quality improvement detection engine may initiate a retraining of the trained cognitive computing system and/or the trained computer machine learning model via a machine learning training engine utilizing data that more accurate reflects the set of user-defined monitored and reference groupings or another set of monitored and reference groupings based on the needs of the client. The operation terminates thereafter.") the generative dialog model being configured to generate the reply corresponding to the dialog input. ("[0037]… The trained artificial intelligence based computing system or computer model may be implemented as part of or utilized by a cognitive computing system that employs the trained artificial intelligence (AI) based computing system or computer model to generate results upon which the cognitive computing system operates to generate cognitive computing responses to user requests, for example, e.g., answering natural language questions, performing image recognition, generating recommendations, decision support operations, or any other cognitive computing operation..." ) Chamarthy does not explicitly disclose that the number of updated dialog inputs is larger than the number without the update. Beaver discloses: the first dialog input set meets the following predetermined condition: a number proportion of first-class dialog inputs is larger than that of second-class dialog inputs, the first-class dialog inputs are the dialog inputs corresponding to the updated combination, and the second-class dialog inputs are the dialog inputs corresponding to the combination without the update. (“[0023] … In the case of single class bias, the system may recommend to the human user that either a calculated percentage of samples containing that value (e.g., the overrepresented sentiment) be removed or more samples from the other values be added... If a value has a significantly (for example two orders of magnitude or other appropriate threshold) smaller distribution than the other values, the system may determine that that value is underrepresented in the training data. In this case the system may recommend to the human user that either more samples for that value (e.g., the underrepresented sentiment) be added or a calculated percentage of samples for each of the other values be removed. In the case of removal, the system can randomly select a percentage of the samples and remove them automatically or allow the human to choose. For example, if 70% of output samples have the value of “negative” for the feature of sentiment, the system would recommend adding 20% more samples of “positive” or remove 20% of the existing samples that are “negative”. – adding more samples would mean that the total number of samples corresponding to the update would be greater than the previous training set.) Chamarthy and Beaver are considered analogous art to the claimed invention because they disclose methods of removing bias in machine learning systems. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Chamarthy to add additional samples for underrepresented groups as taught by Beaver. Doing so would have been beneficial to remove bias. (Beaver [0023]) This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Regarding claim 2, Chamarthy discloses: 2. The method according to claim 1, wherein the update of the safety specification comprises one or any combination of: addition of a combination and an evaluation specification of at least one corresponding evaluation dimension, addition of an evaluation dimension and a corresponding evaluation specification for a previous combination, and adjustment of the previous evaluation specification. (Fig. 4 shows that at step 422 the group is updated and the fairness metric recalculated – this is an adjustment of the previous evaluation specification.) Regarding claim 4, Chamarthy discloses: 4. The method according to claim 1, wherein optimizing the generative dialog model comprises: selecting some or all of the dialog inputs from the first dialog input set to form a second dialog input set, the second dialog input set meeting the predetermined condition; generating replies corresponding respectively to the dialog inputs in the second dialog input set by the generative dialog model to form a first reply set, and optimizing the generative dialog model and a detection model according to the first reply set and the target safety specification; selecting some or all of the dialog inputs from the first dialog input set to form a third dialog input set, the third dialog input set meeting the predetermined condition; generating replies corresponding respectively to the dialog inputs in the third dialog input set by the optimized generative dialog model to form a second reply set, and optimizing the optimized generative dialog model again according to the second reply set and the optimized detection model, the detection model being configured to carry out safety detection on the replies generated. (Claim 4 amounts to performing two iterations of the improvement process. Fig. 4 of Chamarthy shows that the process continues to iterate until the threshold is met, which reads on at least 2 iterations of the process.) Chamarthy does not explicitly disclose that the input sets meet the predetermined condition disclosed in claim 1. Beaver discloses that the second and third dialog input sets meet the predetermined condition. See claim 1 for mapping and motivation statement. Regarding claim 13, Chamarthy discloses: 13. A generative dialog implementing method, comprising: obtaining a to-be-processed dialog input; and generating a reply corresponding to the to-be-processed dialog input by using a generative dialog model, ([0062] discloses deploying the system described above in claim 1) the generative dialog model being obtained after N optimization iterations and conforming to a launch requirement, N being a positive integer greater than one, and each optimization iteration comprising: (Fig. 4 shows that the flow iterates through multiple iterations until the threshold is met.) optimization performed on the generative dialog model according to a determined dialog input and a principle that a reply generated by the generative dialog model conforms to a target safety specification in response to determination of an update of a safety specification, the target safety specification being an updated safety specification, (Fig. 4 shows that in each iteration the Fairness Metric (safety specification) is recalculated based on an updated group if the metric does not meet the threshold (launch requirement)) the update being performed on a previous safety specification when a generative dialog model after last optimization is determined not to meet a launch requirement, (Fig. 4 shows the steps continue to 420 after step 412 fails to meet the threshold (launch requirement).) wherein the safety specification comprises an evaluation specification of at least one evaluation dimension corresponding to different combinations respectively, ("[0028] Evaluating the fairness metric on the reference/monitored group as a whole provides an indication of whether erroneous inclinations exists or not based on the accepted ranges of fairness metrics indicating fairness or erroneous inclination…" – fairness and erroneous inclinations are both dimensions) any one combination consists of one content field and one application scenario, (see mapping of next two limitations) the content field is a safe content field involved by a generative dialog, and ("[0017] There are multiple techniques that exist to detect erroneous inclinations in machine learning models for both regression and classification type models. These techniques compute whether the machine learning model is exhibiting erroneous inclination against a monitored group as compared its counterpart a reference group. For example, a loan machine learning model may be erroneously inclined to give 90% of favorable outcomes to loan applicants within group B as compared to only 75% of favorable outcomes to group A applicants. If a fairness calculation is determined using a disparate impact ratio, then the metric would turn out to be 75/90=83%. If a machine learning model validator has a fairness threshold as 90% (meaning at least 90% of the monitored population should get 90% of favorable outcomes) then an inference may be made that the machine learning model is erroneous inclined against the group A." – each group reads on a content field) the application scenario is an application scenario of the generative dialog; ("[0044] The logic of the cognitive system implements the cognitive operation(s), examples of which include, but are not limited to, question answering, identification of related concepts within different portions of content in a corpus, intelligent search algorithms, such as Internet web page searches, for example, medical diagnostic and treatment recommendations, financial trend analysis, financial investment recommendations, credit scoring and credit/loan approval recommendations, and other types of recommendation generation, e.g., items of interest to a particular user, potential new contact recommendations, or the like..." – any of the applications listed reads on an application scenario) the determined dialog input corresponding to a current optimization determined according to the target safety specification, (Fig. 4 shows that in each iteration the Fairness Metric (safety specification) is recalculated based on an updated group if the metric does not meet the threshold (launch requirement). The current iteration reads on the current optimization.) the determined dialog input corresponding to the current optimization comprising: dialog inputs of a first dialog input set; (“[0083] Utilizing the identified set of user-defined monitored and reference groupings, the machine learning model data quality improvement detection engine looks at the monitored behavior data (step 406) and computes the percentage of favorable outcomes for the set of user-defined monitored and reference groupings (step 408)…) wherein the first dialog input set at least comprises dialog inputs corresponding to an updated combination, and (Fig. 4 shows step 422 updates the groups which would update the inputs corresponding to an updated combination) the first dialog input set meets the following predetermined condition: a number proportion of first-class dialog inputs is larger than that of second-class dialog inputs, the first-class dialog inputs are the dialog inputs corresponding to the updated combination, and the second-class dialog inputs are the dialog inputs corresponding to the combination without the update. (not explicitly disclosed) Chamarthy does not explicitly disclose that the number of updated dialog inputs is larger than the number without the update. Beaver discloses: the first dialog input set meets the following predetermined condition: a number proportion of first-class dialog inputs is larger than that of second-class dialog inputs, the first-class dialog inputs are the dialog inputs corresponding to the updated combination, and the second-class dialog inputs are the dialog inputs corresponding to the combination without the update. (“[0023] … In the case of single class bias, the system may recommend to the human user that either a calculated percentage of samples containing that value (e.g., the overrepresented sentiment) be removed or more samples from the other values be added... If a value has a significantly (for example two orders of magnitude or other appropriate threshold) smaller distribution than the other values, the system may determine that that value is underrepresented in the training data. In this case the system may recommend to the human user that either more samples for that value (e.g., the underrepresented sentiment) be added or a calculated percentage of samples for each of the other values be removed. In the case of removal, the system can randomly select a percentage of the samples and remove them automatically or allow the human to choose. For example, if 70% of output samples have the value of “negative” for the feature of sentiment, the system would recommend adding 20% more samples of “positive” or remove 20% of the existing samples that are “negative”. – adding more samples would mean that the total number of samples corresponding to the update would be greater than the previous training set.) Chamarthy and Beaver are considered analogous art to the claimed invention because they disclose methods of removing bias in machine learning systems. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Chamarthy to add additional samples for underrepresented groups as taught by Beaver. Doing so would have been beneficial to remove bias. (Beaver [0023]) This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Claim 14 is a device claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally, at least one processor and a memory of the Claim are taught by Chamarthy (PROCESSING UNIT(S) 306, MAIN MEMORY 308, Fig. 3). Claim 15 is a device claim with limitations corresponding to the limitations of Claim 2 and is rejected under similar rationale. Claim 17 is a device claim with limitations corresponding to the limitations of Claim 4 and is rejected under similar rationale. Claim 18 is a device claim with limitations corresponding to the limitations of Claim 13 and is rejected under similar rationale. Additionally, at least one processor and a memory of the Claim are taught by Chamarthy (PROCESSING UNIT(S) 306, MAIN MEMORY 308, Fig. 3). Claim 19 is a medium claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale. Additionally, a non-transitory computer readable storage medium storing computer instructions of the Claim are taught by Chamarthy (“[0033]… Thus, the mechanisms described herein may be implemented as specialized hardware, software executing on general purpose hardware, software instructions stored on a medium such that the instructions are readily executable by specialized or general purpose hardware, a procedure or method for executing the functions, or a combination of any of the above.”) Claim 20 is a medium claim with limitations corresponding to the limitations of Claim 13 and is rejected under similar rationale. Additionally, a non-transitory computer readable storage medium storing computer instructions of the Claim are taught by Chamarthy (see claim 19) Claim(s) 5-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chamarthy in view of Beaver as applied in claim 4 above, further in view of Brown et al (US 20200364511 A1). Regarding claim 5, Chamarthy and Beaver do not disclose the additional limitations. Brown discloses: 5. The method according to claim 4, wherein the first reply set comprises M replies generated for each dialog input in the second dialog input set, and M is a positive integer greater than one; (“[0046] As shown in block 303, the system (e.g., computer 102 and/or machine learning system 124 shown in FIG. 1) then provides alternative answers/utterances to the user (e.g., utterances U2b-a, U2b-b, U2b-c shown in FIG. 2), which are alternative answers/responses to the user's original question/utterance U1a.” – see also Fig. 2 which shows 3 replies generated for the dialog input.) wherein optimizing the generative dialog model and the detection model according to the first reply set and the target safety specification comprises: performing the following processing on any one dialog input in the second dialog input set: taking the dialog input as the to-be-processed dialog input, and acquiring each candidate reply corresponding to the to-be-processed dialog input and an manual annotation result of each candidate reply, (“[0047] As described in block 305, the user then selects one of the alternative answers/responses to the user's original question/utterance U1a. That is, the user selects one of the utterances U2b-a, U2b-b, U2b-c shown in FIG. 2 and sends it to the system (e.g., by “clicking” on the selected utterance), thus indicating that the selected answer/utterance provides the response/answer that the user was asking for in (or which proved to a successful resolution to) the user's original question/utterance U1a.” – a user selecting the utterance as correct reads on manually annotating it.) a number of the candidate replies being greater than or equal to M, (Fig. 2 shows 3 candidate replies) the candidate replies comprising the replies generated for the to-be-processed dialog input and/or replies obtained by manually modifying the replies generated for the to-be-processed dialog input, and (Fig. 2 shows that the candidate replies are generated for the dialog input) the manual annotation result of each candidate reply comprising an annotation result obtained after safety annotation is manually performed on the candidate reply according to the target safety specification; (the annotation is performed according to the target safety specification of accurate answers) constructing a training sample according to the to-be-processed dialog input, each candidate reply and the manual annotation result of each candidate reply, and optimizing the generative dialog model and the detection model by using the training sample. (“[0051] The rule for the user's original question/utterance U1a and/or a substantially similar question/answer as the user's original question/utterance U1a is thus modified to not only extract certain terms and provide a classification for the user's original question/utterance U1a and/or a substantially similar question/answer as the user's original question/utterance U1a, but also (in one or more embodiments of the present invention) provides rules for retraining machine learning…”) Chamarthy, Beaver, and Brown are considered analogous art to the claimed invention because they disclose methods of increasing accuracy in machine learning systems. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to allow the user to select the correct answer from a group of candidates as taught by Brown. This would have been beneficial so the user could select the answer that the user was asking for or which proved successful. (Brown [0047]) Regarding claim 6, Chamarthy and Beaver do not disclose the additional limitations. Brown discloses: 6. The method according to claim 5, wherein, for any one candidate reply, the annotation result after safety annotation comprises: evaluation labels of the candidate reply corresponding to different evaluation dimensions (Fig. 4 shows that the utterances are classified by intent labels (402) which correspond to different evaluation dimensions.) manually annotated according to the evaluation specifications of different evaluation dimensions of the combination corresponding to the to-be-processed dialog input, (“[0047] As described in block 305, the user then selects one of the alternative answers/responses to the user's original question/utterance U1a. That is, the user selects one of the utterances U2b-a, U2b-b, U2b-c shown in FIG. 2 and sends it to the system (e.g., by “clicking” on the selected utterance), thus indicating that the selected answer/utterance provides the response/answer that the user was asking for in (or which proved to a successful resolution to) the user's original question/utterance U1a.” – this provides a manual annotation which is used to classify the utterance according to the classes (dimensions); see also “[0077] … That is, the answer labels 502 identify responses to utterances (e.g., answers to a question, responses to a feedback, etc. —see FIG. 2) which most closely align with the intent and entities 500 for answering a particular question. Thus, in the example shown in FIG. 5, neuron 504, which is associated with the answer associated with the answer label “Product A Tech answer” has a highest output value. As such, “Product A Tech answer” is the label of the answer/response that is most appropriate for an utterance that include the intent and entities 500 found in the user's question and/or the system's response(s).”) and the evaluation label indicates conformance to the corresponding evaluation specification or non-conformance to the corresponding evaluation specification. (“[0034] A feedback loop for a conversational system starts with the user's input about the answer to their question. The system will then store this positive or negative response to be reviewed later by a subject matter expert (SME). If the feedback is negative, the SME will then provide the system enough information to retrain the system. In a conversational system, this involves correcting the initial classification of the question as understood by the system.” – providing a correct classification is a conformance to the evaluation specification.) See claim 5 for motivation statement. Regarding claim 7, Chamarthy discloses: 7. The method according to claim 6, wherein constructing the training sample according to the to-be-processed dialog input, each candidate reply and the manual annotation result of each candidate reply, and optimizing the generative dialog model and the detection model by using the training sample comprises: constructing first-class training samples and second-class training samples; optimizing the generative dialog model by the first-class training samples in a supervised learning mode; and optimizing the detection model by the second-class training samples in a supervised learning mode. (both first class and second class samples are used to optimize the model in a supervised learning mode; therefore, any 2 classes of samples reads on the claims, for example samples from the reference (class 1) and monitored groups (class 2). [0037] discloses supervised machine learning.) Chamarthy and Beaver do not disclose the additional limitations of claim 8. Brown discloses: 8. The method according to claim 7, wherein constructing the first-class training samples comprises: selecting candidate replies meeting the following condition from the candidate replies: the evaluation labels of different evaluation dimensions all indicate conformance to the corresponding evaluation specification; and forming the first-class training samples by the selected candidate replies and the to-be-processed dialog input. (“[0059] As shown in block 325, a vector distance between the question feature vector and the answer feature vector is calculated, in order to determine which answer most closely answers the user's question. That is, the question is described as a vector of intent and entities, which together describe the class of the question. Each proposed answer is also described as a vector of intent and entities, which together describe the class of each proposed answer. The difference in values between the intent and entities (class) of the question and the intent and entities (class) of the answer result in a vector distance between the classes of the questions and answers, thus describing “how closely” each response answers the user's question. A minimum vector distance is applied, such that only answers who classification vectors are close enough to the particular question (i.e., have intents and entities that closely match those of that particular question) are returned to the user.” Closely answering the user’s question is a measure of the accuracy for each category, which is the evaluation specification for each dimension.) See claim 5 for motivation statement. Regarding claim 9, Chamarthy discloses a comprehensive detection model. (ERRONEOUS INCLINATION DETECTION TOOL 112, Fig. 1) Chamarthy and Beaver do not disclose the additional limitations. Brown discloses: 9. The method according to claim 7, further comprising: obtaining a comprehensive score of each candidate reply, the higher the comprehensive score, the higher the safety; (“[0072] ... In an exemplary input, the input to input layer 403 contains values that describe that type of utterance. If DNN 424 has been properly trained (by adjusting the mathematical function (s), output value(s), weight(s), and biases in one or more of the electronic neurons within DNN 424) to output a 5-tuple output vector (e.g., 0.2, 0.9, 0.2, 0.3, 0.4) to the output layer 407, indicating that the neuron 404 that is associated with the label “tech support” has the highest value (0.9), then it indicates that the utterance 400 describes (is related) to a question or answer about tech support for a product.” – the higher value indicates a more accurate (safe) result)) wherein the detection model comprises a comprehensive detection model and classification detection models corresponding respectively to different evaluation dimensions; (“[0048] As described in block 307, the system then recommends a new classification, entity extraction, and/or rule for the user's original question/utterance U1a and/or a substantially similar question/answer as the user's original question/utterance U1a, in order to be able to appropriately respond in the future to the user's original question/utterance U1a and/or a similar version of the user's original question/utterance U1a.” – the classification reflects a different dimension depending on the question.) the second-class training sample comprises a first-sub-class training sample and a second-sub-class training sample, the first-sub-class training sample comprises two candidate replies with different comprehensive scores, the to-be-processed dialog input, and a sample label, the sample label is used to indicate the candidate reply with a higher comprehensive score in the two candidate replies, and the second-sub-class training sample comprises one candidate reply, the to-be-processed dialog input and an evaluation label for the candidate reply; (“[0059] As shown in block 325, a vector distance between the question feature vector and the answer feature vector is calculated, in order to determine which answer most closely answers the user's question. That is, the question is described as a vector of intent and entities, which together describe the class of the question. Each proposed answer is also described as a vector of intent and entities, which together describe the class of each proposed answer. The difference in values between the intent and entities (class) of the question and the intent and entities (class) of the answer result in a vector distance between the classes of the questions and answers, thus describing “how closely” each response answers the user's question. A minimum vector distance is applied, such that only answers who classification vectors are close enough to the particular question (i.e., have intents and entities that closely match those of that particular question) are returned to the user.” – the vector distance reads on a comprehensive score. Because only answers meeting a threshold are returned, there would be cases with 1 candidate reply and cases with 2 candidate replies depending on the vector distance.) optimizing the detection models comprises: optimizing the comprehensive detection model by using the first-sub-class training sample, and for any one classification detection model, optimizing the classification detection model by using the second-sub-class training sample comprising the evaluation label of the evaluation dimension corresponding to the classification detection model. (“[0051] The rule for the user's original question/utterance U1a and/or a substantially similar question/answer as the user's original question/utterance U1a is thus modified to not only extract certain terms and provide a classification for the user's original question/utterance U1a and/or a substantially similar question/answer as the user's original question/utterance U1a, but also (in one or more embodiments of the present invention) provides rules for retraining machine learning…”) See claim 5 for motivation statement. Regarding claim 10, Chamarthy discloses: 10. The method according to claim 4, wherein the second reply set comprises the replies generated respectively for the dialog inputs in the third dialog input set ; wherein optimizing the optimized generative dialog model again according to the second reply set and the optimized detection model comprises: performing safety detection on each reply in the second reply set by the optimized detection model, and optimizing the optimized generative dialog model again in a reinforcement learning manner according to a safety detection result of each reply. (Fig. 4 shows that the fairness metric is calculated in each iteration (step 418). See “[0089] The machine learning model data quality improvement detection engine then determines a fairness metric using a disparate impact ratio based on the original record data and the newly considered perturbed record data (step 524)…”; Brown discloses optimization as mapped in claim 5 above.) Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chamarthy in view of Beaver and Brown as applied in claim 10 above, further in view of Chaloulos et al. (US 20200320428 A1). Regarding claim 11, Chamarthy in view of Beaver and Brown discloses: 11. The method according to claim 10, wherein the detection models comprise a comprehensive detection model and classification detection models corresponding respectively to different evaluation dimensions; (see claim 9 mapping) wherein optimizing the optimized generative dialog model again in the reinforcement learning manner according to the safety detection result of each reply comprises: performing the following processing for any one reply: obtaining a comprehensive detection result of the reply and classification detection results corresponding to different classification detection models respectively, (see claim 9 mapping) determining a reward corresponding to the reply by combining the comprehensive detection result and the different classification detection results, and forming a training sample by using the reply, the dialog input corresponding to the reply and the reward; and optimizing the optimized generative dialog model again by using the training samples. (not explicitly disclosed) Chamarthy, Beaver, and Brown do not explicitly disclose determining a reward and optimizing the model based on the reward. Chaloulos discloses: determining a reward corresponding to the reply by combining the comprehensive detection result and the different classification detection results, and forming a training sample by using the reply, the dialog input corresponding to the reply and the reward; and optimizing the optimized generative dialog model again by using the training samples. (“[0045] Thereby, the reward function may be a metric specifying how good that reinforcement engine reaches its target which may be defined by a performance metric, a preciseness metric, and a fairness metric.”; see also “[0061] The specification on this function again depends on the problem. An example of a reward function is the ratio of the performance metric (F-score) to the fairness metric (average relative difference between data samples that only differ in protected attributes), 504. The output of the reward function is then passed to the controlling reinforcement learning engine, which may either begin the next duration or terminate the optimization process, 514.”) Chamarthy, Beaver, Brown, and Chaloulos are considered analogous art to the claimed invention because they disclose methods of increasing accuracy in machine learning systems. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to optimize learning with a reward as taught by Chaloulos. This would have been beneficial to efficiently optimize the parameters and/or hyperparameters of the model. (Chaloulos [0071]) This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chamarthy in view of Beaver, Brown, and Chaloulos as applied in claim 11 above, further in view of Agarwal et al. (US 20190080225 A1 ). Regarding claim 12, Chamarthy, Beaver, Brown, and Chaloulos do not disclose Kullback-Leibler divergence. Agarwal discloses: 12. The method according to claim 11, wherein optimizing the optimized generative dialog model again by using the training samples comprises: using the optimized generative dialog model as a baseline model, and generating a target model identical to the baseline model; and optimizing the target model using the training sample based on a constraint of a Kullback-Leibler divergence introduced between the baseline model and the target model, and using the optimized target model as the generative dialog model optimized again. (“[0038] In an embodiment of the present disclosure, at step 212, a softmax layer of the classification model 304 determines at least one target class of the one or more queries based on the final vector formed and outputs (or provides) a response to the one or more queries based on the determined target class. In an embodiment, the system 100 provides response from one or more pre-defined responses stored in the database 108. In an embodiment, a Square root Kullback-Leibler divergence (KLD) Loss Function is applied to the sequence of vector to optimize the classification model 304. In an embodiment, the crossentropy loss function can be seen as KLdivergence between predicted discrete probability distribution P … and the target distribution T… “) Chamarthy, Beaver, Brown, Chaloulos, and Agarwal are considered analogous art to the claimed invention because they disclose methods of increasing accuracy in machine learning systems. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to optimize learning with Kullback-Leibler divergence as taught by Agarwal. This would have been beneficial to optimize the model. (Agarwal [0038]) This combination falls under combining prior art elements according to known methods to yield predictable results or use of known technique to improve similar devices (methods, or products) in the same way. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Ktena et al. (WO 2024180126 A1). Ktena discloses a method for improving fairness of machine learning models by augmenting the training data to meet a target distribution ([0087]-[0095]). Dey et al. (US 20210035014 A1). Dey discloses a method for training AI models by modifying data points and discarding them if they are out of the desired class (see figs. 3-4). Any inquiry concerning this communication or earlier communications from the examiner should be directed to JON C MEIS whose telephone number is (703)756-1566. The examiner can normally be reached Monday - Thursday, 8:30 am - 5:30 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JON CHRISTOPHER MEIS/Examiner, Art Unit 2654 /HAI PHAN/Supervisory Patent Examiner, Art Unit 2654
Read full office action

Prosecution Timeline

Jun 17, 2024
Application Filed
May 13, 2026
Non-Final Rejection mailed — §103
Jul 20, 2026
Response Filed
Sep 22, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711961
SYSTEM AND METHOD FOR DIGITAL VOICE DATA PROCESSING AND AUTHENTICATION
3y 4m to grant Granted Aug 18, 2026
Patent 12603087
VOICE RECOGNITION USING ACCELEROMETERS FOR SENSING BONE CONDUCTION
3y 8m to grant Granted Apr 14, 2026
Patent 12579975
Detecting Unintended Memorization in Language-Model-Fused ASR Systems
2y 11m to grant Granted Mar 17, 2026
Patent 12482487
MULTI-SCALE SPEAKER DIARIZATION FOR CONVERSATIONAL AI SYSTEMS AND APPLICATIONS
3y 0m to grant Granted Nov 25, 2025
Patent 12475312
FOREIGN LANGUAGE PHRASES LEARNING SYSTEM BASED ON BASIC SENTENCE PATTERN UNIT DECOMPOSITION
2y 9m to grant Granted Nov 18, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
33%
Grant Probability
86%
With Interview (+52.4%)
2y 10m (~7m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 33 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month