Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 1-18 are objected to because of the following informalities:
Claim 1 lines 9-10: “the answers” lacks proper antecedent basis.
Claim 6, line 1: before “one or more”, --the-- should be inserted.
Claim 7, line 2: “scores and answers” should be --the scores and the answers--.
Claim 9 lines 2-3: “the test answer” and “the answer” lack proper antecedent basis.
Claim 10 line 14: “the answers” lacks proper antecedent basis.
Claim 15, line 1: before “one or more”, --the-- should be inserted.
Claim 16, line 2: “scores and answers” should be --the scores and the answers--.
Claim 18 line3: “the test answer” and “the answer” lack proper antecedent basis.
Claims 2-5, 8, 11-14, and 17 are rejected for being dependent on claims 1 and 10.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 5, 7, 9, 10, 12, 14, 16, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Christian et al. (US20190065465, Christian hereinafter), in view of Lakkundi et al. (US11429382, Lakkundi hereinafter), Nizar et al. (US20220245362, Nizar hereinafter), and Ohashi et al. (“Post-processing Networks: Method for Optimizing Pipeline Task-oriented Dialogue Systems using Reinforcement Learning”, 07/25/2022, Ohashi hereinafter).
Regarding claim 1, Christian discloses: A method for evaluating a quality of chatbot answers after changes have been made to (see Christian, paragraph [0008], “…different versions of a chatbot may be tested based on the same conversation with a user…”), (see Christian, paragraph [0007], “…A conversation runner may create a test case from the tag-based conversation file and execute the test case for a different version of the chatbot or a different version of an application integrated with the chatbot…”), comprising:
(see Christian, paragraph [0012], “…The first version may be a version of the chatbot that was previously tested and verified to be operating correctly…”), (see Christian, paragraph [0016], “…The conversation file may be generated from an actual conversation between the user and the chatbot 204… the user may employ an automated test of a second version 206 of the same chatbot 204 integrated with the source application 202 using the conversation file 210… A conversation runner 212 may execute a script from the conversation file 210 and may produce a conversation file 216 representing the same user conversation with the second version of the chatbot 206 using the same user quires from the conversation file 210… comparator 214 may compare the conversation files 210 and 216...”);
obtaining test answers, generated by a new version of the chatbot, to the test questions (see Christian, paragraph [0018], “…The chatbot 206 may represent a new version of the chatbot 204 to be tested. The conversation runner 212 may tag and replay the conversation from the conversation file 210. When the conversation file 210 is replayed, each user tag may trigger a post into the messaging application 208. Each action tag may trigger a method (action call) to the chatbot 206… the conversation runner 212 may generate the conversation file 216 based on execution of the scenario from the conversation file 210 with the current chatbot version…”), (see Christian, paragraph [0016]);
(see Christian, paragraph [0016]), (see Christian, paragraph [0015], “…if the second version of the chatbot is responding differently from the first version to the same request, then the different response is detected…”), (see Christian, paragraph [0012]; and
performing an automated testing process that comprises comparing (see Christian, paragraph [0016]), (see Christian, paragraph [0012]).
Christian does not appear to distinctly disclose:
filtering, based on information identifying changes that have been made to
scoring
However, Lakkundi discloses:
filtering, based on information identifying changes that have been made to (see Lakkundi, col 6 lines 40-42, “…the process at 202 automatically identifies the program code file(s) that were changed when changes are made to program code of an application…”), (see Lakkundi, col 3 lines 63-67, “…Based on observed code changes, it can be determined which program code file(s) were touched as part of those changes, and therefore which features may have been impacted and are therefore to undergo regression testing…”), (see Lakkundi, col 7 lines 48-51, “…the process selects, from a collection of the regression test cases for the application's features, regression test cases to be included in the automated regression testing…”);
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include filtering, based on information identifying changes as taught by Lakkundi, for the result of ensuring that only changed information is tested.
Christian as modified does not appear to distinctly disclose:
scoring
However, Nizar discloses:
scoring (see Nizar, paragraph [0024], “…A natural language model is applied to each expanded text generated from a base text, to determine a respective perplexity score…”), (see Nizar, paragraph [0111], “…the target set augmentation system ranks the expanded texts based on the perplexity scores. The smaller perplexity scores are ranked first; the larger perplexity scores are ranked last.”), (see Nizar, paragraph [0108], “…applying a natural language model to the current expanded text to determine a perplexity score (Operation 606)… The natural language model outputs a perplexity score for the current expanded text…”); and
(see Nizar, paragraph [0111]).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include scoring based on respective perplexity scores as taught by Nizar, for the result of measuring answer quality.
Christian as modified does not appear to distinctly disclose:
However, Ohashi discloses:
A method for (see Ohashi, Abstract, “...neural-based components called post-processing networks (PPNs) are installed inside such a system to post-process the output of each module...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include internal post-processing tasks as taught by Ohashi, for the result of evaluating differences that the modification has on answer quality.
Regarding claim 3, Christian discloses:
wherein a set comprising the test questions and the test answers is built without human involvement (see Christian, paragraph [0018], “The conversation runner 212 may tag and replay the conversation from the conversation file 210. When the conversation file 210 is replayed... Each action tag may trigger a method (action call) to the chatbot 206”), (see Christian, paragraph [0034], “The conversation runner 212 may produce automatically-generated conversation in the same format as a test conversation, in order to generate the test results.”).
Regarding claim 5, Christian discloses:
wherein (see Christian, paragraph [0012], “…The first version may be a version of the chatbot that was previously tested and verified to be operating correctly…”).
Christian does not appear to distinctly disclose:
However, Ohashi discloses:
(see Ohashi, Abstract, “...neural-based components called post-processing networks (PPNs) are installed inside such a system to post-process the output of each module...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include internal post-processing tasks as taught by Ohashi, for the result of evaluating differences that the modification has on answer quality.
Regarding claim 7, Christian discloses:
comparing respective aggregate (see Christian, paragraph [0018], “...the conversation runner 212 may load the conversation file 210 into a conversation files repository...”), (see Christian, paragraph [0016], “…The conversation file may be generated from an actual conversation between the user and the chatbot 204… the user may employ an automated test of a second version 206 of the same chatbot 204 integrated with the source application 202 using the conversation file 210… A conversation runner 212 may execute a script from the conversation file 210 and may produce a conversation file 216 representing the same user conversation with the second version of the chatbot 206 using the same user quires from the conversation file 210… comparator 214 may compare the conversation files 210 and 216...”), and
the reference table comprises the test questions and the answers generated by the reference version of the chatbot (see Christian, paragraph [0018]), (see Christian, paragraph [0016]), (see Christian, paragraph [0012], “…The first version may be a version of the chatbot that was previously tested and verified to be operating correctly…”), and
the test table comprises the test questions and the answers generated by the new version of the chatbot (see Christian, paragraph [0018]), (see Christian, paragraph [0016]).
Christian does not appear to distinctly disclose:
However, Nizar discloses:
comparing respective aggregate scores (see Nizar, paragraph [0111], “…the target set augmentation system ranks the expanded texts based on the perplexity scores. The smaller perplexity scores are ranked first; the larger perplexity scores are ranked last.”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include the comparison of respective aggregate scores as taught by Nizar, for the result of producing a single summary measure for each version so that the overall change in answer quality between versions can be assessed.
Regarding claim 9, Chistian discloses:
wherein, after the comparing, a report is generated that explains, for a given one of the test questions, a change in the test answer relative to the answer generated by the reference version of the chatbot, and also explains any changes in the (see Christian, paragraph [0035], “...a hypertext markup language (HTML) test report 220 may be generated for all test results.”), (see Christian, paragraph [0016], “…The conversation file may be generated from an actual conversation between the user and the chatbot 204… the user may employ an automated test of a second version 206 of the same chatbot 204 integrated with the source application 202 using the conversation file 210… A conversation runner 212 may execute a script from the conversation file 210 and may produce a conversation file 216 representing the same user conversation with the second version of the chatbot 206 using the same user quires from the conversation file 210… comparator 214 may compare the conversation files 210 and 216...”), (see Christian, paragraph [0019], “...The comparator 214 may generate the test report 220... if a threshold is set at 90%, the new version 206 may be considered acceptable and no further actions by the developers may be needed. In other words, a sufficient number of correct responses are produces by the new version of the chatbot 206...”).
Christian does not appear to distinctly disclose:
However, Nizar discloses: (see Nizar, paragraph [0108],“…applying a natural language model to the current expanded text to determine a perplexity score (Operation 606)… The natural language model outputs a perplexity score for the current expanded text…”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include scores as taught by Nizar, for the result of providing a quantitative measure of answer quality that can be tracked and reported.
Regarding claim 10, Christian discloses: A non-transitory storage medium having stored therein instructions that are executable by one or more hardware processors (e.g., see Christian paragraph [0011]) to:
perform operations that implement a method for evaluating a quality of chatbot answers after change have made to (see Christian, paragraph [0008], “…different versions of a chatbot may be tested based on the same conversation with a user…”), (see Christian, paragraph [0007], “…A conversation runner may create a test case from the tag-based conversation file and execute the test case for a different version of the chatbot or a different version of an application integrated with the chatbot…”), the operations comprising:
(see Christian, paragraph [0012], “…The first version may be a version of the chatbot that was previously tested and verified to be operating correctly…”), (see Christian, paragraph [0016], “…The conversation file may be generated from an actual conversation between the user and the chatbot 204… the user may employ an automated test of a second version 206 of the same chatbot 204 integrated with the source application 202 using the conversation file 210… A conversation runner 212 may execute a script from the conversation file 210 and may produce a conversation file 216 representing the same user conversation with the second version of the chatbot 206 using the same user quires from the conversation file 210… comparator 214 may compare the conversation files 210 and 216...”);
obtaining test answers, generated by a new version of the chatbot, to the test questions (see Christian, paragraph [0018], “…The chatbot 206 may represent a new version of the chatbot 204 to be tested. The conversation runner 212 may tag and replay the conversation from the conversation file 210. When the conversation file 210 is replayed, each user tag may trigger a post into the messaging application 208. Each action tag may trigger a method (action call) to the chatbot 206… the conversation runner 212 may generate the conversation file 216 based on execution of the scenario from the conversation file 210 with the current chatbot version…”), (see Christian, paragraph [0016]);
(see Christian, paragraph [0016]), (see Christian, paragraph [0015], “…if the second version of the chatbot is responding differently from the first version to the same request, then the different response is detected…”), (see Christian, paragraph [0012]; and
performing an automated testing process that comprises comparing (see Christian, paragraph [0016]), (see Christian, paragraph [0012]).
Christian does not appear to distinctly disclose:
...
filtering, based on information identifying changes that have been made to
scoring
However, Lakkundi discloses:
filtering, based on information identifying changes that have been made to (see Lakkundi, col 6 lines 40-42, “…the process at 202 automatically identifies the program code file(s) that were changed when changes are made to program code of an application…”), (see Lakkundi, col 3 lines 63-67, “…Based on observed code changes, it can be determined which program code file(s) were touched as part of those changes, and therefore which features may have been impacted and are therefore to undergo regression testing…”), (see Lakkundi, col 7 lines 48-51, “…the process selects, from a collection of the regression test cases for the application's features, regression test cases to be included in the automated regression testing…”);
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include filtering, based on information identifying changes as taught by Lakkundi, for the result of ensuring that only changed information is tested.
Christian as modified does not appear to distinctly disclose:
scoring
However, Nizar discloses:
scoring (see Nizar, paragraph [0024], “…A natural language model is applied to each expanded text generated from a base text, to determine a respective perplexity score…”), (see Nizar, paragraph [0111], “…the target set augmentation system ranks the expanded texts based on the perplexity scores. The smaller perplexity scores are ranked first; the larger perplexity scores are ranked last.”), (see Nizar, paragraph [0108],“…applying a natural language model to the current expanded text to determine a perplexity score (Operation 606)… The natural language model outputs a perplexity score for the current expanded text…”); and
(see Nizar, paragraph [0111]).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include scoring based on respective perplexity scores as taught by Nizar, for the result of measuring answer quality.
Christian as modified does not appear to distinctly disclose:
However, Ohashi discloses:
A method for (see Ohashi, Abstract, “...neural-based components called post-processing networks (PPNs) are installed inside such a system to post-process the output of each module...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include internal post-processing tasks as taught by Ohashi, for the result of evaluating differences that the modification has on answer quality.
Regarding claim 12, Christian discloses:
wherein a set comprising the test questions and the test answers is built without human involvement (see Christian, paragraph [0018], “The conversation runner 212 may tag and replay the conversation from the conversation file 210. When the conversation file 210 is replayed... Each action tag may trigger a method (action call) to the chatbot 206”), (see Christian, paragraph [0034], “The conversation runner 212 may produce automatically-generated conversation in the same format as a test conversation, in order to generate the test results.”).
Regarding claim 14, Christian discloses:
wherein (see Christian, paragraph [0012], “…The first version may be a version of the chatbot that was previously tested and verified to be operating correctly…”).
Christian does not appear to distinctly disclose:
However, Ohashi discloses:
(see Ohashi, Abstract, “...neural-based components called post-processing networks (PPNs) are installed inside such a system to post-process the output of each module...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include internal post-processing tasks as taught by Ohashi, for the result of evaluating differences that the modification has on answer quality.
Regarding claim 16, Christian discloses:
comparing respective aggregate (see Christian, paragraph [0018], “...the conversation runner 212 may load the conversation file 210 into a conversation files repository...”), (see Christian, paragraph [0016], “…The conversation file may be generated from an actual conversation between the user and the chatbot 204… the user may employ an automated test of a second version 206 of the same chatbot 204 integrated with the source application 202 using the conversation file 210… A conversation runner 212 may execute a script from the conversation file 210 and may produce a conversation file 216 representing the same user conversation with the second version of the chatbot 206 using the same user quires from the conversation file 210… comparator 214 may compare the conversation files 210 and 216...”), and
the reference table comprises the test questions and the answers generated by the reference version of the chatbot (see Christian, paragraph [0018]), (see Christian, paragraph [0016]), (see Christian, paragraph [0012], “…The first version may be a version of the chatbot that was previously tested and verified to be operating correctly…”), and
the test table comprises the test questions and the answers generated by the new version of the chatbot (see Christian, paragraph [0018]), (see Christian, paragraph [0016]).
Christian does not appear to distinctly disclose:
However, Nizar discloses:
comparing respective aggregate scores (see Nizar, paragraph [0111], “…the target set augmentation system ranks the expanded texts based on the perplexity scores. The smaller perplexity scores are ranked first; the larger perplexity scores are ranked last.”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include the comparison of respective aggregate scores as taught by Nizar, for the result of producing a single summary measure for each version so that the overall change in answer quality between versions can be assessed.
Regarding claim 18, Chistian discloses:
wherein, after the comparing, a report is generated that explains, for a given one of the test questions, a change in the test answer relative to the answer generated by the reference version of the chatbot, and also explains any changes in the (see Christian, paragraph [0035], “...a hypertext markup language (HTML) test report 220 may be generated for all test results.”), (see Christian, paragraph [0016], “…The conversation file may be generated from an actual conversation between the user and the chatbot 204… the user may employ an automated test of a second version 206 of the same chatbot 204 integrated with the source application 202 using the conversation file 210… A conversation runner 212 may execute a script from the conversation file 210 and may produce a conversation file 216 representing the same user conversation with the second version of the chatbot 206 using the same user quires from the conversation file 210… comparator 214 may compare the conversation files 210 and 216...”), (see Christian, paragraph [0019], “...The comparator 214 may generate the test report 220... if a threshold is set at 90%, the new version 206 may be considered acceptable and no further actions by the developers may be needed. In other words, a sufficient number of correct responses are produces by the new version of the chatbot 206...”).
Christian does not appear to distinctly disclose:
However, Nizar discloses: (see Nizar, paragraph [0108],“…applying a natural language model to the current expanded text to determine a perplexity score (Operation 606)… The natural language model outputs a perplexity score for the current expanded text…”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include scores as taught by Nizar, for the result of providing a quantitative measure of answer quality that can be tracked and reported.
Claims 2, 4, 6, 8, 11, 13, 15, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Christian, Lakkundi, Nizar, and Ohashi as applied to claims 1 and 10 above, and further in view of Hu et al. (US20250097171, Hu hereinafter).
Regarding claim 2, Christian as modified does not appear to distinctly disclose:
wherein the chatbot comprises an LLM (large language model)-based chatbot.
However, Hu discloses:
wherein the chatbot comprises an LLM (large language model)-based chatbot (see Hu, paragraph [0011], “...an LLM may be configured to perform as a ChatBot...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include an LLM as taught by Hu, for the result of generating higher quality responses as compared to traditional chatbots.
Regarding claim 4, Christian as modified does not appear to distinctly disclose:
wherein the comparing indicates whether or not a change has occurred between a quality of the answers generated by the reference version of the chatbot, and a quality of the test answers generated by the new version of the chatbot.
However, Hu discloses:
wherein the comparing indicates whether or not a change has occurred between a quality of the answers generated by the reference version of the chatbot, and a quality of the test answers generated by the new version of the chatbot (see Hu, paragraph [0009], “...the pipeline automatically tests the tuned LLM as a ChatBot to determine the improvement/degradation of the fine-tuned LLM to determine whether or not the LLM weights for ChatBot are improved over prior weights...”), (see Hu, paragraph [0035], “...automatic LLM evaluator 110 is configured to generate an evaluation score 150 for performance of the tuned large language model 134 as a chatbot...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include indicating a change in answer quality between chatbot versions as taught by Hu, for the result of assessing the improvement or lack or improvement between chatbot version.
Regarding claim 6, Christian as modified does not appear to distinctly disclose:
wherein one or more of the test answers are scored using a similarity function.
However, Hu discloses:
wherein one or more of the test answers are scored using a similarity function (see Hu, paragraph [0034], “...a BLEU (Bilingual Evaluation Understudy) score loss that is configured to quantify similarity between generated response(s) 138... based on n-gram overlap...”), (see Hu, paragraph [0031], “...determine a cosine distance between the respective embedded vectors for the collective generated responses and collective example responses...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include scoring using a similarity function as taught by Hu, for the result of determining what degree an answer diverged between chatbot versions.
Regarding claim 8, Christian as modified does not appear to distinctly disclose:
wherein, based on the comparing, either the new version of the chatbot is deployed to a production environment in place of the reference version of the chatbot, or the reference version of the chatbot remains in the production environment and is not replaced with the new version of the chatbot.
However, Hu discloses:
wherein, based on the comparing, either the new version of the chatbot is deployed to a production environment in place of the reference version of the chatbot (see Hu, paragraph [0040], “...deployment decider 112 is configured to automatically determine to deploy 158 the tuned large language model 134 to a production environment 156 for ChatBot in response to the evaluation score satisfying the threshold 154....”), (see Hu, paragraph [0045], “The automated deployment process rolls the tuned LLM 134 out to production environment 156 to replace or supersede a prior LLM as a ChatBot...”), or the reference version of the chatbot remains in the production environment and is not replaced with the new version of the chatbot (see Hu, paragraph [0109], “...If the evaluation shows insufficient improvement in ChatBot performance by the tuned LLM 134, or even decrease in performance, the tuned LLM 134 is not deployed. Instead, the deployment decider 112 initiates further epochs of training...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include deploying or not deploying a new version of the chatbot as taught by Hu, for the result of producing only chatbot versions whose answer quality is verified to be acceptable.
Regarding claim 11, Christian as modified does not appear to distinctly disclose:
wherein the chatbot comprises an LLM (large language model)-based chatbot.
However, Hu discloses:
wherein the chatbot comprises an LLM (large language model)-based chatbot (see Hu, paragraph [0011], “...an LLM may be configured to perform as a ChatBot...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include an LLM as taught by Hu, for the result of generating higher quality responses as compared to traditional chatbots.
Regarding claim 13, Christian as modified does not appear to distinctly disclose:
wherein the comparing indicates whether or not a change has occurred between a quality of the answers generated by the reference version of the chatbot, and a quality of the test answers generated by the new version of the chatbot.
However, Hu discloses:
wherein the comparing indicates whether or not a change has occurred between a quality of the answers generated by the reference version of the chatbot, and a quality of the test answers generated by the new version of the chatbot (see Hu, paragraph [0009], “...the pipeline automatically tests the tuned LLM as a ChatBot to determine the improvement/degradation of the fine-tuned LLM to determine whether or not the LLM weights for ChatBot are improved over prior weights...”), (see Hu, paragraph [0035], “...automatic LLM evaluator 110 is configured to generate an evaluation score 150 for performance of the tuned large language model 134 as a chatbot...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include indicating a change in answer quality between chatbot versions as taught by Hu, for the result of assessing the improvement or lack or improvement between chatbot version.
Regarding claim 15, Christian as modified does not appear to distinctly disclose:
wherein one or more of the test answers are scored using a similarity function.
However, Hu discloses:
wherein one or more of the test answers are scored using a similarity function (see Hu, paragraph [0034], “...a BLEU (Bilingual Evaluation Understudy) score loss that is configured to quantify similarity between generated response(s) 138... based on n-gram overlap...”), (see Hu, paragraph [0031], “...determine a cosine distance between the respective embedded vectors for the collective generated responses and collective example responses...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include scoring using a similarity function as taught by Hu, for the result of determining what degree an answer diverged between chatbot versions.
Regarding claim 17, Christian as modified does not appear to distinctly disclose:
wherein, based on the comparing, either the new version of the chatbot is deployed to a production environment in place of the reference version of the chatbot, or the reference version of the chatbot remains in the production environment and is not replaced with the new version of the chatbot.
However, Hu discloses:
wherein, based on the comparing, either the new version of the chatbot is deployed to a production environment in place of the reference version of the chatbot (see Hu, paragraph [0040], “...deployment decider 112 is configured to automatically determine to deploy 158 the tuned large language model 134 to a production environment 156 for ChatBot in response to the evaluation score satisfying the threshold 154....”), (see Hu, paragraph [0045], “The automated deployment process rolls the tuned LLM 134 out to production environment 156 to replace or supersede a prior LLM as a ChatBot...”), or the reference version of the chatbot remains in the production environment and is not replaced with the new version of the chatbot (see Hu, paragraph [0109], “...If the evaluation shows insufficient improvement in ChatBot performance by the tuned LLM 134, or even decrease in performance, the tuned LLM 134 is not deployed. Instead, the deployment decider 112 initiates further epochs of training...”).
It would have been obvious to one of ordinary skill in the art before the effecting filing
date of the claimed invention to have modified a system for comparing chatbot versions as taught by Christian, to include deploying or not deploying a new version of the chatbot as taught by Hu, for the result of producing only chatbot versions whose answer quality is verified to be acceptable.
Conclusion
Any inquiry concerning this communication or earlier communications from the
examiner should be directed to Joshua Tran whose telephone number is (571)272-5460.
The examiner can normally be reached on M-F 9-5.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s
supervisor, Hyung Sough can be reached on (571)272-6799. The fax phone number for the
organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the
automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JOSHUA TRAN/
Examiner, Art Unit 2192
/S. SOUGH/
spe, art unit 2192