DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Claims 1, 13, and 19 are amended. Claims 2, 6, 14, and 18 are canceled. As such, Claims 1, 3-5, 7-13, 15-17, and 19-20 are presented for examination.
Response to Arguments
Rejection under 35 U.S.C. 101
Applicant's arguments have been fully considered but they are not persuasive. Applicant argues, “claim 1 has been amended to recite the use of embeddings to represent analytical parameters and segments. Such functionality cannot be performed in the human mind. Embeddings are vector representations of features extracted from a word or group of words that represents semantic content of the word or group of words. The process of generating such embeddings is inherently technical and involves the use of a machine learning model trained to generate embeddings from such inputs; it is not process that can be performed in the human mind.”
However, under the broadest reasonable interpretation, an embedding is a vector of numerical values representing an object in relation to other objects. The human mind is capable of assigning numerical values to an object to form a numerical vector, and is thus capable of determining embeddings for analytical parameters and segments of a communication record. Further, the use of trained ML models to generate the embeddings is recited at a high level of generality and merely confines the generation of embeddings to ML models. Additionally, MPEP 2106.05(a) states, “the judicial exception alone cannot provide the improvement. The improvement can be provided by one or more additional elements… In addition, the improvement can be provided by the additional element(s) in combination with the recited judicial exception.” In this case, the “trained ML models” and “trained LLM” represent the additional elements and do not provide an improvement to a computer technology because they generally link the use of a judicial exception (mental process) to the field of machine learning and are used to apply the mental process on a computer.
Rejection under 35 U.S.C. 102/103
Applicant’s arguments with respect to the limitation, “determining a subset of segments semantically associated with the respective analytical parameter based on determined similarities between the analytical parameter embeddings and each of the segment embeddings,” of claim 1 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Applicant's arguments with respect to the limitation, “generating… an evaluation of each segment of the respective subset of segments with respect to the respective analytical parameter,” of claim 1 have been fully considered but they are not persuasive. Applicant argues, “the cited functionality in Anwade only relates to generating evaluation prompts, not generating evaluations of each segment of the respective subset of segments. Nothing in Anwade discloses or makes obvious generating an evaluation of each segment that was identified as being semantically associated with a respective analytical parameter.”
However, Anwade discloses an evaluation for each segment that corresponds to an analytical parameter using the evaluation prompts. For example, paragraph [0118] of Anwade recites, “in case that three evaluation prompts are created, three interaction data items within an interaction may be evaluated and three evaluation scores may be obtained that can be compared, e.g. to pre-set threshold values for the performance indicators.” These “interaction data items” represent a subset of segments that are evaluated to determine evaluation scores, which are then used to determine an evaluation result in response to an evaluation question. Thus, Anwade discloses generating an evaluation of interaction data items, representing each segment of a subset of segments, with respect to an evaluation question, representing an analytical parameter.
Further, applicant argues, “claim 1 recites that a trained LLM generates a response to the respective analytical parameter based on the evaluations of each of the segments discussed above. The cited functionality in Anwade does not perform such functionality... And while the Office Action also cites to paragraph 117, this functionality also does not disclose or make obvious that an LLM generates a response to the respective analytical parameter based on the evaluations of each of the segments discussed above.”
However, Anwade discloses using an LLM to generate an evaluation result as a response to a question using an evaluation of segments. Specifically, paragraph [0179] of Anwade recites, “A QA chain may be a series of steps that may allow evaluating interaction transcripts based on questions, e.g. present in evaluation forms of agents, e.g. evaluation form 804. In some embodiments, OpenAI and LangChain may be used to perform QA on an interaction transcript; other models may be used. In the evaluation, an auto-evaluation and coaching opportunity finder engine 730 may be used to conduct the generation of chunks that include interaction data items from an interaction recording, vectorization, similarity search, and OpenAI's large language models to generate an evaluation result for interaction data items present that are present in an interaction.” The evaluation result, which represents a response to an analytical parameter, is based on performance indicators that are determined from the evaluation of interaction data items. Paragraphs [0118-0119] of Anwade recite, “in case that three evaluation prompts are created, three interaction data items within an interaction may be evaluated and three evaluation scores may be obtained that can be compared, e.g. to pre-set threshold values for the performance indicators… The identification of evaluation results may include the comparison of one or more performance indicators, identified in the interaction data items or determined by calculating the aggregation of each focus area and behavior combination 613, with pre-set threshold values 614 for the performance indicators.” Thus, the generation of an evaluation result relies on the evaluation of interaction data items representing segments and the use of an LLM.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 3-5, 7-13, 15-17, and 19-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding claim 1, the claim recites “(a) receiving a set of communication records, the set of communication records representing one or more communications between a first person and a second person,” “(b) receiving a set of analytical parameters associated with the set of communication records,” “(c) generating a plurality of segments from the set of communication records,” “(d) generating… analytical parameter embeddings for each analytical parameter of the set of analytical parameters,“ “(e) generating… segment embeddings for each segment of the plurality of segments,” (f) for each analytical parameter in the set of analytical parameters: determining a subset of segments semantically associated with the respective analytical parameter,” “(g) generating… an evaluation of each segment of the respective subset of segments with respect to the respective analytical parameter,” “(h) generating… a response to the analytical parameter based on the evaluations of the segments,” and “(i) outputting a full evaluation of the set of communication records based on the set of analytical parameters and the respective generated responses to the analytical parameters”. Limitations (a) – (i) recite mental processes that may be practically performed in the mind using pen and paper. For example, limitation (a) can be done by a person receiving chat logs between two people. Limitation (b) can be done by a person receiving a set of parameters related to the chat logs. Limitation (c) can be done by a person determining different sections of a set of chat logs. Limitation (d) can be done by someone determining numerical vectors to represent analytical parameters. Limitation (e) can be done by someone determining numerical vectors to represent segments of a communication record. Limitation (f) can be done by a person determining a subset of input text is semantically related to a parameter. Limitation (g) can be done by a person evaluating each section of an input text corresponding a specific parameter. Limitation (h) can be done by a person determining a response to a parameter based on evaluation different text sections. Limitation (i) can be done by a person evaluating a chat log based specific parameters and responses to those parameters, and determining an output evaluation. Under its broadest reasonable interpretation when read in light of the specification, the actions of “receiving,” “generating,” “determining,” and “outputting” encompass mental processes practically performed in the human mind by observation, evaluation, and judgement using pen and paper. Accordingly, the claim recites an abstract idea (Step 2A, Prong One).
The judicial exception is not integrated into a practical application. In particular, the claim recites additional element of “(j) using a trained machine learning ("ML") model” and “(k) using a trained large language model (‘LLM’).” Further, limitations (a) - (i) are recited as being performed by a computer. In limitations (a) - (b), the computer is used as a tool to perform the generic computer function of receiving data. In limitations (c) - (i), the computer is used to perform an abstract idea, as discussed above in Step 2A, Prong One, such that it amounts to no more than mere instructions to apply the exception using a generic computer. The limitations (j) – (k) provide nothing more than mere instructions to implement an abstract idea on a generic computer. The ML model and LLM recited in limitations (j) and (k) is used to perform limitations (d) – (e) and (g) – (h), respectively, without placing any limits on how the models functions. Rather, these models only recites the outcomes and do not include any details on how the outcomes are accomplished. Additionally, limitation (j) and (k) merely indicate a field of use or technological environment in which the judicial exception is performed. This type of limitation merely confines the use of the abstract idea to a particular technological environment (machine learning and LLMs) and thus fails to add an inventive concept to the claims. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application (Step 2A, Prong Two: NO), and the claim is directed to an abstract idea (Step 2A: YES).
The claim does not include additional elements that are sufficient to amount to more than the judicial exception. As discussed above, the recitation of a computer to perform limitations (a) – (i) amount to no more than mere instructions to apply the exception using a generic computer component. Even when considered in combination, these additional elements represent mere instructions to implement an abstract idea or other exception on a computer and insignificant extra-solution activity, which do not provide an inventive concept (Step 2B).
Regarding claims 13 and 19, the claims are rejected with similar analysis to claim 1.
Similarly, dependent claims 3-5, 7-12, 15-17, and 20 include additional steps that are considered abstract ideas because they fail to provide meaningful significance that goes beyond generally linking the use of an abstract idea to a particular technological environment and using the computer to perform an abstract idea.
Claims 3, 15, and 20 recite using a generic computer component to perform the mental process of generating a justification for a response.
Claims 4 and 16 read on a person creating a prompt using a text segment, a parameter, and an evaluation format and inputting the prompt into ChatGPT using a generic computer.
Claims 5 and 17 recite a person determining a specific answer format to include in a prompt.
Claims 7, 9, and 11 read on someone using a generic computer to determine a response to a parameter based on a range of determined scores.
Claims 8 and 12 read on someone using a generic computer to determine justifications for the score evaluation (best and worst) of a communication segment corresponding to a parameter.
Claim 10 reads on someone using a generic computer to determine a summary of the justifications for the evaluations of each communication segment corresponding to a parameter.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-5, 11, 13, 15-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Anwade et al. (US 20250200491 A1; hereinafter referred to as Anwade) in view of Kotaru (US 20250259096 A1; hereinafter referred to as Kotaru).
Regarding claim 1, Anwade discloses: a method comprising: receiving a set of communication records, the set of communication records representing one or more communications between a first person and a second person ([0107] An interaction may be an interaction or a recording of an interaction via a digital channel, e.g. a text-based chat such as an online chat or a text chat using an application between an agent device of an agent and a customer device of a customer. An interaction recording can be a textual representation of an interaction such as an audio transcription, a chat transcription an email or a transcript of any other form of digital communication);
receiving a set of analytical parameters associated with the set of communication records ([0117] evaluation questions and evaluation parameters of evaluation forms such as “Number of elevated calls to a supervisor for an agent” may be extracted from an evaluation form and used in the generation of an evaluation prompt. The extraction of content present in evaluation forms may be analyzed using Gen AI-based LLMs);
generating a plurality of segments ([0108] An interaction, e.g. an interaction recording such as a transcript of a conversation between an agent and a customer may include interaction data items. Interaction data items can be excerpts or snippets of an interaction recording, e.g. input from an agent or input from a customer in an interaction) from the set of communication records… ([0108] a plurality of evaluation prompts for evaluating interaction data items of one or more interactions may be created);
for each analytical parameter in the set of analytical parameters: determining a subset of segments semantically associated with the respective analytical parameter… ([0117] a service may identify chunks of an interaction recording that are most relevant to an evaluation prompt created by a user by identifying a similarity of chunks, e.g. chunks that may include interaction data items, and a prompt, e.g. an evaluation prompt. This may be done by computing a similarity between a vector representation of a prompt and vectors of chunks, e.g. to identify a semantic similarity between chunks of an interaction recording and the evaluation prompt. The evaluation prompt can contain parameters.);
generating, using a trained large language model ("LLM") ([0117] Gen AI-based large language models (LLM) provided by Open AI 606A and LangChain 606B by LangChain Inc. as a framework to implement LLM in applications may be used to generate evaluation prompts from evaluation prompt templates 602 and transcripts from interaction analytics 606C which may include interaction data items), an evaluation of each segment of the respective subset of segments ([0118] Evaluation results 608 may include one or more evaluation scores. The number of evaluation scores may depend on the number of created evaluation prompts that are used in the assessment of an interaction. For example, in case that three evaluation prompts are created, three interaction data items within an interaction may be evaluated and three evaluation scores may be obtained that can be compared, e.g. to pre-set threshold values for the performance indicators) with respect to the respective analytical parameter ([0109] Evaluation prompts may include threshold parameters for interaction data items present in interactions and may allow comparing an interaction data item for an interaction, e.g. time for handling a customer agent interaction, with a threshold parameter for the interaction data items);
and generating, using the trained LLM, a response to the analytical parameter ([0109] Generated evaluation results may include answers to evaluation prompts that have been identified in an interaction. They may further include reasons for an evaluation result, may identify interaction data items that are above a threshold, e.g. meet or exceed expectations when compared to a threshold value for an interaction data item) based on the evaluations of the segments ([0118-0119] in case that three evaluation prompts are created, three interaction data items within an interaction may be evaluated and three evaluation scores may be obtained that can be compared, e.g. to pre-set threshold values for the performance indicators… The identification of evaluation results may include the comparison of one or more performance indicators, identified in the interaction data items or determined by calculating the aggregation of each focus area and behavior combination 613, with pre-set threshold values 614 for the performance indicators. The evaluation result is based on the performance indicators from evaluating the interaction data items.);
and outputting a full evaluation of the set of communication records ([0196] An example output of an evaluation result 924 for an evaluation prompt created for the evaluation of an interaction data item may include: an answer to the question present in the prompt, a reason for the answer, positive feedback, negative feedback, suggestions for improvements in dealing with the question and a grade how a question was handled) based on the set of analytical parameters and the respective generated responses to the analytical parameters ([0140] An evaluation using machine learning such as Gen AI in combination with a large language model and word embedding may allow evaluating an interaction. For example, each question of an evaluation form may be embedded into an evaluation prompt that is used to query an interaction transcript with the help of machine learning, e.g. Gen AI. A generated evaluation result may include items for a question such as, for example: an answer to a question, a reason for the answer, positive feedback, negative feedback, suggestions for improvements, and a grade for an assessed interaction date item).
Anwade does not explicitly, but Kotaru teaches: generating, using a trained machine learning ("ML") model, analytical parameter embeddings for each analytical parameter of the set of analytical parameters ([0035-0036] each content segment 206A-C may be fed into a ML model 208 (e.g., a foundation model or a T5-base model) to generate one or more questions 210… Each generated question 210 is then fed into embedding model 414A to transform the generated question 210 into a word embedding (e.g., vector));
generating, using the trained ML model, segment embeddings for each segment of the plurality of segments… ([0034] the text of each content segment 206 may be tokenized and converted to a word embedding vector for consumption by an LLM. For example, embedding model 214A (e.g., sentence-BERT all-MiniLM-L6-v2 embedding model) may be used to transform each of the content segments 206 into word embedding vectors);
based on determined similarities between the analytical parameter embeddings and each of the segment embeddings ([0036] Content extractor 224 then compares the word embedding of the generated question 210A to the word embeddings of each content segment 206 to determine the most relevant (“top-K”) content segments 216 based on semantic similarity (e.g., based on the distance between corresponding word embeddings));
Anwade and Kotaru are considered analogous in the field of machine learning models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Anwade to combine the teachings of Kotaru because doing so would improve LLM evaluation of a corpus of text by using a trained finetuned embedding model to determine semantic similarities between question/query embeddings and text content embeddings, leading to more accurate LLM responses to an analytical parameter (Kotaru [0047] while embedding models may be finetuned using annotated datasets, generating these datasets for niche and specialized domains is an extremely arduous, time-consuming, and expensive task. The present application overcomes these issues by implementing a generative pseudo labeling (GPL) approach to creating labeled data for finetuning an existing embedding model to recognize semantic similarities within a specialized domain while maintaining understanding of semantic similarities in regular language. In this way, the accuracy and relevance of foundation-model responses to domain-specific queries is substantially improved, as illustrated by FIG. 2C).
Regarding claim 3, the combination of Anwade and Kotaru teaches: the method of claim 1. Anwade further teaches: wherein generating the response to the analytical parameter comprises generating, using the trained LLM, a justification associated with the response to the analytical parameter ([0109] Generated evaluation results may include answers to evaluation prompts that have been identified in an interaction. They may further include reasons for an evaluation result).
Regarding claim 4, the combination of Anwade and Kotaru teaches: the method of claim 1. Anwade further teaches: further comprising: generating, for each analytical parameter and associated segment ([0108] An interaction, e.g. an interaction recording such as a transcript of a conversation between an agent and a customer may include interaction data items. Interaction data items can be excerpts or snippets of an interaction recording, e.g. input from an agent or input from a customer in an interaction) of the respective subset of segments, a prompt for the LLM, the prompt comprising the analytical parameter, the segment ([0117] implement LLM in applications may be used to generate evaluation prompts from evaluation prompt templates 602 and transcripts from interaction analytics 606C which may include interaction data items), and an indication of an evaluation format to be generated ([0117] evaluation questions and evaluation parameters of evaluation forms such as “Number of elevated calls to a supervisor for an agent” may be extracted from an evaluation form and used in the generation of an evaluation prompt);
providing the prompt to the LLM ([0117] Input to a LLM may be in form of a prompt, e.g. an evaluation prompt or a training recommendation prompt or may include input that may be derived from an interaction recording);
and wherein generating the respective evaluation of each segment is based on the prompt ([0105] A processor such as processor 403 of computing device 402 processor 411 of device 410, and/or processor 421 of computing device 420 may be configured to generate evaluation results for the interaction data items using the plurality of evaluation prompts and machine learning).
Regarding claim 5, the combination of Anwade and Kotaru teaches: the method of claim 4. Anwade further teaches: wherein the evaluation format comprises one of (a) a “yes” or "no" answer, (b) a selection from multiple choices, (c) a numerical rating ([0109] may identify interaction data items that are below a threshold, e.g. don't meet expectations when compared to a threshold value for an interaction data item and can be summarized in a grade for an evaluation question within a certain scale, e.g. customer satisfaction for handling a call was rated 7 out of 10 (on a scale from 1 to 10, 10 being the highest and 1 being the lowest score)), or (d) a free-form answer ([0108] Evaluation prompts may include, for example, questions that are used in the evaluation of an agent such as “What is the average handling time for a call for an agent?”).
Regarding claim 11, the combination of Anwade and Kotaru teaches: the method of claim 1. Anwade further teaches: wherein generating, using the trained LLM, the response to the analytical parameter comprises: generating the response based on the respective evaluation for the analytical parameter ([0109] Evaluation prompts may include threshold parameters for interaction data items present in interactions and may allow comparing an interaction data item for an interaction, e.g. time for handling a customer agent interaction, with a threshold parameter for the interaction data items, e.g. average handling time of customer agent interactions among all agents of a contact center. Generated evaluation results may include answers to evaluation prompts that have been identified in an interaction) having a worst score of the respective evaluations ([0109] They may further include reasons for an evaluation result, may identify interaction data items that are above a threshold, e.g. meet or exceed expectations when compared to a threshold value for an interaction data item, may identify interaction data items that are below a threshold, e.g. don't meet expectations when compared to a threshold value for an interaction data item and can be summarized in a grade for an evaluation question within a certain scale, e.g. customer satisfaction for handling a call was rated 7 out of 10 (on a scale from 1 to 10, 10 being the highest and 1 being the lowest score). A worst score can be a score for a interaction data item that doesn’t meet a threshold score.).
Regarding claim 13, Anwade discloses: a system comprising: a non-transitory computer-readable medium; and one or more processors communicatively connected to the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to cause the one or more processors… ([0102] Embodiments of the invention may include one or more article(s) (e.g. memory 320 or storage 330) such as a computer or processor non-transitory readable medium, or a computer or processor non-transitory storage medium). The rest of the claim recites similar limitations as claim 1 and therefore is rejected similarly.
Regarding claim 15, it recites similar limitations as claim 3 and therefore is rejected similarly.
Regarding claim 16, it recites similar limitations as claim 4 and therefore is rejected similarly.
Regarding claim 17, it recites similar limitations as claim 5 and therefore is rejected similarly.
Regarding claim 19, Anwade discloses: a non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors… ([0102] Embodiments of the invention may include one or more article(s) (e.g. memory 320 or storage 330) such as a computer or processor non-transitory readable medium, or a computer or processor non-transitory storage medium). The rest of the claim recites similar limitations as claim 1 and therefore is rejected similarly.
Regarding claim 20, it recites similar limitations as claim 3 and therefore is rejected similarly.
Claims 7-8, 10, and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Anwade in view of Kotaru, as applied to claims 1, 3-5, 11, 13, 15-17, and 19-20 above, and further in view of Desgarennes et al. (US 20240289686 A1; hereinafter referred to as Desgarennes).
Regarding claim 7, the combination of Anwade and Kotaru teaches: the method of claim 1. The combination of Anwade and Kotaru does not explicitly, but Desgarennes teaches: wherein generating, using the trained LLM, the response to the analytical parameter comprises: generating the response based on the respective evaluation for the analytical parameter having a best score of the respective evaluations ([0077] the evaluation may be performed by generating a confidence score for one or more components of the output and comparing the one or more confidence scores to a threshold value. The confidence score may be determined using evaluation metrics developed for the ML model and the output. The confidence score may be a measure of the output's responsiveness to the input and how well the output satisfies the environment guidelines based on the one or more evaluation metrics).
Anwade, Kotaru, and Desgarennes are considered analogous in the field of large language models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Anwade and Kotaru to combine the teachings of Desgarennes because doing so would allow for output from a LLM to be evaluated using a scoring method to determine responsiveness to an input prompt, leading to improved responses from the LLM and better evaluations for analytical parameters (Desgarennes [0077] The confidence score may be determined using evaluation metrics developed for the ML model and the output. The confidence score may be a measure of the output's responsiveness to the input and how well the output satisfies the environment guidelines based on the one or more evaluation metrics. At operation 214, the output evaluator may attempt to modify one or more aspects of the model output and/or data mask portions of the output to make the output responsive to the input and/or the guidelines as required).
Regarding claim 8, the combination of Anwade, Kotaru, and Desgarennes teaches: the method of claim 7. Anwade further teaches: wherein generating, using the trained LLM, the response to the analytical parameter further comprises: generating, using the trained LLM, a justification based on the respective evaluation for the analytical parameter having the best score of the respective evaluations ([0109] Generated evaluation results may include answers to evaluation prompts that have been identified in an interaction. They may further include reasons for an evaluation result, may identify interaction data items that are above a threshold, e.g. meet or exceed expectations when compared to a threshold value for an interaction data item);
and wherein outputting the full evaluation comprises outputting the justification ([0140] A generated evaluation result may include items for a question such as, for example: an answer to a question, a reason for the answer, positive feedback, negative feedback, suggestions for improvements, and a grade for an assessed interaction date item).
Regarding claim 10, the combination of Anwade, Kotaru, and Desgarennes teaches: the method of claim 7. Anwade further teaches: wherein generating, using the trained LLM ([0137] A system may use generative AI models and large language models to generate evaluation results (810) from QM evaluation forms and interaction transcripts), the response to the analytical parameter further comprises: generating, using the trained LLM, a justification for each respective evaluation for the analytical parameter ([0109] Generated evaluation results may include answers to evaluation prompts that have been identified in an interaction. They may further include reasons for an evaluation result, may identify interaction data items that are above a threshold, e.g. meet or exceed expectations when compared to a threshold value for an interaction data item, may identify interaction data items that are below a threshold, e.g. don't meet expectations when compared to a threshold value for an interaction data item and can be summarized in a grade for an evaluation question);
generating, using the trained LLM, a summary justification ([0231] Table 1 is an example summary of evaluation results for an agent 1. The evaluation results shown in Table 1 may include focus areas, e.g. CSAT and Productivity and analyzed behaviors of interactions for the focus areas) based on the generated justifications for the respective evaluation for the analytical parameters ([0196] An example output of an evaluation result 924 for an evaluation prompt created for the evaluation of an interaction data item may include: an answer to the question present in the prompt, a reason for the answer, positive feedback, negative feedback, suggestions for improvements in dealing with the question and a grade how a question was handled);
and wherein outputting the full evaluation comprises outputting the justification ([0140] A generated evaluation result may include items for a question such as, for example: an answer to a question, a reason for the answer, positive feedback, negative feedback, suggestions for improvements, and a grade for an assessed interaction date item).
Regarding claim 12, the combination of Anwade, Kotaru, and Desgarennes teaches: the method of claim 7. Anwade further teaches: wherein generating, using the trained LLM, the response to the analytical parameter further comprises: generating, using the trained LLM, a justification based on the respective evaluation for the analytical parameter having the worst score of the respective evaluations ([0109] They may further include reasons for an evaluation result, may identify interaction data items that are above a threshold, e.g. meet or exceed expectations when compared to a threshold value for an interaction data item, may identify interaction data items that are below a threshold, e.g. don't meet expectations when compared to a threshold value for an interaction data item and can be summarized in a grade for an evaluation question within a certain scale, e.g. customer satisfaction for handling a call was rated 7 out of 10 (on a scale from 1 to 10, 10 being the highest and 1 being the lowest score). A worst score can be a score for a interaction data item that doesn’t meet a threshold score.);
and wherein outputting the full evaluation comprises outputting the justification ([0140] A generated evaluation result may include items for a question such as, for example: an answer to a question, a reason for the answer, positive feedback, negative feedback, suggestions for improvements, and a grade for an assessed interaction date item).
Claims 9 is rejected under 35 U.S.C. 103 as being unpatentable over Anwade in view of Kotaru, as applied to claims 1, 3-5, 11, 13, 15-17, and 19-20 above, and further in view of Fallon (US 12058091 B1).
Regarding claim 9, the combination of Anwade and Kotaru teaches: the method of claim 1. The combination of Anwade and Kotaru does not explicitly, but Fallon teaches: wherein generating, using the trained LLM, the response to the analytical parameter comprises: generating the response based on an averaging of the respective evaluations for the analytical parameter ([col 4, lines 59-67] When the communication spaces and conversations of the identified communication group have been processed as determined at operations 235 or 250, an overall score for the identified communication group is determined at operation 255 based on the similarity scores of the communication spaces and conversations of the identified communication group. The similarity scores may be combined in any fashion (e.g., summed, averaged, etc.) to produce the overall score).
Anwade, Kotaru, and Fallon are considered analogous in the field of large language models. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Anwade and Kotaru to combine the teachings of Fallon because doing so would allow for evaluation scores for a set communications to be combined to determine an overall score for evaluation, leading to more accurate LLM output that is based on the average score (Fallon [col 8, lines 4-10] The communication spaces of each communication group are compared to the communication space to determine the weighted attribute scores for the communication group. The attribute scores for communication spaces and conversations of the communication group are combined (e.g., summed, averaged, etc.) to produce the overall score for the communication group as described above).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Nathan Tengbumroong whose telephone number is (703)756-1725. The examiner can normally be reached Monday - Friday, 11:30 am - 8:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NATHAN TENGBUMROONG/Examiner, Art Unit 2654
/HAI PHAN/Supervisory Patent Examiner, Art Unit 2654