DETAILED ACTION
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over U.S Pub. No. 2023/0291835 A1 to Ramprashad et al. (hereinafter “Ramprashad”) in view of U.S Patent No. 12,126,769 B1 to Koul et al. (hereinafter “Koul”).
Regarding claim 1, Ramprashad teaches an artificial intelligence (AI)-based call response system for providing a context-based recommendation during a monitored conversation (paragraphs [0006]- [0007] and [0046]; systems/processes for providing real-time contact center monitoring, alerting and analytics, while ensuring appropriate treatment of sensitive customer information), comprising:
one or more processors and a non-transitory computer readable medium operably coupled thereto, the non-transitory computer readable medium comprising a plurality of instructions stored in association therewith that are accessible to, and executable by, the one or more processors, to perform conversation analysis operations (paragraphs [0020] and [0049]; computer executable instructions, embodied in non-transitory media, for implementing parts or all of the systems and processes), which comprise:
determining transcribed words for the monitored conversation (paragraphs [0016] and [0048]; systems/processes for telephonic contact center monitoring in which: (a) at least the following steps are performed within a first (higher) security zone: (i) receiving, in real time, contact center telephony data indicative of multiple agent-caller communications; (ii) separating, in real time, the received telephony data into tagged utterances, each representing a single utterance spoken by either an agent or a caller);
analyzing the transcribed words using one or more machine learning models to produce a score associated with a model identifier (ID) identifying a machine learning model of the one or more machine learning models (paragraphs [0045] and [0046]; utilizing natural language processing (“NLP”)/machine learning (“ML”) technique is used to identify critical calls (e.g., customers likely to leave, angry customers, agent misbehavior, etc.) immediately upon their transcription. (In fact, such determination need not await complete transcription of the call, but may proceed in real time while the call is still in progress.) Because the critical call classifier makes its determination based upon the sanitized ASR transcripts, it can be alternatively located within the lower security zone);
comparing the score to a predefined threshold of the machine learning model; generating an alert when the score meets or exceeds the predefined threshold, the alert comprising the model ID and a call identifier (ID) identifying the monitored conversation (paragraphs [0046]- [0048]; once a call is identified as critical, an immediate alert is sent to a critical response team that operates within the high security zone).
Ramprashad does not explicitly teach creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response; retrieving the response for each of the one or more prompts; and providing the response to a user.
In the same field of endeavor, Koul discloses creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response (column 11, lines 3-25; a LLM model can be used to create this reference dataset from customer conversational data or recordings. Once the ground-truth data is generated, it is used to refine or fine-tune smaller versions of the transcription/summarization models and possibly other language models (LLMs). Fine-tuning involves adjusting the model's parameters to make it perform better on specific tasks or data. In this case, the goal is to improve the model's performance on customer conversational data. The fine-tuned models are expected to provide improved accuracy when processing customer conversational data, which can be crucial for applications like natural language understanding and processing. The refined models can be optimized for real-time processing, meaning they can quickly analyze and understand customer conversations as they happen); retrieving the response for each of the one or more prompts; and providing the response to a user (column 6, lines 33-54 and column 8, lines 51-65; summarization model 310 is a generative artificial intelligence model trained to provide a summary of a conversation based on a prompt. According to an embodiment, the prompt comprises at least context-aware transcription data. In some implementations, summarization model 310 leverages LLMs to provide GAI capabilities. In some implementations, based on CTI event data (or other contact center data), the prompt generated for summarization model 310 will explicitly identify the direction of the call, which the LLM model can automatically include in the summarization).
At the time of the effective filing date of the invention, it would have been obvious to a person of ordinary skilled in the art to modify Ramprashad’s teaching with a feature of creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response; retrieving the response for each of the one or more prompts; and providing the response to a user as taught by Koul in order to provide real time assistance to an agent (Abstract, Koul).
Regarding claim 2, Ramprashad teaches the AI-based call response system of claim 1, wherein the conversation analysis operations further comprise:
registering the transcribed words with the call ID; and storing the transcribed words registered with the call ID in a storage (paragraphs [0007] and [0016]; using a privacy-filtering ASR engine to process each utterance, in real time, into a corresponding sanitized ASR transcription; and receiving, in real time, a critical call alert; and at least the following steps are performed within a second lower security zone: updating, in real time, a database to include each sanitized ASR transcription).
Regarding claim 3, Ramprashad teaches the AI-based call response system of claim 2, wherein the creating the one or more prompts comprises: retrieving, from the storage, a model description of the model ID associated with the alert, the stored transcribed words corresponding to the call ID associated with the alert, or a combination thereof; and generating the executable instruction based on the model description, the transcribed words, or the combination thereof (paragraphs [0007] and [0016]; using a privacy-filtering ASR engine to process each utterance, in real time, into a corresponding sanitized ASR transcription; and receiving, in real time, a critical call alert; and at least the following steps are performed within a second lower security zone: updating, in real time, a database to include each sanitized ASR transcription).
Regarding claim 4, Ramprashad teaches the AI-based call response system of claim 1, wherein the response comprises:
a summary of an interaction between a customer and an agent during the monitored conversation, an insight of the monitored conversation that includes an in-context explanation of the interaction capturing a behavior of the agent, or a recommendation that includes one or more in-context responses that follow definitions based on an experience of the customer during the monitored conversation (paragraphs [0004] and [0016]; timely monitoring and reporting of customer-agent interactions is more important than ever. Preferably, such monitoring should include both analytics to gauge overall customer sentiment, agent performance and to spot trends. Furthermore, for optimal results, such monitoring should be available in real time or near real time).
Regarding claim 5, Ramprashad teaches the AI-based call response system of claim 4, wherein the providing the response to the user comprises:
communicating the response to an external application, wherein the response comprises the summary, the insight, the recommendation, or a combination thereof (paragraphs [0004] and [0016]; timely monitoring and reporting of customer-agent interactions is more important than ever. Preferably, such monitoring should include both analytics to gauge overall customer sentiment, agent performance and to spot trends. Furthermore, for optimal results, such monitoring should be available in real time or near real time).
Regarding claim 6, Ramprashad teaches the AI-based call response system of claim 4, wherein the monitored conversation is a phone call, and wherein the providing the response to the user comprises: providing the recommendation to the agent in a written text during the monitored conversation (paragraphs [0049] and [0062]; privacy-filtering ASR engine is configured to output both unredacted and redacted text. The unredacted text is maintained within the higher security zone, where it can be fed to the critical call classifier, also located in the higher security zone. The purpose of this arrangement is to facilitate quicker and more accurate identification of critical calls, by reducing the informational “noise” or uncertainty that redaction can add. Additionally, in this embodiment, the critical response team has access to unredacted ASR text, via the speech browser).
Regarding claim 7, Ramprashad teaches the Al-based call response system of claim 4, wherein the monitored conversation is a chat, and wherein the providing the response to the user comprises: providing the recommendation to the customer in a written text during the monitored conversation (paragraphs [0049] and [0055]; privacy-filtering ASR engine is configured to output both unredacted and redacted text. The unredacted text is maintained within the higher security zone, where it can be fed to the critical call classifier, also located in the higher security zone. The purpose of this arrangement is to facilitate quicker and/or more accurate identification of critical calls, by reducing the informational “noise” or uncertainty that redaction can add. Additionally, in this embodiment, the critical response team has access to unredacted ASR text, via the speech browser).
Regarding claim 8, Ramprashad teaches the AI-based call response system of claim 1, wherein the conversation analysis operations further comprise:
receiving a new set of transcribed words after a new word is transcribed during the monitored conversation (paragraphs [0007] and [0049]; using a privacy-filtering ASR engine to process each utterance, in real time, into a corresponding sanitized ASR transcription);
analyzing the new set of transcribed words to produce an updated score (paragraphs [0007] and [0011]; updating, in real time, a database to include the sanitized ASR transcription);
generating, based on the updated score, an updated alert comprising a different model ID with a different model description; creating a new set of prompts based on the updated alert with each new prompt comprising a new executable instruction that prompts the LLM for a new response; and generating the new response different from the response (paragraphs [0046]- [0048]; once a call is identified as critical, an immediate alert is sent to a critical response team that operates within the high security zone).
Regarding claim 9, Ramprashad teaches A method for providing a context-based recommendation during a monitored conversation (paragraphs [0006]- [0007] and [0046]; systems/processes for providing real-time contact center monitoring, alerting and analytics, while ensuring appropriate treatment of sensitive customer information), the method comprising:
determining, via an automatic speech recognition system, transcribed words for the monitored conversation (paragraphs [0016] and [0048]; systems/processes for telephonic contact center monitoring in which: (a) at least the following steps are performed within a first (higher) security zone: (i) receiving, in real time, contact center telephony data indicative of multiple agent-caller communications; (ii) separating, in real time, the received telephony data into tagged utterances, each representing a single utterance spoken by either an agent or a caller);
analyzing the transcribed words using one or more machine learning models to produce a score associated with a model identifier (ID) identifying a machine learning model of the one or more machine learning models (paragraphs [0045] and [0046]; utilizing natural language processing (“NLP”)/machine learning (“ML”) techniques—is used to identify critical calls (e.g., customers likely to leave, angry customers, agent misbehavior, etc.) immediately upon their transcription. (In fact, such determination need not await complete transcription of the call, but may proceed in real time while the call is still in progress.) Because the critical call classifier makes its determination based upon the sanitized ASR transcripts, it can be alternatively located within the lower security zone);
comparing the score to a predefined threshold of the machine learning model; generating an alert when the score meets or exceeds the predefined threshold, the alert comprising the model ID and a call identifier (ID) identifying the monitored conversation (paragraphs [0046]- [0048]; once a call is identified as critical, an immediate alert is sent to a critical response team that operates within the high security zone).
Ramprashad does not explicitly teach creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response; retrieving the response for each of the one or more prompts; and providing the response to a user.
In the same field of endeavor, Koul discloses creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response; retrieving the response for each of the one or more prompts; and providing the response to a user (column 11, lines 3-25; a LLM model can be used to create this reference dataset from customer conversational data or recordings. Once the ground-truth data is generated, it is used to refine or fine-tune smaller versions of the transcription/summarization models and possibly other language models (LLMs). Fine-tuning involves adjusting the model's parameters to make it perform better on specific tasks or data. In this case, the goal is to improve the model's performance on customer conversational data. The fine-tuned models are expected to provide improved accuracy when processing customer conversational data, which can be crucial for applications like natural language understanding and processing. The refined models can be optimized for real-time processing, meaning they can quickly analyze and understand customer conversations as they happen); retrieving the response for each of the one or more prompts; and providing the response to a user (column 6, lines 33-54 and column 8, lines 51-65; summarization model 310 is a generative artificial intelligence model trained to provide a summary of a conversation based on a prompt. According to an embodiment, the prompt comprises at least context-aware transcription data. In some implementations, summarization model 310 leverages LLMs to provide GAI capabilities. In some implementations, based on CTI event data (or other contact center data), the prompt generated for summarization model 310 will explicitly identify the direction of the call, which the LLM model can automatically include in the summarization).
At the time of the effective filing date of the invention, it would have been obvious to a person of ordinary skilled in the art to modify Ramprashad’s teaching with a feature of creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response; retrieving the response for each of the one or more prompts; and providing the response to a user as taught by Koul in order to provide real time assistance to an agent (Abstract, Koul).
Regarding claim 10, Ramprashad teaches the method of claim 9, further comprising:
registering the transcribed words with the call ID; and storing the transcribed words registered with the call ID in a storage (paragraphs [0007] and [0016]; using a privacy-filtering ASR engine to process each utterance, in real time, into a corresponding sanitized ASR transcription; and receiving, in real time, a critical call alert; and at least the following steps are performed within a second lower security zone: updating, in real time, a database to include each sanitized ASR transcription).
Regarding claim 11, Ramprashad teaches the method of claim 10, wherein the creating the one or more prompts comprises:
retrieving, from the storage, a model description of the model ID associated with the alert, the stored transcribed words corresponding to the call ID associated with the alert, or a combination thereof; and generating the executable instruction based on the model description, the transcribed words, or the combination thereof (paragraphs [0007] and [0016]; using a privacy-filtering ASR engine to process each utterance, in real time, into a corresponding sanitized ASR transcription; and receiving, in real time, a critical call alert; and at least the following steps are performed within a second lower security zone: updating, in real time, a database to include each sanitized ASR transcription).
Regarding claim 12, Ramprashad teaches the method of claim 9, wherein the response comprises:
a summary of an interaction between a customer and an agent during the monitored conversation, an insight of the monitored conversation that includes an in-context explanation of the interaction capturing a behavior of the agent, or a recommendation that includes one or more in-context responses that follow definitions based on an experience of the customer during the monitored conversation (paragraphs [0004] and [0016]; timely monitoring and reporting of customer-agent interactions is more important than ever. Preferably, such monitoring should include both analytics to gauge overall customer sentiment, agent performance and to spot trends. Furthermore, for optimal results, such monitoring should be available in real time or near real time).
Regarding claim 13, Ramprashad teaches the method of claim 12, wherein the providing the response to the user comprises: communicating the response to an external application, wherein the response comprises the summary, the insight, the recommendation, or a combination thereof (paragraphs [0004] and [0016]; timely monitoring and reporting of customer-agent interactions is more important than ever. Preferably, such monitoring should include both analytics to gauge overall customer sentiment, agent performance and to spot trends. Furthermore, for optimal results, such monitoring should be available in real time or near real time).
Regarding claim 14, Ramprashad teaches the method of claim 12, wherein the monitored conversation is a phone call, and wherein the providing the response to the user comprises: providing the recommendation to the agent in a written text during the monitored conversation (paragraphs [0049] and [0062]; privacy-filtering ASR engine is configured to output both unredacted and redacted text. The unredacted text is maintained within the higher security zone, where it can be fed to the critical call classifier, also located in the higher security zone. The purpose of this arrangement is to facilitate quicker and more accurate identification of critical calls, by reducing the informational “noise” or uncertainty that redaction can add. Additionally, in this embodiment, the critical response team has access to unredacted ASR text, via the speech browser).
Regarding claim 15, Ramprashad teaches the method of claim 12, wherein the monitored conversation is a chat, and wherein the providing the response to the user comprises: providing the recommendation to the customer in a written text during the monitored conversation (paragraphs [0049] and [0055]; privacy-filtering ASR engine is configured to output both unredacted and redacted text. The unredacted text is maintained within the higher security zone, where it can be fed to the critical call classifier, also located in the higher security zone. The purpose of this arrangement is to facilitate quicker and/or more accurate identification of critical calls, by reducing the informational “noise” or uncertainty that redaction can add. Additionally, in this embodiment, the critical response team has access to unredacted ASR text, via the speech browser).
Regarding claim 16, Ramprashad teaches the method of claim 9, further comprising:
receiving a new set of transcribed words after a new word is transcribed during the monitored conversation (paragraphs [0007] and [0049]; using a privacy-filtering ASR engine to process each utterance, in real time, into a corresponding sanitized ASR transcription);
analyzing the new set of transcribed words to produce an updated score (paragraphs [0007] and [0011]; updating, in real time, a database to include the sanitized ASR transcription);
generating, based on the updated score, an updated alert comprising a different model ID with a different model description; creating a new set of prompts based on the updated alert with each new prompt comprising a new executable instruction that prompts the LLM for a new response; and generating the new response different from the response (paragraphs [0046]- [0048]; once a call is identified as critical, an immediate alert is sent to a critical response team that operates within the high security zone).
Regarding claim 17, Ramprashad teaches a non-transitory computer-readable medium having stored thereon computer-readable instructions executable to provide a context-based recommendation during a monitored conversation using an artificial intelligence (AI)-based call response system, the computer-readable instructions executable to perform conversation analysis operations (paragraphs [0020] and [0049]; computer executable instructions, embodied in non-transitory media, for implementing parts or all of the systems and processes);which comprise:
determining transcribed words for the monitored conversation (paragraphs [0016] and [0048]; systems/processes for telephonic contact center monitoring in which: (a) at least the following steps are performed within a first (higher) security zone: (i) receiving, in real time, contact center telephony data indicative of multiple agent-caller communications; (ii) separating, in real time, the received telephony data into tagged utterances, each representing a single utterance spoken by either an agent or a caller);
analyzing the transcribed words using one or more machine learning models to produce a score associated with a model identifier (ID) identifying a machine learning model of the one or more machine learning models (paragraphs [0045] and [0046]; utilizing natural language processing (“NLP”)/machine learning (“ML”) technique is used to identify critical calls (e.g., customers likely to leave, angry customers, agent misbehavior, etc.) immediately upon their transcription. (In fact, such determination need not await complete transcription of the call, but may proceed in real time while the call is still in progress.) Because the critical call classifier makes its determination based upon the sanitized ASR transcripts, it can be alternatively located within the lower security zone);
comparing the score to a predefined threshold of the machine learning model; generating an alert when the score meets or exceeds the predefined threshold, the alert comprising the model ID and a call identifier (ID) identifying the monitored conversation (paragraphs [0046]- [0048]; once a call is identified as critical, an immediate alert is sent to a critical response team that operates within the high security zone).
Ramprashad does not explicitly teach creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response; retrieving the response for each of the one or more prompts; and providing the response to a user.
In the same field of endeavor, Koul discloses creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response; retrieving the response for each of the one or more prompts; and providing the response to a user (column 11, lines 3-25; a LLM model can be used to create this reference dataset from customer conversational data or recordings. Once the ground-truth data is generated, it is used to refine or fine-tune smaller versions of the transcription/summarization models and possibly other language models (LLMs). Fine-tuning involves adjusting the model's parameters to make it perform better on specific tasks or data. In this case, the goal is to improve the model's performance on customer conversational data. The fine-tuned models are expected to provide improved accuracy when processing customer conversational data, which can be crucial for applications like natural language understanding and processing. The refined models can be optimized for real-time processing, meaning they can quickly analyze and understand customer conversations as they happen); retrieving the response for each of the one or more prompts; and providing the response to a user (column 6, lines 33-54 and column 8, lines 51-65; summarization model 310 is a generative artificial intelligence model trained to provide a summary of a conversation based on a prompt. According to an embodiment, the prompt comprises at least context-aware transcription data. In some implementations, summarization model 310 leverages LLMs to provide GAI capabilities. In some implementations, based on CTI event data (or other contact center data), the prompt generated for summarization model 310 will explicitly identify the direction of the call, which the LLM model can automatically include in the summarization).
At the time of the effective filing date of the invention, it would have been obvious to a person of ordinary skilled in the art to modify Ramprashad’s teaching with a feature of creating, based on the alert, one or more prompts with each prompt comprising an executable instruction that prompts, queries, or requests an output from a large language model (LLM) for a response; retrieving the response for each of the one or more prompts; and providing the response to a user as taught by Koul in order to provide real time assistance to an agent (Abstract, Koul).
Regarding claim 18, Ramprashad teaches the non-transitory computer-readable medium of claim 17, wherein the conversation
analysis operations further comprise: registering the transcribed words with the call ID; and storing the transcribed words registered with the call ID in a storage (paragraphs [0007] and [0016]; using a privacy-filtering ASR engine to process each utterance, in real time, into a corresponding sanitized ASR transcription; and receiving, in real time, a critical call alert; and at least the following steps are performed within a second lower security zone: updating, in real time, a database to include each sanitized ASR transcription).
Regarding claim 19, Ramprashad teaches the non-transitory computer-readable medium of claim 18, wherein the creating the one
or more prompts comprises: retrieving, from the storage, a model description of the model ID associated with the alert, the stored transcribed words corresponding to the call ID associated with the alert, or a combination thereof; and generating the executable instruction based on the model description, the transcribed words, or the combination thereof (paragraphs [0007] and [0016]; using a privacy-filtering ASR engine to process each utterance, in real time, into a corresponding sanitized ASR transcription; and receiving, in real time, a critical call alert; and at least the following steps are performed within a second lower security zone: updating, in real time, a database to include each sanitized ASR transcription).
Regarding claim 20, Ramprashad teaches the non-transitory computer-readable medium of claim 17, wherein the response comprises:
a summary of an interaction between a customer and an agent during the monitored conversation, an insight of the monitored conversation that includes an in-context explanation of the interaction capturing a behavior of the agent, or a recommendation that includes one or more in-context responses that follow definitions based on an experience of the customer during the monitored conversation (paragraphs [0004] and [0016]; timely monitoring and reporting of customer-agent interactions is more important than ever. Preferably, such monitoring should include both analytics to gauge overall customer sentiment, agent performance and to spot trends. Furthermore, for optimal results, such monitoring should be available in real time or near real time).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AKELAW A TESHALE whose telephone number is (571)270-5302. The examiner can normally be reached 9 am -6pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, FAN TSANG can be reached at (571) 272-7547. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
AKELAW TESHALE
Primary Examiner
Art Unit 2694
/AKELAW TESHALE/Primary Examiner, Art Unit 2694