DETAILED ACTION
This Office Action is sent in response to the Applicant’s Communication received on 07/23/2026 for application number 18/454,882. The Office hereby acknowledges receipt of the following and placed of record in file: Specification, Drawings, Abstract, Oath/Declaration, IDS, and Claims.
Claims 1, 3, 8, 10, 15, and 17 are amended.
Claims 1-20 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
35 USC 101
In summary of pages 9-13 of the remarks section, the Applicant argues that the newly amended claim limitations ("collecting, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service" and "identifying context of the prompt, by analyzing collected data from the plurality of communication sources, wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image") encompass AI in a manner that cannot be
practically performed in the human mind.
The Examiner respectfully disagrees. The further narrowing from the newly amended claim limitations put the Application in a better position for allowability but fail to add sufficient detail that would overcome the 35 USC 101 rejection. Specifically, the limitation "collecting, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service" was analyzed was falling under an additional element in Step 2A Prong Two. The limitation was interpreted to be mere data gathering recited at a high level of generality. According to MPEP 2106.05(g), this is insignificant extra-solution activity that does not integrate the judicial exception into a practical application and, whether considered or individually, cannot provide for an inventive concept under Step 2B. The combined limitation "identifying context of the prompt, by analyzing collected data from the plurality of communication sources” was analyzed under Step 2A Prong One was being an abstract idea because the actions of “identifying” and “analyzing” are recited at a high level of generality. Therefore, under the Broadest Reasonable Interpretation (BRI), the limitation was interpreted as reciting a mental process. The limitation “wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image" was analyzed under Step 2A Prong Two as being an additional element there merely applies the judicial exception to a computer. The details of the claim limitation are recited at high level of generality. When considered individually or in combination, the claim limitation cannot integrate the judicial exception into a practical application and cannot provide for inventive concept.
35 USC 103
On pages 13 and 14 of the remarks section, the Applicant argues that the cited sections of the applied references, whether taken alone or in any reasonable combination, do not disclose at least "collecting, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service" and "identifying context of the prompt, by analyzing collected data from the plurality of communication sources, wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image," as recited in the amended claim 1. Independent claims 8 and 15, as amended, recite similar features.
Applicant’s arguments with respect to claim(s) 1, 8, and 15 have been considered but are moot because the new ground of rejection relies on a new set of references.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claims 1-7 are directed to a method. Claims 8-14 are directed to a computer program product. Claims 15-20 are directed to a system. Therefore, all claims are directed to one of the four statutory categories of patent eligible subject matter.
Claim 1
Step 2A Prong 1:
Claim 1 recites:
“identifying context of the prompt, by analyzing collected data from the plurality of communication sources;” Identifying context of the prompt, by analyzing collected data from various sources is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“identifying an intent of the prompt, [by using natural language processing] to disambiguate the prompt;” Identifying an intent of the prompt to disambiguate the prompt is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“providing a modified prompt, according to the context and the intent;” Providing a modified prompt, according to the context and the intent is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“selecting datasets relevant to the context;” Selecting datasets relevant to the context is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“generating a response to the prompt, based on the modified prompt and the datasets;” Generating a response to the prompt, based on the modified prompt and the datasets is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“A computer-implemented method for training a context-aware artificial intelligence (AI) chatbot;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“receiving, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“collecting, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“by using natural language processing;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
“wherein the context-aware AI chatbot responds to the prompt with the response;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“A computer-implemented method for training a context-aware artificial intelligence (AI) chatbot;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“receiving, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“collecting, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“by using natural language processing;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
“wherein the context-aware AI chatbot responds to the prompt with the response;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 2
Step 2A Prong 1:
Claim 2 recites:
“determining the level of the ambiguity of the prompt;” Determining the level of the ambiguity of the prompt is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
Step 2A Prong Two and Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Claim 3
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“pre-processing the collected data to provide pre-processed data with correct formats and free from irrelevant words or phrases;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“pre-processing the collected data to provide pre-processed data with correct formats and free from irrelevant words or phrases;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 4
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“applying one or more of tokenization, stemming, lemmatization, and stop word removal to pre-process the collected data;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“applying one or more of tokenization, stemming, lemmatization, and stop word removal to pre-process the collected data;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 5
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“applying one or more of autoregressive integrated moving average, part-of-speech tagging, and named entity recognition to identify the context of the prompt;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“applying one or more of autoregressive integrated moving average, part-of-speech tagging, and named entity recognition to identify the context of the prompt;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 6
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“applying one or more of sentiment analysis, intent recognition, and entity extraction to identify the intent of the prompt;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“applying one or more of sentiment analysis, intent recognition, and entity extraction to identify the intent of the prompt;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 7
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“applying one or more of topic modelling, text categorization, and clustering to select the datasets relevant to the context;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“applying one or more of topic modelling, text categorization, and clustering to select the datasets relevant to the context;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 8
Step 2A Prong 1:
Claim 8 recites:
“identify context of the prompt, by analyzing collected data from the plurality of communication sources;” Identifying context of the prompt, by analyzing collected data from various sources is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“identify an intent of the prompt, [by using natural language processing] to disambiguate the prompt;” Identifying an intent of the prompt to disambiguate the prompt is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“provide a modified prompt, according to the context and the intent;” Providing a modified prompt, according to the context and the intent is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“select datasets relevant to the context;” Selecting datasets relevant to the context is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“generate a response to the prompt, based on the modified prompt and the datasets;” Generating a response to the prompt, based on the modified prompt and the datasets is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“A computer program product for training a context-aware artificial intelligence (AI) chatbot, the computer program product comprising a computer readable storage medium having program instructions stored therewith, the program instructions executable by one or more processors;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“receive, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“by using natural language processing;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
“wherein the context-aware AI chatbot responds to the prompt with the response;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“A computer program product for training a context-aware artificial intelligence (AI) chatbot, the computer program product comprising a computer readable storage medium having program instructions stored therewith, the program instructions executable by one or more processors;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“receive, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“by using natural language processing;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
“wherein the context-aware AI chatbot responds to the prompt with the response;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claims 9-14 are computer program product claims that recite similar limitations to method claims 2-7, respectively. Therefore, claims 9-14 are rejected using the same rationales as claims 2-7, respectively.
Claim 15
Step 2A Prong 1:
Claim 15 recites:
“identify context of the prompt, by analyzing collected data from the plurality of communication sources;” Identifying context of the prompt, by analyzing collected data from various sources is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“identify an intent of the prompt, [by using natural language processing] to disambiguate the prompt;” Identifying an intent of the prompt to disambiguate the prompt is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“provide a modified prompt, according to the context and the intent;” Providing a modified prompt, according to the context and the intent is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“select datasets relevant to the context;” Selecting datasets relevant to the context is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
“generate a response to the prompt, based on the modified prompt and the datasets;” Generating a response to the prompt, based on the modified prompt and the datasets is an action that can be performed mentally with the aid of pen and paper, and is therefore a mental process.
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“A computer system for training a context-aware artificial intelligence (AI) chatbot, the computer system comprising one or more processors, one or more computer readable tangible storage devices, and program instructions stored on at least one of the one or more computer readable tangible storage devices for execution by at least one of the one or more processors;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
“receive, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“by using natural language processing;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
“wherein the context-aware AI chatbot responds to the prompt with the response;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“A computer system for training a context-aware artificial intelligence (AI) chatbot, the computer system comprising one or more processors, one or more computer readable tangible storage devices, and program instructions stored on at least one of the one or more computer readable tangible storage devices for execution by at least one of the one or more processors;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image;” Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)) which cannot provide an inventive concept.
“receive, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“by using natural language processing;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
“wherein the context-aware AI chatbot responds to the prompt with the response;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 16 is a system claim that recites similar limitations to method claims 2. Therefore, claim 16 is rejected using the same rationale as claim 2.
Claim 17
Step 2A Prong Two
This judicial exception is not integrated into a practical application because the additional elements are as follows:
“pre-processing the collected data to provide pre-processed data with correct formats and free from irrelevant words or phrases;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
“wherein the various sources of the collected data include video conferencing services, instant messages, and email messages;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity (MPEP 2106.05(g)).
“wherein one or more of tokenization, stemming, lemmatization, and stop word removal to pre-process the collected data;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements are as follows:
“pre-processing the collected data to provide pre-processed data with correct formats and free from irrelevant words or phrases;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
“wherein the various sources of the collected data include video conferencing services, instant messages, and email messages;” Mere data gathering recited at a high level of generality, and thus are insignificant extra-solution activity. See MPEP 2106.05(g). The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and are well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
“wherein one or more of tokenization, stemming, lemmatization, and stop word removal to pre-process the collected data;” The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)) and cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claims 18-20 are system claims that recite similar limitations to method claims 5-7, respectively. Therefore, claims 18-20 are rejected using the same rationales as claims 5-7, respectively.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 2, 5-9, 12-16, 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Dhrif (US 12592226 B1), hereinafter Dhrif, in view of Paranjape et al. (ART: Automatic multi-step reasoning and tool-use for large language models, published 16 Mar 2023) and Austraat (US 20240232765 A1), hereinafter Austraat.
Regarding claim 1, Dhrif teaches,
A computer-implemented method for training a context-aware artificial intelligence (AI) chatbot [Col 3, lines 55-57, Natural language generation (NLG) is a computer-based process that may be used to produce natural language output... NLG… may be used together as part of a natural language interface system; Col 4, lines 29-32, The various techniques described herein may be used in a variety of contexts, including in natural language processing enabled devices (e.g., devices employing voice control and/or speech processing “voice assistants”); Col 12, lines 16-22, generating pre-computed features based on user feedback data by ranking and arbitration component 140 may generate increasingly contextually rich feature data that may be used to train various machine learning models used to route speech processing request data (e.g., ranking component 120, embedding-based retrieval component 110, etc.)], the method comprising:
Receiving, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity [Col 16, lines 28-37, it may be unclear, based on a user's request, whether the user intended to invoke a movie playback skill or a music playback skill, as a movie and a soundtrack for the movie may be identified using the same name. Accordingly, the decider component 132 may determine that the top two hypotheses of ranking component 120 are equally likely (or approximately equally likely) and may determine that a question should be asked to disambiguate between the two possible actions];
identifying context of the prompt, by analyzing collected data from communication sources [Col 3, lines 66-67 – col 4, lines 1-10, user utterances, input text data, and/or any form of data input to a natural language processing system (“input data”) may be described by “request data” and/or “user request data.” Such request data may change forms many times during processing of the request data by various components of the natural language processing system. For example, initially the request data may be audio data and/or input text data representing a user question… The text data and/or other ASR output data may be transformed into intent data by an NLU component of the speech processing system];
identifying an intent of the prompt, by using natural language processing to disambiguate the prompt [Col 16, lines 28-37, it may be unclear, based on a user's request, whether the user intended to invoke a movie playback skill or a music playback skill, as a movie and a soundtrack for the movie may be identified using the same name. Accordingly, the decider component 132 may determine that the top two hypotheses of ranking component 120 are equally likely (or approximately equally likely) and may determine that a question should be asked to disambiguate between the two possible actions];
Dhrif teaches the above limitations of claim 1 including the context (Dhrif, col 3, lines 66-67 – col 4, lines 1-10), the intent (Dhrif, Col 16, lines 28-37), and the context-aware AI chatbot (Dhrif, Col 4, lines 29-32).
Dhrif does not teach collecting, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service; wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image; providing a modified prompt, according to context and intent; selecting datasets relevant to the context; and generating a response to the prompt, based on the modified prompt and the datasets, wherein chatbot responds to the prompt with the response.
Parajape teaches,
providing a modified prompt, according to context and intent [See Fig 2; Sect 3.1, para 1-2, In Figure 2, ART is presented with a new task description and input instance… ART retrieves similar tasks from a task library (Figure 2(A); Section 3.2), and adds instances of those tasks as demonstrations in the prompt. A demonstration in the task library is written in a specific format, defined by a custom parsing expression grammar (PeG) (Section 3.2). The grammar is defined such that each task instance is decomposed into a sequence of sub-steps];
selecting datasets relevant to the context [Sect 3.1, para 2, ART retrieves similar tasks from a task library (Figure 2(A); Section 3.2), and adds instances of those tasks as demonstrations in the prompt. A demonstration in the task library is written in a specific format, defined by a custom parsing expression grammar (PeG) (Section 3.2). The grammar is defined such that each task instance is decomposed into a sequence of sub-steps; Sect 3.3, para 2, We use SerpAPI, which provides an API for Google search. The input to search is the sequence generated by the LLM after “Qi: [search]”. We extract answer box snippets when they are available or combine the top-2 search result snippets together]; and
generating a response to the prompt, based on the modified prompt and the datasets, wherein the context-aware chatbot responds to the prompt with the response [Sect 3.3, para 3-4, We use the Codex… model for code generation. Input to code generation is the sequence generated by the LM after the sub-task query symbol “Qi : [generate python code]”. This argument is an instruction for code generation and is prompted to Codex as a multi-line comment in Python. For example, in Figure 2, Codex is prompted the instruction ““Use the formula Fx = Ftens * cosine(θ) to solve...”” as a comment and generates “T = 72.0, theta = 35.0, ..., Fx = T*math.cos(radians)”… We run Python code in a virtual Python environment with arithmetic, symbolic, and scientific computing packages pre-installed. The argument to code execute is the previous sub-task’s answer sequence “#(i - 1) :…”, i.e. the python code snippet to be executed. For i = 1, the task input is used as the argument since it potentially contains the code snippet to be executed. In Figure 2, the code snippet generated in the previous step is executed and the value of variable “Fx” is added to the incomplete program; Sect 3.3, para 2, We use SerpAPI, which provides an API for Google search. The input to search is the sequence generated by the LLM after “Qi: [search]”. We extract answer box snippets when they are available or combine the top-2 search result snippets together. For PQA in Figure 2(B), the search query is the original input followed by “What is the formula for the horizontal component of tension force?”, and the output is ““... horizontal component (Fx) can be calculated as Ftens*cosine(θ) ...””.].
Paranjape is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Paranjape and provide modified prompts, and responses to those prompts, in order to output complex answers to user queries by breaking down intricate tasks into a series of simpler tasks. Additionally, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Paranjape and provide selecting datasets relevant to the context in order to gather the pertinent information that is stored in public domains.
Dhriff-Parajape do not teach collecting, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service; wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image.
Austraat teaches,
collecting, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service [Abstract, Systems and methods receive, from a user device through a communication channel, and process, in real-time, a natural language input comprising unstructured data that is derived from an audio signal; Para 0031, The term “transcript data” may be used to refer to a written digital record, in text form, of a single speaker or a written or verbal interaction between multiple participants in a conversation or discussion about various information or content. The transcript data may generally refer to alphanumeric text in digital form. Content can be generated using automatic speech recognition (ASR) and natural language understanding (NLU), which transcribes oral interaction or audio signals during a communication by telephone or video conference and may produce a full-text transcript inclusive of all utterances and disfluencies and/or a summarized version of the interaction or interactions with a focus on only the most semantically or transactionally salient aspects. Alternatively, content may be generated during written exchanges by email, instant messaging, chat messaging, short message service (SMS) text, or other messages exchanged through various online platforms or social media software applications];
wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image [Para 0095, the machine-learning algorithm may include one or more image recognition algorithms suitable to determine one or more categories to which an input, such as data communicated from a visual sensor or a file in JPEG, PNG or other format, representing an image or portion thereof, belongs; Para 0153, the AI program 502 may include a deep neural network (e.g., a front-end network 504 configured to perform pre-processing, such as feature recognition; Para 0154, Additionally… the front-end program 504 can include one or more AI algorithms… For example, a CNN 508 and/or AI algorithm 510 may be used for image recognition].
Austraat is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Austraat and provide a visual analysis on data collected data from video conferencing services, instant messages, and email messages in order to introduce complex analytical tools to inexperienced users for enhancing capabilities in commonplace domains.
Regarding claim 2, Dhrif-Paranjape-Austraat teach the limitations of claim 1.
Dhrif further teaches,
determining the level of the ambiguity of the prompt [Col 10, lines 49-51, The NLU component 260 attempts to make a semantic interpretation of the phrases or statements represented in the text data (and/or other ASR output data) input therein; Col 11, lines 1-12, in addition to the NLU intent and slot data, the NLU component 260 may generate other metadata associated with the request (e.g., with the audio data 211). Examples of such metadata include, an NLU confidence score for the top intent hypothesis, NLU classification type (e.g., statistical vs. deterministic), NLU slot presence (e.g., data indicating that a particular slot was present), NLU confidence score for the overall top hypothesis (e.g., including the relevant speech processing application, intent, and/or slot), entity recognition confidence scores, entity recognition match types (e.g., exact match, prefix match, suffix match, etc.), etc; Col 15, lines 44-54, Accordingly, the decider component 132 may compare the results of the ranking component 120 to one or more predefined policies that may indicate whether or not request data should be sent to top-ranked result of the ranking component 120 or whether some other action should be taken. For example, if the phrase “Arm the security system” is interpreted by ASR/NLU as the current utterance, the decider component may comprise a policy indicating that the ranking component results should be ignored and that the utterance should always be passed to a security system skill used to control security system hardware].
Regarding claim 5, Dhrif-Paranjape-Austraat teach the limitations of claim 1.
Dhrif further teaches,
applying one or more of autoregressive integrated moving average (alternate), part-of-speech tagging (alternate), and named entity recognition to identify the context of the prompt [Col 11, lines 1-12, in addition to the NLU intent and slot data, the NLU component 260 may generate other metadata associated with the request (e.g., with the audio data 211). Examples of such metadata include, an NLU confidence score for the top intent hypothesis, NLU classification type (e.g., statistical vs. deterministic), NLU slot presence (e.g., data indicating that a particular slot was present), NLU confidence score for the overall top hypothesis (e.g., including the relevant speech processing application, intent, and/or slot), entity recognition confidence scores, entity recognition match types (e.g., exact match, prefix match, suffix match, etc.), etc; Col 15, lines 44-54, Accordingly, the decider component 132 may compare the results of the ranking component 120 to one or more predefined policies that may indicate whether or not request data should be sent to top-ranked result of the ranking component 120 or whether some other action should be taken. For example, if the phrase “Arm the security system” is interpreted by ASR/NLU as the current utterance, the decider component may comprise a policy indicating that the ranking component results should be ignored and that the utterance should always be passed to a security system skill used to control security system hardware].
Regarding claim 6, Dhrif-Paranjape-Austraat teach the limitations of claim 1.
Dhrif further teaches,
applying one or more of sentiment analysis (alternate), intent recognition [Col 4, lines 9-11, The text data and/or other ASR output data may be transformed into intent data by an NLU component of the speech processing system], and entity extraction (alternate) to identify the intent of the prompt [Col 16, lines 28-37, it may be unclear, based on a user's request, whether the user intended to invoke a movie playback skill or a music playback skill, as a movie and a soundtrack for the movie may be identified using the same name. Accordingly, the decider component 132 may determine that the top two hypotheses of ranking component 120 are equally likely (or approximately equally likely) and may determine that a question should be asked to disambiguate between the two possible actions].
Regarding claim 7, Dhrif-Paranjape-Austraat teach the limitations of claim 1.
Paranjape further teaches,
applying one or more of topic modelling (alternate), text categorization (Sect 3.3, para 2, extract answer box snippets), and clustering (Sect 3.3, para 2, combine the top-2 search result snippets together) to select the datasets relevant to the context [Sect 3.1, para 2, ART retrieves similar tasks from a task library (Figure 2(A); Section 3.2), and adds instances of those tasks as demonstrations in the prompt. A demonstration in the task library is written in a specific format, defined by a custom parsing expression grammar (PeG) (Section 3.2). The grammar is defined such that each task instance is decomposed into a sequence of sub-steps; Sect 3.3, para 2, We use SerpAPI, which provides an API for Google search. The input to search is the sequence generated by the LLM after “Qi: [search]”. We extract answer box snippets when they are available or combine the top-2 search result snippets together].
Paranjape is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Paranjape and provide efficient methods of selecting datasets relevant to the context in order to gather the pertinent information that is stored in public domains.
Regarding claim 8, Dhrif teaches,
A computer program product for training a context-aware artificial intelligence (AI) chatbot [Col 3, lines 55-57, Natural language generation (NLG) is a computer-based process that may be used to produce natural language output... NLG… may be used together as part of a natural language interface system; Col 4, lines 29-32, The various techniques described herein may be used in a variety of contexts, including in natural language processing enabled devices (e.g., devices employing voice control and/or speech processing “voice assistants”); Col 12, lines 16-22, generating pre-computed features based on user feedback data by ranking and arbitration component 140 may generate increasingly contextually rich feature data that may be used to train various machine learning models used to route speech processing request data (e.g., ranking component 120, embedding-based retrieval component 110, etc.)], the computer program product comprising a computer readable storage medium having program instructions stored therewith, the program instructions executable by one or more processors [Col 18, lines 29-42, The storage element 402 can include one or more different types of memory, data storage, or computer-readable storage media devoted to different purposes within the architecture 400. For example, the storage element 402 may comprise flash memory, random-access memory, disk-based storage, etc. Different portions of the storage element 402, for example, may be used for program instructions for execution by the processing element 404, storage of images or other digital works, and/or a removable storage for transferring data to other devices, etc. In various examples, the storage element 402 may comprise embedding-based retrieval component 110 and/or other components of natural language processing system 200], the program instructions executable to:
Receive, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity [Col 16, lines 28-37, it may be unclear, based on a user's request, whether the user intended to invoke a movie playback skill or a music playback skill, as a movie and a soundtrack for the movie may be identified using the same name. Accordingly, the decider component 132 may determine that the top two hypotheses of ranking component 120 are equally likely (or approximately equally likely) and may determine that a question should be asked to disambiguate between the two possible actions];
identify context of the prompt, by analyzing collected data from communication sources [col 3, lines 66-67 – col 4, lines 1-10, user utterances, input text data, and/or any form of data input to a natural language processing system (“input data”) may be described by “request data” and/or “user request data.” Such request data may change forms many times during processing of the request data by various components of the natural language processing system. For example, initially the request data may be audio data and/or input text data representing a user question… The text data and/or other ASR output data may be transformed into intent data by an NLU component of the speech processing system];
identify an intent of the prompt, by using natural language processing to disambiguate the prompt [Col 16, lines 28-37, it may be unclear, based on a user's request, whether the user intended to invoke a movie playback skill or a music playback skill, as a movie and a soundtrack for the movie may be identified using the same name. Accordingly, the decider component 132 may determine that the top two hypotheses of ranking component 120 are equally likely (or approximately equally likely) and may determine that a question should be asked to disambiguate between the two possible actions];
Dhrif teaches the above limitations of claim 8 including the context (Dhrif, col 3, lines 66-67 – col 4, lines 1-10), the intent (Dhrif, Col 16, lines 28-37), and the context-aware AI chatbot (Dhrif, Col 4, lines 29-32).
Dhrif does not teach collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service; wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image; provide a modified prompt, according to context and intent; select datasets relevant to the context; and generate a response to the prompt, based on the modified prompt and the datasets, wherein chatbot responds to the prompt with the response.
Parajape teaches,
provide a modified prompt, according to context and intent [See Fig 2; Sect 3.1, para 1-2, In Figure 2, ART is presented with a new task description and input instance… ART retrieves similar tasks from a task library (Figure 2(A); Section 3.2), and adds instances of those tasks as demonstrations in the prompt. A demonstration in the task library is written in a specific format, defined by a custom parsing expression grammar (PeG) (Section 3.2). The grammar is defined such that each task instance is decomposed into a sequence of sub-steps];
select datasets relevant to the context [Sect 3.1, para 2, ART retrieves similar tasks from a task library (Figure 2(A); Section 3.2), and adds instances of those tasks as demonstrations in the prompt. A demonstration in the task library is written in a specific format, defined by a custom parsing expression grammar (PeG) (Section 3.2). The grammar is defined such that each task instance is decomposed into a sequence of sub-steps; Sect 3.3, para 2, We use SerpAPI, which provides an API for Google search. The input to search is the sequence generated by the LLM after “Qi: [search]”. We extract answer box snippets when they are available or combine the top-2 search result snippets together]; and
generate a response to the prompt, based on the modified prompt and the datasets, wherein the context-aware chatbot responds to the prompt with the response [Sect 3.3, para 3-4, We use the Codex… model for code generation. Input to code generation is the sequence generated by the LM after the sub-task query symbol “Qi : [generate python code]”. This argument is an instruction for code generation and is prompted to Codex as a multi-line comment in Python. For example, in Figure 2, Codex is prompted the instruction ““Use the formula Fx = Ftens * cosine(θ) to solve...”” as a comment and generates “T = 72.0, theta = 35.0, ..., Fx = T*math.cos(radians)”… We run Python code in a virtual Python environment with arithmetic, symbolic, and scientific computing packages pre-installed. The argument to code execute is the previous sub-task’s answer sequence “#(i - 1) :…”, i.e. the python code snippet to be executed. For i = 1, the task input is used as the argument since it potentially contains the code snippet to be executed. In Figure 2, the code snippet generated in the previous step is executed and the value of variable “Fx” is added to the incomplete program; Sect 3.3, para 2, We use SerpAPI, which provides an API for Google search. The input to search is the sequence generated by the LLM after “Qi: [search]”. We extract answer box snippets when they are available or combine the top-2 search result snippets together. For PQA in Figure 2(B), the search query is the original input followed by “What is the formula for the horizontal component of tension force?”, and the output is ““... horizontal component (Fx) can be calculated as Ftens*cosine(θ) ...””.].
Paranjape is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Paranjape and provide modified prompts, and responses to those prompts, in order to output complex answers to user queries by breaking down intricate tasks into a series of simpler tasks. Additionally, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Paranjape and provide selecting datasets relevant to the context in order to gather the necessary information that is stored in public domains.
Dhriff-Parajape do not teach collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service; wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image.
Austraat teaches,
collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service [Abstract, Systems and methods receive, from a user device through a communication channel, and process, in real-time, a natural language input comprising unstructured data that is derived from an audio signal; Para 0031, The term “transcript data” may be used to refer to a written digital record, in text form, of a single speaker or a written or verbal interaction between multiple participants in a conversation or discussion about various information or content. The transcript data may generally refer to alphanumeric text in digital form. Content can be generated using automatic speech recognition (ASR) and natural language understanding (NLU), which transcribes oral interaction or audio signals during a communication by telephone or video conference and may produce a full-text transcript inclusive of all utterances and disfluencies and/or a summarized version of the interaction or interactions with a focus on only the most semantically or transactionally salient aspects. Alternatively, content may be generated during written exchanges by email, instant messaging, chat messaging, short message service (SMS) text, or other messages exchanged through various online platforms or social media software applications];
wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image [Para 0095, the machine-learning algorithm may include one or more image recognition algorithms suitable to determine one or more categories to which an input, such as data communicated from a visual sensor or a file in JPEG, PNG or other format, representing an image or portion thereof, belongs; Para 0153, the AI program 502 may include a deep neural network (e.g., a front-end network 504 configured to perform pre-processing, such as feature recognition; Para 0154, Additionally… the front-end program 504 can include one or more AI algorithms… For example, a CNN 508 and/or AI algorithm 510 may be used for image recognition].
Austraat is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Austraat and provide a visual analysis on data collected data from video conferencing services, instant messages, and email messages in order to introduce complex analytical tools to inexperienced users for enhancing capabilities in commonplace domains.
Claims 9 and 12-14 are computer program product claims that recite similar limitations to method claims 2 and 5-7, respectively. Therefore, claims 9 and 12-14 are rejected using the same rationales as claims 2 and 5-7, respectively.
Regarding claim 15, Dhrif teaches,
A computer system for training a context-aware artificial intelligence (AI) chatbot [Col 3, lines 55-57, Natural language generation (NLG) is a computer-based process that may be used to produce natural language output... NLG… may be used together as part of a natural language interface system; Col 4, lines 29-32, The various techniques described herein may be used in a variety of contexts, including in natural language processing enabled devices (e.g., devices employing voice control and/or speech processing “voice assistants”); Col 12, lines 16-22, generating pre-computed features based on user feedback data by ranking and arbitration component 140 may generate increasingly contextually rich feature data that may be used to train various machine learning models used to route speech processing request data (e.g., ranking component 120, embedding-based retrieval component 110, etc.)], the computer system comprising one or more processors, one or more computer readable tangible storage devices, and program instructions stored on at least one of the one or more computer readable tangible storage devices for execution by at least one of the one or more processors [Col 18, lines 29-42, The storage element 402 can include one or more different types of memory, data storage, or computer-readable storage media devoted to different purposes within the architecture 400. For example, the storage element 402 may comprise flash memory, random-access memory, disk-based storage, etc. Different portions of the storage element 402, for example, may be used for program instructions for execution by the processing element 404, storage of images or other digital works, and/or a removable storage for transferring data to other devices, etc. In various examples, the storage element 402 may comprise embedding-based retrieval component 110 and/or other components of natural language processing system 200], the program instructions executable to:
Receive, from a user, a prompt directed to the context-aware AI chatbot, wherein the prompt has a level of ambiguity [Col 16, lines 28-37, it may be unclear, based on a user's request, whether the user intended to invoke a movie playback skill or a music playback skill, as a movie and a soundtrack for the movie may be identified using the same name. Accordingly, the decider component 132 may determine that the top two hypotheses of ranking component 120 are equally likely (or approximately equally likely) and may determine that a question should be asked to disambiguate between the two possible actions];
identify context of the prompt, by analyzing collected data from communication sources [col 3, lines 66-67 – col 4, lines 1-10, user utterances, input text data, and/or any form of data input to a natural language processing system (“input data”) may be described by “request data” and/or “user request data.” Such request data may change forms many times during processing of the request data by various components of the natural language processing system. For example, initially the request data may be audio data and/or input text data representing a user question… The text data and/or other ASR output data may be transformed into intent data by an NLU component of the speech processing system];
identify an intent of the prompt, by using natural language processing to disambiguate the prompt [Col 16, lines 28-37, it may be unclear, based on a user's request, whether the user intended to invoke a movie playback skill or a music playback skill, as a movie and a soundtrack for the movie may be identified using the same name. Accordingly, the decider component 132 may determine that the top two hypotheses of ranking component 120 are equally likely (or approximately equally likely) and may determine that a question should be asked to disambiguate between the two possible actions];
Dhrif teaches the above limitations of claim 15 including the context (Dhrif, col 3, lines 66-67 – col 4, lines 1-10), the intent (Dhrif, Col 16, lines 28-37), and the context-aware AI chatbot (Dhrif, Col 4, lines 29-32).
Dhrif does not teach collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service; wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image; provide a modified prompt, according to context and intent; select datasets relevant to the context; and generate a response to the prompt, based on the modified prompt and the datasets, wherein chatbot responds to the prompt with the response.
Parajape teaches,
provide a modified prompt, according to context and intent [See Fig 2; Sect 3.1, para 1-2, In Figure 2, ART is presented with a new task description and input instance… ART retrieves similar tasks from a task library (Figure 2(A); Section 3.2), and adds instances of those tasks as demonstrations in the prompt. A demonstration in the task library is written in a specific format, defined by a custom parsing expression grammar (PeG) (Section 3.2). The grammar is defined such that each task instance is decomposed into a sequence of sub-steps];
select datasets relevant to the context [Sect 3.1, para 2, ART retrieves similar tasks from a task library (Figure 2(A); Section 3.2), and adds instances of those tasks as demonstrations in the prompt. A demonstration in the task library is written in a specific format, defined by a custom parsing expression grammar (PeG) (Section 3.2). The grammar is defined such that each task instance is decomposed into a sequence of sub-steps; Sect 3.3, para 2, We use SerpAPI, which provides an API for Google search. The input to search is the sequence generated by the LLM after “Qi: [search]”. We extract answer box snippets when they are available or combine the top-2 search result snippets together]; and
generate a response to the prompt, based on the modified prompt and the datasets, wherein the context-aware chatbot responds to the prompt with the response [Sect 3.3, para 3-4, We use the Codex… model for code generation. Input to code generation is the sequence generated by the LM after the sub-task query symbol “Qi : [generate python code]”. This argument is an instruction for code generation and is prompted to Codex as a multi-line comment in Python. For example, in Figure 2, Codex is prompted the instruction ““Use the formula Fx = Ftens * cosine(θ) to solve...”” as a comment and generates “T = 72.0, theta = 35.0, ..., Fx = T*math.cos(radians)”… We run Python code in a virtual Python environment with arithmetic, symbolic, and scientific computing packages pre-installed. The argument to code execute is the previous sub-task’s answer sequence “#(i - 1) :…”, i.e. the python code snippet to be executed. For i = 1, the task input is used as the argument since it potentially contains the code snippet to be executed. In Figure 2, the code snippet generated in the previous step is executed and the value of variable “Fx” is added to the incomplete program; Sect 3.3, para 2, We use SerpAPI, which provides an API for Google search. The input to search is the sequence generated by the LLM after “Qi: [search]”. We extract answer box snippets when they are available or combine the top-2 search result snippets together. For PQA in Figure 2(B), the search query is the original input followed by “What is the formula for the horizontal component of tension force?”, and the output is ““... horizontal component (Fx) can be calculated as Ftens*cosine(θ) ...””.].
Paranjape is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Paranjape and provide modified prompts, and responses to those prompts, in order to output complex answers to user queries by breaking down intricate tasks into a series of simpler tasks. Additionally, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Paranjape and provide selecting datasets relevant to the context in order to gather the necessary information that is stored in public domains.
Dhriff-Parajape do not teach collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service; wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image.
Austraat teaches,
collect, in real time, data from a plurality of communication sources associated with the user, the plurality of communication sources comprising a video conferencing service, an instant messaging service, and an email service [Abstract, Systems and methods receive, from a user device through a communication channel, and process, in real-time, a natural language input comprising unstructured data that is derived from an audio signal; Para 0031, The term “transcript data” may be used to refer to a written digital record, in text form, of a single speaker or a written or verbal interaction between multiple participants in a conversation or discussion about various information or content. The transcript data may generally refer to alphanumeric text in digital form. Content can be generated using automatic speech recognition (ASR) and natural language understanding (NLU), which transcribes oral interaction or audio signals during a communication by telephone or video conference and may produce a full-text transcript inclusive of all utterances and disfluencies and/or a summarized version of the interaction or interactions with a focus on only the most semantically or transactionally salient aspects. Alternatively, content may be generated during written exchanges by email, instant messaging, chat messaging, short message service (SMS) text, or other messages exchanged through various online platforms or social media software applications];
wherein analyzing the collected data includes performing a visual analysis of an image associated with the plurality of communication sources by applying a convolutional neural network (CNN) to analyze the image and identify one or more features in the image [Para 0095, the machine-learning algorithm may include one or more image recognition algorithms suitable to determine one or more categories to which an input, such as data communicated from a visual sensor or a file in JPEG, PNG or other format, representing an image or portion thereof, belongs; Para 0153, the AI program 502 may include a deep neural network (e.g., a front-end network 504 configured to perform pre-processing, such as feature recognition; Para 0154, Additionally… the front-end program 504 can include one or more AI algorithms… For example, a CNN 508 and/or AI algorithm 510 may be used for image recognition].
Austraat is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Austraat and provide a visual analysis on data collected data from video conferencing services, instant messages, and email messages in order to introduce complex analytical tools to inexperienced users for enhancing capabilities in commonplace domains.
Claims 16 and 18-20 are system claims that recite similar limitations to method claims 2 and 5-7, respectively. Therefore, claims 16 and 18-20 are rejected using the same rationales as claims 2 and 5-7, respectively.
Claim(s) 3, 4, 10, 11, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Dhrif in view of Paranjape and Austraat, and in further view of Vamvourellis et al. (US 12517984 B1), hereinafter V.
Regarding claim 3, Dhrif-Paranjape-Austraat teach the limitations of claim 1.
Dhrif-Paranjape-Austraat do not teach pre-processing the collected data to provide pre-processed data with correct formats and free from irrelevant words or phrases; and wherein the various sources of the collected data include video conferencing services, instant messages, and email messages.
V teaches,
pre-processing the collected data to provide pre-processed data with correct formats and free from irrelevant words or phrases [Col 15, lines 21-27, unstructured text data is pre-processed. For example, pre-processing layer 204 may receive text data that is received in process 302 and/or filtered in processes 306. Pre-processing layer 204 may further divide text data into tokens, filter stop-word and remove frequent words such as articles, prepositions, and the like, converting text to lowercase, and n-gram and lemmatize text].
V is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of V and provide pre-processed data with correct formats and free from irrelevant words or phrases in order to improve efficiency by focusing on meaningful content.
Regarding claim 4, Dhrif-Paranjape-Austraat-V teach the limitations of claim 3.
V further teaches,
applying one or more of tokenization, stemming, lemmatization, and stop word removal to pre-process the collected data Col 15, lines 21-27, unstructured text data is pre-processed. For example, pre-processing layer 204 may receive text data that is received in process 302 and/or filtered in processes 306. Pre-processing layer 204 may further divide text data into tokens, filter stop-word and remove frequent words such as articles, prepositions, and the like, converting text to lowercase, and n-gram and lemmatize text].
V is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of V and provide stop word removal to pre-process the collected data in order to improve efficiency by focusing on meaningful content.
Claims 10 and 11 are computer program product claims that recite similar limitations to method claims 3 and 4 respectively. Therefore, claims 10 and 11 are rejected using the same rationales as claims 3 and 4 respectively.
Regarding claim 17, Dhrif-Paranjape-Austraat teach the limitations of claim 15.
Dhrif-Paranjape do not teach pre-process the collected data to provide pre-processed data with correct formats and free from irrelevant words or phrases; wherein the various sources of the collected data include video conferencing services, instant messages, and email messages; and wherein one or more of tokenization, stemming, lemmatization, and stop word removal are applied for pre-processing the collected data.
V teaches,
pre-process the collected data to provide pre-processed data with correct formats and free from irrelevant words or phrases; and wherein one or more of tokenization, stemming, lemmatization, and stop word removal are applied for pre-processing the collected data [Col 15, lines 21-27, unstructured text data is pre-processed. For example, pre-processing layer 204 may receive text data that is received in process 302 and/or filtered in processes 306. Pre-processing layer 204 may further divide text data into tokens, filter stop-word and remove frequent words such as articles, prepositions, and the like, converting text to lowercase, and n-gram and lemmatize text].
V is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of V and provide pre-processed data with correct formats and free from irrelevant words or phrases in order to improve efficiency by focusing on meaningful content.
Dhrif-Paranjape-Austraat-V do not teach wherein various sources of collected data include video conferencing services, instant messages, and email messages.
Austraat teaches,
wherein various sources of collected data include video conferencing services, instant messages, and email messages [Para 0031, The term “transcript data” may be used to refer to a written digital record, in text form, of a single speaker or a written or verbal interaction between multiple participants in a conversation or discussion about various information or content. The transcript data may generally refer to alphanumeric text in digital form. Content can be generated using automatic speech recognition (ASR) and natural language understanding (NLU), which transcribes oral interaction or audio signals during a communication by telephone or video conference and may produce a full-text transcript inclusive of all utterances and disfluencies and/or a summarized version of the interaction or interactions with a focus on only the most semantically or transactionally salient aspects. Alternatively, content may be generated during written exchanges by email, instant messaging, chat messaging, short message service (SMS) text, or other messages exchanged through various online platforms or social media software applications].
Austraat is analogous to the claimed invention as they both relate to natural language processing. Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Dhrif’s teachings to incorporate the teachings of Austraat and provide collected data from video conferencing services, instant messages, and email messages.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SYED RAYHAN AHMED whose telephone number is (571)270-0286. The examiner can normally be reached Mon-Fri ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SYED RAYHAN AHMED/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126