DETAILED ACTION
Claims 1-20 are pending. Claims 1, 9, and 17 are independent. Claims are not amended.
This Application was published as U.S. 20240354515.
Apparent priority: 24 April 2023.
Applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 4/10/2026 has been entered.
Response to Amendment and Arguments.
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Applicant has argued regarding Jothilingam (Raghu) that Jothilingam does not teach the LLM that is claimed because the reference uses a “traditional machine learning classification” which employs “a (e.g., statistical) model to classify components:
PNG
media_image1.png
150
612
media_image1.png
Greyscale
PNG
media_image2.png
112
602
media_image2.png
Greyscale
Response 8-9.
In Reply, please note the Specification of the instant Application which describes its LLM as “statistical models”:
“[0008] Various machine learning (ML) and artificial intelligence (AI) models are capable of analyzing and/or generating text. One example of such a model is a large language model (LLM). An LLM is a statistical model that predicts the next word in a sequence, given the previous words (often referred to as a “prompt”). …”
However, Jothilingam was not cited for teaching an LLM and a different reference was and is cited for this feature.
Additionally, please note that the independent Claims include a reference to a “description of the potential action”: “generating, by the large language model, a description of a potential action identified by the large language model.” The phrase “description of potential action” is not shown in the drawings nor described with any detail in the Specification. The phrase may be supported by “notification 306,” Figure 3 of the instant Application below:
PNG
media_image3.png
692
1004
media_image3.png
Greyscale
Note the similarity of the above drawing to the following Figure 3 of Jothilingam:
PNG
media_image4.png
370
528
media_image4.png
Greyscale
It is suggested that if the “notification 306” of Figure 3 of the instant Application is intended as the definition of “description of the potential action,” then the phrase “notification” which is better supported is used in the Claim. Additionally, in order to distinguish the phrase from the reference, the claim language needs to define the phrase with adequate particularity.
See also the following Figure of the instant Application showing the order of operations:
PNG
media_image5.png
758
396
media_image5.png
Greyscale
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1,2, 3, 4, 5 8, 9, 10, 11, 12, 13, 16, 17, 18, 19, and 20 are rejected under 35 U.S.C. 103 as obvious over Jothilingam (US 20170193349) in view of Dan (US 20240305589) in further view of Lang (US 20230259714)
Claim 1,9,17
Jothilingam:
PNG
media_image6.png
300
454
media_image6.png
Greyscale
PNG
media_image4.png
370
528
media_image4.png
Greyscale
PNG
media_image7.png
780
440
media_image7.png
Greyscale
Regarding Claim 1,9,17, Jothilingam teaches
1. A method comprising:
identifying, by a processor, a plurality of electronic messages addressed to a user;
Figure 1: “Processors 104.”
Figure 2: “Electronic Communication 202” teaches the “plurality of electronic messages” of the Claim.
“[0040] FIG. 2 is a block diagram illustrating electronic communication 202 subjected to an example task extraction process 204. For example, process 204 may involve any of a number of techniques for detecting whether task content is included in incoming or outgoing communications or in a database. Process 204 may also involve techniques for automatically marking, annotating, or otherwise identifying the message as containing task content. In some examples, process 204 may include techniques that extract a summary (not illustrated) of tasks for presentation and follow-up tracking and analysis. Task 206 may be extracted from multiple forms of content of electronic communication 202. Such content may include interpersonal communications such as email, SMS text or images, instant messaging, posts in social media, meeting notes, database content, and so on. Such content may also include content composed using email applications or word-processing applications, among other possibilities.”
extracting, from the plurality of electronic messages, at least one of an action, a subject, and a keyword associated with a potential action:
(Fig 2 shows extracting from a communication a task (208), further a task parameter extraction for a action, subject, and keyword
Fig 3 shows an example of extracting a task, action, subject, keyword.
Paragraph 17 "… For example, the system may examine other messages exchanged by one or both of the authors of the email exchange or by other people. The system may also examine larger corpora of email and other messages….")
analyzing, by a large language model executed by the processor, the plurality of electronic messages and the extracted at least one of the action, subject, and keyword to identify potential actions;
Figure 1, the “extraction module 116” conducts the “perform task extraction process 1004” of Figure 10 which teaches the “analyzing …” of this limitation because it “identify potential actions” as shown in Figure 10: “generate task-oriented processes 1006.”
“[0032] … Extraction module 116 may selectively apply any of a number of statistical models or predictive models (e.g., via machine learning module 114) stored in computer-readable media 108 to apply to input data.”
“[0078] At block 1006, task operations module 402 may generate one or more task-oriented actions based, at least in part, on the determined task content. ….”
generating, by the large language model, a description of a potential action identified by the large language model;
Figure 10, block 1006 generates a list of identified processes or tasks.
“[0078] At block 1006, task operations module 402 may generate one or more task-oriented actions based, at least in part, on the determined task content….”
Generating a list at least suggests the “description of a potential action” depending on how the action is named. The name of the action may be descriptive and therefore would provide “a description of a potential action.”
Further, Figure 3 of the reference, as provided above, shows that at 308 the result of the “Extraction” is: “Task: Write a presentation for the meeting on May 9th.”
“[0049] … In the example illustrated in FIG. 3, text 306 by the other user includes a task 308 that the user writes a presentation for a meeting on May 9.sup.th. ….”
(Figure 3 of the instant Application which shows the “Notification 306” is very similar to Figure 3 of Jothilingam.)
suggesting, by the processor, to the user, the generated description of a potential action identified by the large language model by displaying the description in a user interface;
Figure 10: “provide task-oriented process list to user 1008.”
“[0079] At block 1008, task operations module 402 may provide a list of the task-oriented actions to the user for inspection or review….”
“[0076] FIG. 9 is a view of a display 900 showing an example task list 902, which may include a prioritization field 904 of a list of tasks 906. A system may use a prioritization engine, such as 510, to prioritize the tasks by using task parameters (e.g., 310, illustrated in FIG. 3) and results of machine learning, as described above….’
receiving, by the processor, user input related to the potential action via a chat interface, wherein the user input comprises a modification of the potential action;
Figure 10, “Modify task-oriented processes 1010.”
“[0079] … At diamond 1010, the user may select among choices of different possible actions to be performed by task operations module 402, refine possible actions, delete actions, manually add actions, and so on….”
generating, using the large language model, a new potential action based on the modification; and
Figure 10: when the response out of the decision diamond 101 is YES it means that the tasks have been modified by the user and the process goes back to 1004.
“[0079] … If there are any such changes, then process 1000 may return to block 1004 where task operations module 402 may re-generate task-oriented processes in view of the user's edits of the task-oriented process list.. ….”
performing, by the processor, on behalf of the user, a subsequent action conforming to the user input wherein performing the subsequent action comprises performing the new potential action via a large language model agent that interoperates with the large language model.
Figure 10, 1012 and 1014.
“[0079] … .On the other hand, if the user approves the list, then process 1000 may proceed to block 1012 where task operations module 402 performs the task-oriented processes. At block 1014, the task operations module may generate and display a visual cue and productivity report, for example.”
(Paragraph 79 "At block 1008, task operations module 402 may provide a list of the task-oriented actions to the user for inspection or review. For example, a task-oriented action may be to find or locate digital artefacts (e.g., documents) related to a particular task to support completion of, or user comprehension of, a task activity. At diamond 1010, the user may select among choices of different possible actions to be performed by task operations module 402, refine possible actions, delete actions, manually add actions, and so on. If there are any such changes, then process 1000 may return to block 1004 where task operations module 402 may re-generate task-oriented processes in view of the user's edits of the task-oriented process list. On the other hand, if the user approves the list, then process 1000 may proceed to block 1012 where task operations module 402 performs the task-oriented processes. At block 1014, the task operations module may generate and display a visual cue and productivity report, for example.")
Jothilingam does not explicitly teach all of (the bolded )
analyzing, by a large language model executed by the processor, the plurality of electronic messages and the extracted at least one of the action, subject, and keyword to identify potential actions;
generating, by the large language model, a description of a potential action identified by the large language model;
suggesting, by the processor, to the user, the generated description of a potential action identified by the large language model by displaying the description in a user interface;
generating, using the large language model, a new potential action based on the modification; and
performing, by the processor, on behalf of the user, a subsequent action conforming to the user input wherein performing the subsequent action comprises performing the new potential action via a large language model agent that interoperates with the large language model.
Further, Jothlingam teaches the entire framework of the Claim.
Although the machine learning model that Jothlingam uses can be an LLM and any of the actions identified could be performed by an LLM agent.
Although Jothlingam teaches generating a “description” of the action in “generating, by the large language model, a description of a potential action” it does not expressly teach “suggesting … the generated description … in a user interface.” However, as provided above, the name of the action could be descriptive and provide a description of the task. For example, Email is descriptive as the name and the description of the task of emailing.
However, DAN. teach
analyzing, by a large language model executed by the processor, the plurality of electronic messages and the extracted at least one of the action, subject, and keyword to identify potential actions;
(paragraph 3 "Relatively recently, generative models have been developed, where generative models include generative language models (GLMs) (also referred to as large language models (LMS)), models that generate images based upon input (where the input may be text, voice, an image, etc.), models that generate video based upon input, and so forth. An example of a GLM is the Generative Pre-trained Transformer 4 (GPT-4) model. Another example of a GLM is the BigScience Large Open-science Open-access Multilingual language (BLOOM) model, which is also a transformer-based model. Briefly, a generative model is configured to generate an output (such as text in human-readable language, source code, music, video, and the like) based upon a prompt provided as input to the generative model, where the generative model generates the output in near real-time (e.g., within a few seconds of receiving the prompt)."
paragraph 35 "The prompt generator module 304 then includes a transcript of messages in the group conversation in the prompt, where the transcripts include content of the messages, timestamps that indicate when the messages were obtained by the messaging application 114, and identities of users who submitted the messages. The prompt generator module 304 also identifies the user who requested invocation of the bot 116 in the prompt 400, and optionally includes a location of the user in the prompt 400. The prompt generator module 304 can also include instructions in the prompt 400 that are to be provided to the bot 116 to restrict output of the bot 116. For instance, the instructions can indicate that the user who is requesting invocation of the bot 116 is interested in answers from the provided context of the group conversation (e.g., from the transcription of the group conversation). Further, the prompt 400 can include a request for the bot 116 to not extrapolate, to not provide personal information about people who are participating in the group conversation, to not make inferences about users participating in the group conversation, and so forth. These instructions can facilitate reduction in hallucinations output by the bot 116, as the bot 116 is limited to responding through use of information within the context."
Paragraph 37 "Optionally, the prompt generator model 304 can include other information about users in the group conversation in the prompt 400, such as information retrieved from profiles of the users, including topics of interest to the users, keywords and facts mentioned by the users in other conversations, demographic information of the users (such as age) to allow the bot 116 to tailor responses, and so forth."
Paragraph 51 "The bot 116 can assist users in the group conversations with various tasks. Examples of tasks that the bot 116 can assist with include, but are not limited to: 1) summarizing the current conversation; 2) answering a question about what a person said about a topic in the conversation; 3) proposing follow up action items based upon information in the group conversation; 4) helping the group plan events, such as vacations; 5) generating images, avatars, memes, etc. pertaining to messages included in the group conversation; 6) helping to rewrite and check text that the group is working on-making text more professional, making text more humorous, fixing typos in the text, translating text to different languages, and so forth; 7) answering questions about the contents of a webpage pointed to by a URL that has been placed in the group conversation, such as summarizing news articles, identifying facts mentioned in a news article, reading structured data such as company earnings documents, answering specific questions about content of the web page, etc.; 8) brainstorming ideas, such as producing road maps for product development; 9) translating text from one language to another, such that languages of participants in the group does not matter, as the bot 116 can translate and convey information between participants in languages requested by the participants; and 10) starting and running text and image-based games to entertain users in the group conversation.")
generating, by the large language model, a description of a potential action identified by the large language model;
(paragraph 3 provisional paragraph 2 "Relatively recently, generative models have been developed, where generative models include generative language models (GLMs) (also referred to as large language models (LMS)), models that generate images based upon input (where the input may be text, voice, an image, etc.), models that generate video based upon input, and so forth. An example of a GLM is the Generative Pre-trained Transformer 4 (GPT-4) model. Another example of a GLM is the BigScience Large Open-science Open-access Multilingual language (BLOOM) model, which is also a transformer-based model. Briefly, a generative model is configured to generate an output (such as text in human-readable language, source code, music, video, and the like) based upon a prompt provided as input to the generative model, where the generative model generates the output in near real-time (e.g., within a few seconds of receiving the prompt)."
Paragraph 37 provisional paragraph 36 "Optionally, the prompt generator model 304 can include other information about users in the group conversation in the prompt 400, such as information retrieved from profiles of the users, including topics of interest to the users, keywords and facts mentioned by the users in other conversations, demographic information of the users (such as age) to allow the bot 116 to tailor responses, and so forth."
Paragraph 51 provisional paragraph 50 "The bot 116 can assist users in the group conversations with various tasks. Examples of tasks that the bot 116 can assist with include, but are not limited to: 1) summarizing the current conversation; 2) answering a question about what a person said about a topic in the conversation; 3) proposing follow up action items based upon information in the group conversation; 4) helping the group plan events, such as vacations; 5) generating images, avatars, memes, etc. pertaining to messages included in the group conversation; 6) helping to rewrite and check text that the group is working on-making text more professional, making text more humorous, fixing typos in the text, translating text to different languages, and so forth; 7) answering questions about the contents of a webpage pointed to by a URL that has been placed in the group conversation, such as summarizing news articles, identifying facts mentioned in a news article, reading structured data such as company earnings documents, answering specific questions about content of the web page, etc.; 8) brainstorming ideas, such as producing road maps for product development; 9) translating text from one language to another, such that languages of participants in the group does not matter, as the bot 116 can translate and convey information between participants in languages requested by the participants; and 10) starting and running text and image-based games to entertain users in the group conversation.")
suggesting, by the processor, to the user, the generated description of a potential action identified by the large language model by displaying the description in a user interface;
(paragraph 25 "The memory 112 further includes a bot 116 that comprises or has access to a generative model 118. In an example, the generative model 118 is a transformer-based model. The generative model 118 can be a GLM, such that the generative model 118 is configured to receive text input and generate text output. In such case, the bot 116 can be a chatbot. In other examples, the generative model 118 can generate audio, images, video (with audio) or other suitable outputs based upon a variety of inputs, such as voice, text, video, images, audio, etc. While the bot 116 is illustrated as being external to the messaging application 114, it is contemplated that the bot 116 may be included in the messaging application 114. As will be described in greater detail herein, the messaging application 114 receives messages from client computing devices 104-106 that are part of a group conversation between users of the client computing devices 104-106. The bot 116, through use of the generative model 118, generates output based upon the received messages and the output is provided to the client computing devices 104-106 as part of the group conversation. Hence, the bot 116 can participate in the group conversation."
Paragraph 50-51 "As expressed above, the technologies described herein allow for use cases that were heretofore not possible in connection with messaging applications that allow for group conversations. For instance, the bot 116 can answer questions related to the context of the group conversation, as well as general questions. In some implementations, the bot 116 generates output only when invoked by a user who is participating in the group conversation. In another example, the bot 116 is provided with input for each turn of a group conversation, and the bot 116 decides when it would be useful to inject itself into the group conversation by generating an output. In still yet another example, the messaging application 114 is associated with a UX canvas that is displayed on the displays of the client computing devices 104-106, where the UX canvas is continuously updated with outputs of the bot 116 (where the outputs are suggestions for inclusion in the group conversation). When a user taps on a suggestion, such suggestion is entered into the group conversation as a turn (and identified as being generated by the bot 116).
[0051] The bot 116 can assist users in the group conversations with various tasks. Examples of tasks that the bot 116 can assist with include, but are not limited to: 1) summarizing the current conversation; 2) answering a question about what a person said about a topic in the conversation; 3) proposing follow up action items based upon information in the group conversation; 4) helping the group plan events, such as vacations; 5) generating images, avatars, memes, etc. pertaining to messages included in the group conversation; 6) helping to rewrite and check text that the group is working on-making text more professional, making text more humorous, fixing typos in the text, translating text to different languages, and so forth; 7) answering questions about the contents of a webpage pointed to by a URL that has been placed in the group conversation, such as summarizing news articles, identifying facts mentioned in a news article, reading structured data such as company earnings documents, answering specific questions about content of the web page, etc.; 8) brainstorming ideas, such as producing road maps for product development; 9) translating text from one language to another, such that languages of participants in the group does not matter, as the bot 116 can translate and convey information between participants in languages requested by the participants; and 10) starting and running text and image-based games to entertain users in the group conversation.")
generating, using the large language model, a new potential action based on the modification: and
(paragraph 25 provisional paragraph 24 "The memory 112 further includes a bot 116 that comprises or has access to a generative model 118. In an example, the generative model 118 is a transformer-based model. The generative model 118 can be a GLM, such that the generative model 118 is configured to receive text input and generate text output. In such case, the bot 116 can be a chatbot. In other examples, the generative model 118 can generate audio, images, video (with audio) or other suitable outputs based upon a variety of inputs, such as voice, text, video, images, audio, etc. While the bot 116 is illustrated as being external to the messaging application 114, it is contemplated that the bot 116 may be included in the messaging application 114. As will be described in greater detail herein, the messaging application 114 receives messages from client computing devices 104-106 that are part of a group conversation between users of the client computing devices 104-106. The bot 116, through use of the generative model 118, generates output based upon the received messages and the output is provided to the client computing devices 104-106 as part of the group conversation. Hence, the bot 116 can participate in the group conversation.")
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jothilingam to incorporate the teachings of DAN to provide a “analyzing, by a large language model executed by the processor, the plurality of electronic messages and the extracted at least one of the action, subject, and keyword to identify potential actions; suggesting, by the processor, to the user, the generated description of a potential action identified by the large language model by displaying the description in a user interface; ” Doing so would Help the model be less confused about the participants, as recognized by DAN. (paragraph 34).
However, Jothilingamin view of DAN. do not explicitly teach all of (the bolded)
performing, by the processor, on behalf of the user, a subsequent action conforming to the user input wherein performing the subsequent action comprises performing the new potential action via a large language model large language model agent that interoperates with the large language model.
However, Lange:
PNG
media_image8.png
720
554
media_image8.png
Greyscale
PNG
media_image9.png
380
462
media_image9.png
Greyscale
However, Lange teaches:
analyzing, by a large language model executed by the processor, the plurality of electronic messages and the extracted at least one of the action, subject, and keyword to identify potential actions;
Figure 3, 320:
“[0088] The GAIN system processes the user input through a language model trained to receive the user input and generate one or more function calls, according to block 320. The one or more function calls can include a first function call which, when executed by the one or more processors, causes the one or more processors to perform one or more predetermined actions associated with a node of a conversation graph specified in the function call. The one or more function calls can be API calls.”
performing, by the processor, on behalf of the user, a subsequent action conforming to the user input wherein performing the subsequent action comprises performing the new potential action via a large language model agent that interoperates with the large language model.
Figure 3, 330:
“[0093] In response to the user input, the GAIN system performs one or more predetermined actions associated with the node in the conversation graph, according to block 330. As described herein with reference to FIGS. 1-2B, the one or more predetermined actions can include one or more of sending a prompt to a user computing device for more information for responding to the user input; providing information responsive to the user input; updating one or more parameter values with information provided from the user input, the one or more parameters saved in one or more memory devices by the one or more processors; and updating the current node in the conversation graph to a different node and performing one or more predetermined actions associated with the different node.”
(paragraph 49 "FIG. 1 is a block diagram of a Graph AI Navigator (GAIN) system 100. The GAIN system 100 includes a user frontend 110 and an agent 105. The agent 105 includes a state handler 115 implementing an Application Program Interface (API) 117 and a language model 120. The GAIN system can be implemented on one or more processors, for example on one or more computing devices of a computing platform 101, described herein with reference to FIG. 9."
Paragraph 52 "The LM 120 can be any of a variety of different machine learning or statistical models, such as deep neural networks, recurrent neural networks, transformers, etc. As described herein, the LM is trained to receive user input and the state of a conversation represented in a conversation graph and generate one or more API calls for causing the agent to perform one or more actions in response to the user input. The LM can be a large language model, for example initially trained and then fine-tuned or retrained with session log data, as described herein with reference to FIG. 4. Although examples are provided herein for training a LM, it is understood that any of a variety of different machine learning or statistical models that can be trained as described herein, may be used, not limited to language models or large language models."
Paragraph 54 "The state handler 115 can refer to one or more components of the system 100 for communicating with a user frontend and the LM 120. The state handler 115 also stores the state of the conversation between the user and the conversation agent. Note that the state handler is an intermediary, preventing the output of the LM from reaching the user frontend directly, and vice versa. The state handler can be at least partially implemented as one or more processors, configured to execute software for receiving and acting on API calls received from the conversational agent. The state handler can define an API of potential operations that can be performed, invoked through one or more function calls.")
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jothilingamin view of DAN to incorporate the teachings of Lange to provide a “large language model agent that interoperates with the large language model.” Doing so would Make the agent more accurate and efficient to execute tasks, as recognized by Lange. (Paragraph 33).
Claim 9
Regarding Claim 9, Jothilingam in view of DAN, in further view of Lange, further more Jothilingam teaches A non-transitory computer-readable storage medium for tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions defining steps of:
(Paragraph 31 “In some examples, as shown regarding device 102d, computer-readable media 108 can store instructions executable by the processor(s) 104 including an operating system (OS) 112, a machine learning module 114, an extraction module 116, a task operations module 118, a graphics generator 120, and programs or applications 122 that are loadable and executable by processor(s) 104. The one or more processors 104 may include one or more central processing units (CPUs), graphics processing units (GPUs), video buffer processors, and so on. In some implementations, machine learning module 114 comprises executable code stored in computer-readable media 108 and is executable by processor(s) 104 to collect information, locally or remotely by computing device 102, via input/output 106. The information may be associated with one or more of applications 122. Machine learning module 114 may selectively apply any of a number of machine learning decision models stored in computer-readable media 108 (or, more particularly, stored in machine learning module 114) to apply to input data.”)
Similar analysis for the remaining claim limitations is shown in claim 1.
Claim 17
Regarding Claim 17, Jothilingam in view of DAN, in further view of Lange, further more Jothilingam teaches a storage medium for tangibly storing thereon logic for execution by the processor, the logic comprising instructions for:
(Paragraph 37 "In contrast, communication media embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, or other transmission mechanism. As defined herein, computer storage media does not include communication media. In various examples, memory 108 is an example of computer storage media storing computer-executable instructions. When executed by processor(s) 104, the computer-executable instructions configure the processor(s) to, among other things, receive a task; extract at least one of an action, a subject, and a keyword from the task; search a history of execution of tasks (e.g., task types) that are similar to the task in a database; and categorize the task based, at least in part, on the history of execution of the similar tasks.")
Similar analysis for the remaining claim limitations is shown in claim 1.
Claim 2,10, 18
Regarding Claim 2,10,18, Jothilingam in view of DAN, in further view of Lange, further more Jothilingam teaches the method of claim 1, wherein performing the subsequent action comprises performing the subsequent action.
(Paragraph 79 "At block 1008, task operations module 402 may provide a list of the task-oriented actions to the user for inspection or review. For example, a task-oriented action may be to find or locate digital artefacts (e.g., documents) related to a particular task to support completion of, or user comprehension of, a task activity. At diamond 1010, the user may select among choices of different possible actions to be performed by task operations module 402, refine possible actions, delete actions, manually add actions, and so on. If there are any such changes, then process 1000 may return to block 1004 where task operations module 402 may re-generate task-oriented processes in view of the user's edits of the task-oriented process list. On the other hand, if the user approves the list, then process 1000 may proceed to block 1012 where task operations module 402 performs the task-oriented processes. At block 1014, the task operations module may generate and display a visual cue and productivity report, for example.")
Jothilingam do not explicitly teach all of via a large language model.
However, Jothilingam in view of DAN, in further view of Lange, further more DAN. teach the method of claim 1, wherein performing the subsequent action comprises performing the subsequent action via a large language model.
(paragraph 50 "As expressed above, the technologies described herein allow for use cases that were heretofore not possible in connection with messaging applications that allow for group conversations. For instance, the bot 116 can answer questions related to the context of the group conversation, as well as general questions. In some implementations, the bot 116 generates output only when invoked by a user who is participating in the group conversation. In another example, the bot 116 is provided with input for each turn of a group conversation, and the bot 116 decides when it would be useful to inject itself into the group conversation by generating an output. In still yet another example, the messaging application 114 is associated with a UX canvas that is displayed on the displays of the client computing devices 104-106, where the UX canvas is continuously updated with outputs of the bot 116 (where the outputs are suggestions for inclusion in the group conversation). When a user taps on a suggestion, such suggestion is entered into the group conversation as a turn (and identified as being generated by the bot 116)."
Paragraph 52 "Further, the bot 116 can generate text in the output, such as answers to questions in plain text, following the flow of a conversation, and understanding pasted structured text such as tables. In another example, the bot 116 accepts audio as input, thereby allowing users to provide audio notes in the group conversation and the bot 116 can understand speech of the user using speech to text technologies. The bot 116 can reply to the input text in either text or using text to speech. The bot 116, through receipt of audio, can listen to and understand an audio/video call, can take a URL pasted into a group conversation as input, such that the bot can follow the URL, download HTML, and answer questions about contents of the page. In another example, the bot 116 accesses links to documents in shared storage spaces. With respect to output, the bot 116 can generate output in any of the modes referenced above (text, audio, images, and so forth). The bot 116 can generate output in images that it has generated, and can fill in meme templates, generate graphs, and so forth. The bot 116 can generate outputs in a tone selected by users in the group chat, and outputs of the bot 116 can be shared to other applications.")
See claim 1 for rationale.
Jothilingam in view of DAN, do not explicitly teach all of large language model agent that interoperates with the large language model.
However, Lange Teach large language model agent that interoperates with the large language model.
(paragraph 49 "FIG. 1 is a block diagram of a Graph AI Navigator (GAIN) system 100. The GAIN system 100 includes a user frontend 110 and an agent 105. The agent 105 includes a state handler 115 implementing an Application Program Interface (API) 117 and a language model 120. The GAIN system can be implemented on one or more processors, for example on one or more computing devices of a computing platform 101, described herein with reference to FIG. 9."
Paragraph 52 "The LM 120 can be any of a variety of different machine learning or statistical models, such as deep neural networks, recurrent neural networks, transformers, etc. As described herein, the LM is trained to receive user input and the state of a conversation represented in a conversation graph and generate one or more API calls for causing the agent to perform one or more actions in response to the user input. The LM can be a large language model, for example initially trained and then fine-tuned or retrained with session log data, as described herein with reference to FIG. 4. Although examples are provided herein for training a LM, it is understood that any of a variety of different machine learning or statistical models that can be trained as described herein, may be used, not limited to language models or large language models."
Paragraph 54 "The state handler 115 can refer to one or more components of the system 100 for communicating with a user frontend and the LM 120. The state handler 115 also stores the state of the conversation between the user and the conversation agent. Note that the state handler is an intermediary, preventing the output of the LM from reaching the user frontend directly, and vice versa. The state handler can be at least partially implemented as one or more processors, configured to execute software for receiving and acting on API calls received from the conversational agent. The state handler can define an API of potential operations that can be performed, invoked through one or more function calls.")
See claim one for rationale.
Claim 3,11,19
Regarding Claim 3,11,19, Jothilingam in view of DAN, in further view of Lange, further more Jothilingam teaches the method of claim 1, wherein identifying the plurality of electronic messages comprises identifying a thread of messages about a same subject and analyzing messages within the thread in context with other messages within the thread.
(Paragraph 17 " Various examples describe techniques and architectures for a system that performs, among other things, collection or extraction of tasks from databases, user accounts, and electronic communications, such as messages between or among one or more users (e.g., a single user may send a message to oneself or to one or more other users). For example, a system may extract a set of tasks from a calendar application associated with one or more users. In another example, an email exchange between two people may include text from a first person sending a request to a second person to perform a task. The email exchange may convey enough information for the system to automatically determine the presence of the request to perform the task. In some implementations, the email exchange does not convey enough information to determine the presence of a task. Whether or not this is the case, the system may query other sources of information that may be related to one or more portions of the email exchange. For example, the system may examine other messages exchanged by one or both of the authors of the email exchange or by other people. The system may also examine larger corpora of email and other messages. Beyond other messages, the system may query a calendar or database of one or both of the authors of the email exchange for additional information. In some implementations, the system may, among other things, query traffic or weather conditions at respective locations of one or both of the authors.")
Claim 4,12,20
Regarding Claim 4,12,20, Jothilingam in view of DAN, in further view of Lange, further more Jothilingam teaches the method of claim 1, wherein receiving user input related to the potential action comprises receiving a modification of the potential action provided by the user, wherein performing the subsequent action comprises performing the new potential action.
(Paragraph 79 "At block 1008, task operations module 402 may provide a list of the task-oriented actions to the user for inspection or review. For example, a task-oriented action may be to find or locate digital artefacts (e.g., documents) related to a particular task to support completion of, or user comprehension of, a task activity. At diamond 1010, the user may select among choices of different possible actions to be performed by task operations module 402, refine possible actions, delete actions, manually add actions, and so on. If there are any such changes, then process 1000 may return to block 1004 where task operations module 402 may re-generate task-oriented processes in view of the user's edits of the task-oriented process list. On the other hand, if the user approves the list, then process 1000 may proceed to block 1012 where task operations module 402 performs the task-oriented processes. At block 1014, the task operations module may generate and display a visual cue and productivity report, for example.")
Jothilingam do not explicitly teach all of (the bolded ) via a chat interface and generating a new potential action using the large language model
Although the machine learning model that Jothlingam uses can be an LLM and any of the actions identified could be performed by an LLM agent.
Although Jothlingam teaches generating a “description” of the action in “generating, by the large language model, a description of a potential action” it does not expressly teach “suggesting … the generated description … in a user interface.” However, as provided above, the name of the action could be descriptive and provide a description of the task. For example, Email is descriptive as the name and the description of the task of emailing.
However, Jothilingam in view of DAN, in further view of Lange, further more DAN. teach via a chat interface and generating a new potential action using the large language model
(paragraph 25 provisional paragraph 24 "The memory 112 further includes a bot 116 that comprises or has access to a generative model 118. In an example, the generative model 118 is a transformer-based model. The generative model 118 can be a GLM, such that the generative model 118 is configured to receive text input and generate text output. In such case, the bot 116 can be a chatbot. In other examples, the generative model 118 can generate audio, images, video (with audio) or other suitable outputs based upon a variety of inputs, such as voice, text, video, images, audio, etc. While the bot 116 is illustrated as being external to the messaging application 114, it is contemplated that the bot 116 may be included in the messaging application 114. As will be described in greater detail herein, the messaging application 114 receives messages from client computing devices 104-106 that are part of a group conversation between users of the client computing devices 104-106. The bot 116, through use of the generative model 118, generates output based upon the received messages and the output is provided to the client computing devices 104-106 as part of the group conversation. Hence, the bot 116 can participate in the group conversation.")
See claim 1 for rationale.
Claim 5 and 13
Regarding Claim 5 and 13, Jothilingam in view of DAN, in further view of Lange, further more Jothilingam teaches the method of claim 1, wherein performing the subsequent action comprises creating a digital calendar event on behalf of the user.
(Paragraph 22 "Once identified and extracted by a computing system, a task (e.g., the proposal or affirmation of a commitment or request) of a communication may be further processed or analyzed to identify or infer semantics of the commitment or request including: identifying the primary owners of the request or commitment (e.g., if not the parties in the communication); the nature (e.g., type) of the task and its properties (e.g., its description or summarization); specified or inferred pertinent dates (e.g., deadlines for completing the commitment or request); relevant responses such as initial replies or follow-up messages and their expected timing (e.g., per expectations of courtesy or around efficient communications for task completion among people or per an organization); and information resources to be used to satisfy the request. Such information resources, for example, may provide information about time, people, locations, and so on. The identified task and inferences about the task may be used to drive automatic (e.g., computer generated) services such as reminders, revisions (e.g., and displays) of to-do lists, prioritization of tasks, appointments, meeting requests, and other time management activities. In some examples, such automatic services may be applied during the composition of a message (e.g., typing an email or text), reading the message, or at other times, such as during offline processing of email on a server or client device. The initial extraction and inferences about a task may also invoke automated services that work with one or more participants to confirm or refine current understandings or inferences about the task and the status of the task based, at least in part, on the identification of missing information or of uncertainties about one or more properties detected or inferred from the communication.")
Claim 8 and 16
Regarding Claim 8, 16, Jothilingam in view of DAN, in further view of Lange, further more Jothilingam teaches the method of claim 1, wherein suggesting the potential action to the user comprises generating, a description of the potential action and displaying the description in a user interface.
(Paragraph 79 "At block 1008, task operations module 402 may provide a list of the task-oriented actions to the user for inspection or review. For example, a task-oriented action may be to find or locate digital artefacts (e.g., documents) related to a particular task to support completion of, or user comprehension of, a task activity. At diamond 1010, the user may select among choices of different possible actions to be performed by task operations module 402, refine possible actions, delete actions, manually add actions, and so on. If there are any such changes, then process 1000 may return to block 1004 where task operations module 402 may re-generate task-oriented processes in view of the user's edits of the task-oriented process list. On the other hand, if the user approves the list, then process 1000 may proceed to block 1012 where task operations module 402 performs the task-oriented processes. At block 1014, the task operations module may generate and display a visual cue and productivity report, for example.")
Jothilingam do not explicitly teach all of by the large language model,
However, Jothilingam in view of DAN, in further view of Lange, further more DAN. teach by the large language model,
(paragraph 25 "The memory 112 further includes a bot 116 that comprises or has access to a generative model 118. In an example, the generative model 118 is a transformer-based model. The generative model 118 can be a GLM, such that the generative model 118 is configured to receive text input and generate text output. In such case, the bot 116 can be a chatbot. In other examples, the generative model 118 can generate audio, images, video (with audio) or other suitable outputs based upon a variety of inputs, such as voice, text, video, images, audio, etc. While the bot 116 is illustrated as being external to the messaging application 114, it is contemplated that the bot 116 may be included in the messaging application 114. As will be described in greater detail herein, the messaging application 114 receives messages from client computing devices 104-106 that are part of a group conversation between users of the client computing devices 104-106. The bot 116, through use of the generative model 118, generates output based upon the received messages and the output is provided to the client computing devices 104-106 as part of the group conversation. Hence, the bot 116 can participate in the group conversation."
Paragraph 50-51 "As expressed above, the technologies described herein allow for use cases that were heretofore not possible in connection with messaging applications that allow for group conversations. For instance, the bot 116 can answer questions related to the context of the group conversation, as well as general questions. In some implementations, the bot 116 generates output only when invoked by a user who is participating in the group conversation. In another example, the bot 116 is provided with input for each turn of a group conversation, and the bot 116 decides when it would be useful to inject itself into the group conversation by generating an output. In still yet another example, the messaging application 114 is associated with a UX canvas that is displayed on the displays of the client computing devices 104-106, where the UX canvas is continuously updated with outputs of the bot 116 (where the outputs are suggestions for inclusion in the group conversation). When a user taps on a suggestion, such suggestion is entered into the group conversation as a turn (and identified as being generated by the bot 116).
[0051] The bot 116 can assist users in the group conversations with various tasks. Examples of tasks that the bot 116 can assist with include, but are not limited to: 1) summarizing the current conversation; 2) answering a question about what a person said about a topic in the conversation; 3) proposing follow up action items based upon information in the group conversation; 4) helping the group plan events, such as vacations; 5) generating images, avatars, memes, etc. pertaining to messages included in the group conversation; 6) helping to rewrite and check text that the group is working on-making text more professional, making text more humorous, fixing typos in the text, translating text to different languages, and so forth; 7) answering questions about the contents of a webpage pointed to by a URL that has been placed in the group conversation, such as summarizing news articles, identifying facts mentioned in a news article, reading structured data such as company earnings documents, answering specific questions about content of the web page, etc.; 8) brainstorming ideas, such as producing road maps for product development; 9) translating text from one language to another, such that languages of participants in the group does not matter, as the bot 116 can translate and convey information between participants in languages requested by the participants; and 10) starting and running text and image-based games to entertain users in the group conversation.")
See claim 1 for rationale.
Claims 6, 14, 7, and 15 are rejected under 35 U.S.C. 103 as obvious over Jothilingam, Dan, and Lange in further view of Chow (US Patent US 20220215351).
Claim 6 and 14
Regarding Claim 6 and 14 , Jothilingam in view of DAN, in further view of Lange do not explicitly teach all of the method of claim 1, wherein performing the subsequent action comprises sending an electronic message on behalf of the user.
However, Chow teaches the method of claim 1, wherein performing the subsequent action comprises sending an electronic message on behalf of the user.
(Paragraph 27 "The email client 104 can send and receive emails 115 by utilizing an email server 110. The email server 110 can be a remote server that stores information about a user's email account, such as the emails 115 currently in the user's inbox, and provides that information to the user device 100. The email server 110 can also send outgoing emails on behalf of the user device 100, to be received by a separate server or device on the other end of the communication chain. In some examples, a secure email gateway (“SEG”) 120 can route email 115 to the email server 110. The SEG 120 can be either incorporated into the email server 110 or provided as a standalone gateway server. The SEG 120 can communicate with the management server 130 and implement rules for allowing or disallowing access to the email server 110. In one example, the SEG 120 can route email 115 to the mail server 110, but also send a copy of the email 115 to the management server 130 for task scheduling purposes."
Paragraph 52 " At stage 315, the SEG 120 or mail server 110 can filter for task-related emails. This can include parsing the email by applying an ML model 135 to determine intents and slots, or simply looking for keywords. A task can be identified based on recognition of a relationship between sender and recipient, reference to a backend system 140, reference to a project, or reference to a document, among other ways. If a task is identified, the task-related email can be processed for task scheduling. In one example, this can include sending a copy of that email to the management server 130 for task processing.")
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified , Jothilingam in view of DAN, in further view of Lange to incorporate the teachings of Chow to provide a “The method of claim 1, wherein performing the subsequent action comprises sending an electronic message on behalf of the user.” Doing so would enable the individual to automate messages and other actions, as recognized by Chow. (Paragraph 27).
Claim 7 and 15
Regarding Claim 7 and 15 , Jothilingam in view of DAN, in further view of Lange do not explicitly teach all of the method of claim 1, wherein performing the subsequent action comprises sending input to an application programming interface.
However, Chow teaches the method of claim 1, wherein performing the subsequent action comprises sending input to an application programming interface.
(Paragraph 8 "The task recognized by the service can then be used as an input to the retrieved ML model, along with the available time slots. The ML model can then output a time within the available timeslots that the user is most likely to perform the task. The service can then schedule the task within the time slot based on the result from the machine learning model. To schedule the task, the service can make an application programming interface (“API”) call to the calendar application or backend database, in an example. This can cause the task to show up on the user device within a calendar or to-do list that displays on the user device. A notification can also display on the user device as the time draws near for performing the task.")
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified , Jothilingam in view of DAN, in further view of Lange to incorporate the teachings of Chow to provide a “the method of claim 1, wherein performing the subsequent action comprises sending input to an application programming interface.” Doing so would remind the user to complete a task, as recognized by Chow. (Paragraph 8).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALI M HASSAN whose telephone number is (571)272-5331. The examiner can normally be reached Monday - Friday 8:00am - 4:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras Shah can be reached at (571)270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALI M HASSAN/ Examiner, Art Unit 2653
/FARIBA SIRJANI/Primary Examiner, Art Unit 2659