Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the amendment and remarks filed July 1st, 2025. In the
amendment, claims 1-2, 7, 9, and 13-14 were amended, claims 6 and 18 were cancelled, and claims 21-22 were added. As such, claims 1-5, 7-17, and 19-22 are pending.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on June 9th, 2026, May 15th 2026, May 12th 2026, December 16th 2025, December 1st 2025, October 31st 2025, July 11th 2025, June 10th 2025, are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Response to Arguments
Applicant’s arguments, see Pages 7 and 10-11, filed July 1st, 2025, with respect to the double patenting and 35 U.S.C 102 rejections have been fully considered and are persuasive. The double patenting and 35 U.S.C 102 rejections have been withdrawn.
Applicant’s arguments, see Page 7, filed July 1st, 2025, with respect to the 35 U.S.C 112 (b) rejections have been fully considered and are persuasive. Amendments to the claims obviate the rejections of record. The rejections of claims have been withdrawn. See updated rejections below.
Applicant’s arguments, in view of the amendment filed July 1st, 2025, with respect to the rejections of claims 1-5, 7-17, and 19-22 under 35 U.S.C § 101 and 103, are maintained and not persuasive for the following reasons:
35 U.S.C. 101:
Applicant argues that the claims do not recite a mental process because the recited limitations cannot practically be performed in the human mind, relying on SiRF Tech. and MPEP § 2106.04(a)(2)(III)(A) (Pages 7-8 of Remarks). The Examiner respectfully disagrees. SiRF Tech. is distinguishable because the claims there required a GPS receiver that calculated coordinates estimating the distance from the receiver to a plurality of satellites where the manipulation of satellite signal data was tethered to a physical device and could not be performed in the mind. Here, by contrast, generating a prompt, determining a model output, and processing that output are evaluation and judgment steps. The claim also doesn’t recite any mathematical relationships, formulas, equations, or calculations that would place the operation beyond human mental capability, as explained under MPEP § 2106.04(a)(2)(III). The multimodal machine learning model is merely invoked as a tool, and the claims do not remove a limitation from the mental processes grouping simply because a generic computer or model performs it (MPEP § 2106.04(a)(2)(III)(C)).
Applicant argues that the claims integrate any alleged abstract idea into a practical application because they reflect a technical improvement to productivity applications (Pages 8-9 of Remarks). The Examiner respectfully disagrees. The purported improvement is an improvement to the abstract idea itself rather than an improvement to the functioning of the computer or to the operation of the machine learning model. The claims recite a multimodal machine learning model at a high level that generates output from user input and a prompt in order to modify the productivity object of a productivity application. As set forth in the rejection, the processor, memory, and multimodal model amount to mere instructions to apply the exception using a computer as a tool (MPEP § 2106.05(f)), while controlling functionality of the productivity application merely links the exception to a particular technological environment or field of use (MPEP § 2106.05(h)). The receiving and modifying steps are insignificant extra-solution activity in the form of data gathering and data output (MPEP § 2106.05(g)).
Applicant argues that the additional elements are unconventional and amount to more than well-understood, routine, conventional activity (Pages 9-10 of Remarks). The Examiner respectfully disagrees. The features Applicant identifies as the point of novelty are part of the recited abstract idea, and a judicial exception cannot itself supply the inventive concept required at Step 2B. The additional elements of at least one processor, memory storing instructions, and the multimodal machine learning model are recited at a high level of generality and use generic computer components as tools that are applied to the abstract idea. The model receives a natural language input at a high level which is analyzed under well-understood, routine, and conventional (MPEP § 2106.05(d)). The receiving step amounts to mere data gathering over a network which is analyzed under well-understood, routine, and conventional (MPEP § 2106.05(d)). The elements considered individually and as an ordered whole does not add significantly more than the abstract idea. Therefore, the rejection under 35 U.S.C. 101 is maintained.
35 U.S.C. 103:
Applicant’s arguments with respect to claim(s) 1-5, 7-17, and 19-22 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Specification
The disclosure is objected to because of the following informalities:
Paragraph 17 states that "[a] generative multimodal machine learning model (also generally referred to herein as a multimodal machine learning model herein)," recites "herein" twice. The sentence should instead read "[a] generative multimodal machine learning model (also generally referred to herein as a multimodal machine learning model)."
Paragraph 30: "Returning to the example or programmatic output" should read "Returning to the example of programmatic output"
Paragraph 50: "aspects of productivity application 118 … may instead by implemented remotely" should read "may instead be implemented remotely"
Paragraph 111 recites "In an example, determining the model output comprises: providing, to a multimodal generative platform, an indication of the user input; and receiving, from the multimodal generative platform, the model output," and then repeats the identical sentence again as "In another example, determining the model output comprises…". One of the duplicate recitations should be deleted or amended to recite a distinct example.
Appropriate correction is required.
Claim Objections
Claims 1, 13, 19, 21 and their corresponding dependent claims are objected to because of the following informalities:
Claim 1 recites "memory storing instructions that, when executed by the at least one processor, causes the system to perform," which should read "memory storing instructions that, when executed by the at least one processor, cause the system to perform…"
It is unclear in claim 19 if "a prompt associated with the productivity application" is meant to refer to "the generated prompt" recited in claim 13. The examiner suggests amending the limitation to read "the generated prompt associated with the productivity application."
Claim 21 recites "application programing interface (API)," which should read "application programming interface (API)."
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 5 and 17 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 5 and 17 recite the limitation "and the first model output and the second model output are different" in the second limitation. There is insufficient antecedent basis for this limitation in the claim. Examiner suggests removing the word first from “the first model output.”
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-5, 7-17, and 19-22 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claim 1
Step 1: The claim recites a system comprising at least one processor; therefore, it is directed to the statutory category of a machine.
Step2A Prong 1: The claim recites, inter alia:
generating a prompt for priming a … model, wherein the prompt is associated with at least one of the productivity application or a productivity object of the productivity application: This limitation recites a mental process because it involves the evaluation/judgement/opinion to generate a prompt associated with an application for priming a model, which can be performed by pen and paper.
determining, using the generated prompt and the received user input, a model output associated with the … model: This limitation recites a mental process because it involves the evaluation/judgement/opinion to determine a model output based on the received input.
processing the model output…: This limitation recites a mental process because it involves the evaluation/judgement/opinion to process/analyze the model output.
as a result of processing the model output, modifying the productivity object of the productivity application: This limitation recites a mental process because it involves the evaluation/judgement/opinion to process the model output and modify a productivity object, which can be performed mentally.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
[a] system comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, causes the system to perform a set of operations, the set of operations comprising: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
receiving, at a productivity application, a natural language user input: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
…a multimodal machine learning model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
…to control functionality of the productivity application: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
[a] system comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, causes the system to perform a set of operations, the set of operations comprising: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
receiving, at a productivity application, a natural language user input: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and is well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
a multimodal machine learning model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f))
.…to control functionality of the productivity application: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
The elements in combination as an ordered whole still do not amount to significantly more than the judicial exception (i.e., the mental processes for generating a prompt, determining a model output, and processing that model output). The claim merely describes a process of applying known machine learning techniques (priming a multimodal machine learning model with a generated prompt and user input and determining a model output) together with generic data gathering steps (receiving a natural language user input). The recitation of at least one processor and memory storing instructions, a productivity application, a productivity object, and a multimodal machine learning model merely indicates a technological environment in which the abstract ideas are applied, without improving the functioning of a computer or the machine learning model itself.
Therefore, the claim as a whole remains focused on the abstract idea and fails Step 2B of the eligibility
analysis.
Claim 2
Step 1: A machine, as above.
Step2A Prong 1:
determining the model output comprises…: This limitation recites a mental process because it involves the evaluation/judgement/opinion to determine a model output.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
providing, to a multimodal generative platform, an indication of the generated prompt and the user input: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
receiving, from the multimodal generative platform, the model output: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
a multimodal generative platform: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
providing, to a multimodal generative platform, an indication of the generated prompt and the user input: The additional element of “providing” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of providing steps amounts to no more than mere data gathering. This element amounts to sending/transmitting data over a network and is well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
receiving, from the multimodal generative platform, the model output: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and is well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
a multimodal generative platform: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 3
Step 1: A machine, as above.
Step2A Prong 1: This claim does not recite any additional abstract ideas but depends on claim 1 which recites an abstract idea.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
wherein the set of operations further comprises: storing the received user input and the determined model output as part of a context: This limitation is merely a post-solution step of storing the data—a nominal addition to the claim that does not meaningfully limit the claim. The method storing is recited at a high level of generality. Simply implementing the abstract idea in a generic method is not a practical application of the abstract idea. Therefore, storing step is an insignificant extra-solution activity. See MPEP 2106.05(g).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
wherein the set of operations further comprises: storing the received user input and the determined model output as part of a context: These elements amount to storing… information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93; See MPEP 2106.05(d) (II)(iv). The courts have recognized the computer functions of storing as well‐understood, routine, and conventional function when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity.
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 4
Step 1: A machine, as above.
Step2A Prong 1: The claim recites, inter alia:
determining, based on the second user input and the context, a second model output associated with the multimodal machine learning model: This limitation recites a mental process because it involves the evaluation/judgement/opinion to determine a model output based on the received input.
processing the second model output…: This limitation recites a mental process because it involves the evaluation/judgement/opinion to process/analyze the second model output.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
the set of operations further comprises: receiving a second user input: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
…to control functionality of the productivity application to modify the productivity object: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
the set of operations further comprises: receiving a second user input: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and is well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
…to control functionality of the productivity application to modify the productivity object: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 5
Step 1: A machine, as above.
Step2A Prong 1: This claim does not recite any additional abstract ideas but depends on claim 4, which recites an abstract idea.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
the received user input and the second user input are the same; and the first model output and the second model output are different: The limitation merely describes the type of data being processed and thus amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
the received user input and the second user input are the same; and the first model output and the second model output are different: The limitation merely describes the type of data being processed and thus amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 7
Step 1: A machine, as above.
Step2A Prong 1: The claim recites, inter alia:
the model output is further determined based on a context associated with a different productivity application: This limitation recites a mental process because it involves determining the model output being based on contextual information associated with an application.
Step 2A Prong Two and Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception under step 2B. Thus, the judicial exception is not integrated into a practical application (see MPEP 2106.04(d) I.), failing step 2A prong 2. The claim is ineligible.
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 8
Step 1: A machine, as above.
Step2A Prong 1: This claim does not recite any additional abstract ideas but depends on claim 1 which recites an abstract idea.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
modifying the productivity object of the productivity application comprises one or more of: changing formatting in the productivity object; changing a transition in the productivity object; adding graphical content to the productivity object; adding audio content to the productivity object; adding textual content to the productivity object; or generating a formula in the productivity object: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
modifying the productivity object of the productivity application comprises one or more of: changing formatting in the productivity object; changing a transition in the productivity object; adding graphical content to the productivity object; adding audio content to the productivity object; adding textual content to the productivity object; or generating a formula in the productivity object: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 9
Step 1: The claim recites a method; therefore, it is directed to the statutory category of a process.
Step2A Prong 1: The claim recites, inter alia:
generating a prompt for priming the multimodal machine learning model, wherein the prompt is associated with at least one of the productivity application or the document: This limitation recites a mental process because it involves the evaluation/judgement/opinion to generate a prompt associated with an application for priming a model, which can be performed by pen and paper.
determining, based on the generated prompt and the received user input, a model output of the multimodal machine learning model: This limitation recites a mental process because it involves the evaluation/judgement/opinion to determine a model output based on the received input.
processing the model output…: This limitation recites a mental process because it involves the evaluation/judgement/opinion to process/analyze the model output.
as a result of processing the model output, modifying the document: This limitation recites a mental process because it involves the evaluation/judgement/opinion to process the model output and modify a document, which can be performed in the human mind and by pen and paper.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
receiving, at a productivity application, user input associated with a document: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
a multimodal machine learning model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
wherein the model output does not include content to include in the document: The limitation merely describes the type of data being processed and thus amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
…to control functionality of the productivity application: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
receiving, at a productivity application, user input associated with a document: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and is well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
a multimodal machine learning model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f))
wherein the model output does not include content to include in the document: The limitation merely describes the type of data being processed and thus amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
…to control functionality of the productivity application: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 10
Step 1: A process, as above.
Step2A Prong 1: The claim recites, inter alia:
determining, based on the second user input and the context, a second model output associated with the multimodal machine learning model: This limitation recites a mental process because it involves the evaluation/judgement/opinion to determine a model output based on the received input.
and processing the second model output to modify the document: This limitation recites a mental process because it involves the evaluation/judgement/opinion to process/analyze the second model output.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
the model output is a first model output and the method further comprises: storing the received user input and the determined model output as part of a context: This limitation is merely a post-solution step of storing the data—a nominal addition to the claim that does not meaningfully limit the claim. The method storing is recited at a high level of generality. Simply implementing the abstract idea in a generic method is not a practical application of the abstract idea. Therefore, storing steps is an insignificant extra-solution activity. See MPEP 2106.05(g).
receiving a second user input: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
the model output is a first model output and the method further comprises: storing the received user input and the determined model output as part of a context: These elements amount to storing… information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93; See MPEP 2106.05(d) (II)(iv). The courts have recognized the computer functions of storing as well‐understood, routine, and conventional function when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity.
receiving a second user input: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and is well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 11 is a process claim that recites similar limitations to claim 5. Therefore, claim 11 is rejected using the same rationale as claim 5.
Claim 12
Step 1: A process, as above.
Step2A Prong 1: This claim does not recite any additional abstract ideas but depends on claim 9 which recites an abstract idea.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
the model output includes a set of programmatic steps associated with the functionality of the productivity application: The limitation merely describes the type of data being processed and thus amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
processing the model output comprises executing the set of programmatic steps to control functionality of the productivity application: Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
the model output includes a set of programmatic steps associated with the functionality of the productivity application: The limitation merely describes the type of data being processed and thus amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
processing the model output comprises executing the set of programmatic steps to control functionality of the productivity application: Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 13
Step 1: The claim recites a method; therefore, it is directed to the statutory category of a process.
Step2A Prong 1: The claim recites, inter alia:
generating a prompt for priming the … model, wherein the prompt is associated with at least one of the productivity application or a document of the productivity application: This limitation recites a mental process because it involves the evaluation/judgement/opinion to generate a prompt associated with an application for priming a model, which can be performed by pen and paper.
determining, using the generated prompt and the received user input, a model output associated with the multimodal machine learning model: This limitation recites a mental process because it involves the evaluation/judgement/opinion to determine a model output based on the received input.
processing the model output…: This limitation recites a mental process because it involves the evaluation/judgement/opinion to process/analyze the model output.
as a result of processing the model output, modifying the productivity object of the productivity application: This limitation recites a mental process because it involves the evaluation/judgement/opinion to process the model output and modify a productivity object, which can be performed mentally.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
receiving, at a productivity application, a natural language user input: Mere data gathering recited at a high level of generality, and thus is an insignificant extra-solution activity (MPEP 2106.05(g)).
…a multimodal machine learning model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
…to control functionality of the productivity application: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
receiving, at a productivity application, a natural language user input: The additional element of “receiving” does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of receiving steps amounts to no more than mere data gathering. This element amounts to receiving data over a network and is well-understood, routine, conventional activity. See MPEP 2106.05(d), subsection II (i). This cannot provide an inventive concept.
a multimodal machine learning model: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f))
.…to control functionality of the productivity application: The limitation amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 14 is a method claim that recites similar limitations to claim 2. Therefore, claim 14 is rejected using the same rationale as claim 2.
Claim 15 is a method claim that recites similar limitations to claim 3. Therefore, claim 15 is rejected using the same rationale as claim 3.
Claim 16 is a method claim that recites similar limitations to claim 4. Therefore, claim 16 is rejected using the same rationale as claim 4.
Claim 17 is a method claim that recites similar limitations to claim 5. Therefore, claim 17 is rejected using the same rationale as claim 5.
Claim 19 is a method claim that recites similar limitations to claim 7. Therefore, claim 19 is rejected using the same rationale as claim 7.
Claim 20 is a method claim that recites similar limitations to claim 8. Therefore, claim 20 is rejected using the same rationale as claim 8.
Claim 21
Step 1: A process, as above.
Step2A Prong 1: This claim does not recite any additional abstract ideas but depends on claim 12, which depends on claim 9, which recites an abstract idea.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
the set of programmatic steps comprises a call to an application programming interface (API) of the productivity application: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
the set of programmatic steps comprises a call to an application programming interface (API) of the productivity application: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim 22
Step 1: A machine, as above.
Step2A Prong 1: This claim does not recite any additional abstract ideas but depends on claim 1 which recites an abstract idea.
Step2A Prong 2: This judicial exception is not integrated into a practical application because the
additional elements are as follows:
the model output comprises a set of programmatic steps: The limitation merely describes the type of data being processed and thus amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself (MPEP 2106.05(h)).
modifying the productivity object comprises executing the set of programmatic steps to modify the productivity object of the productivity application: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).
Step 2B: The claim does not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the additional elements are as follows:
the model output comprises a set of programmatic steps: The limitation merely describes the type of data being processed and thus amounts to merely indicating a field of use or technological environment in which to apply a judicial exception. This does not amount to significantly more than the exception itself which cannot provide inventive concept (MPEP 2106.05(h)).
modifying the productivity object comprises executing the set of programmatic steps to modify the productivity object of the productivity application: Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea and cannot provide inventive concept (MPEP 2106.05(f)).
Even when considered in combination, these additional elements represent mere instructions to apply
an exception and therefore do not provide an inventive concept. The claim is ineligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-5, 8-17, and 20-22 are rejected under 35 U.S.C. 103 as being unpatentable over Kovacs (US 20190197402 A1) in view of Keskar (“CTRL: A CONDITIONAL TRANSFORMER LANGUAGE
MODEL FOR CONTROLLABLE GENERATION”, 2019), and in further view of Singh (US 20220230061 A1).
Regarding claim 1,
Kovacs teaches [a] system comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, causes the system to perform a set of operations (Paragraph 33 of Kovacs, "The invention can be implemented in numerous ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer readable storage medium; and/or a processor, such as a processor configured to execute instructions stored on and/or provided by a memory coupled to the processor… being configured to perform a task may be implemented as a general component...");
receiving, at a productivity application, a natural language user input (See Figure 19, Paragraph 162 of Kovacs, "At 1905, natural language (NL) input is received. For example, the output of speech recognition at 1903 is a natural language input. In some embodiments, the speech is converted to natural language input so that it can be used to update the belief set of the agent and to determine possible actions.", Paragraph 106 of Kovacs, "In various embodiments, conversational artificial intelligence (AI) server 1051 is utilized by a computer-based game to store, process, and/or transmit conversations between game entities including between players such as a mix of human players and/or non-player characters (NPCs)... In various embodiments, dialog of a conversation is processed as the auditory input and/or output for each participant of the conversation. For example, an autonomous AI NPC receives and sends auditory input via conversational AI server 1051 and in-game environment 1031."
The game application is executed on a game device which receives a spoken/typed sentence from the player, converts it to natural language input at element 1905, and passes it into the application's AI logic. Kovacs's game application is a productivity application because it is an application in which a user provides natural-language input to a conversational agent to control the application and act on its objects. The player directs input to an NPC agent, and the agent's resulting action modifies an object in the game environment. The received sentence corresponds to the natural language user input, and the game application executing on game device is the receiving application.);
generating a prompt for … a … model, wherein the prompt is associated with at least one of the productivity application or a productivity object of the productivity application (Paragraph 163 of Kovacs, "In some embodiments, natural language speech segments are sent to the conversational AI server of 1921 with the identity (e.g., name or user identifier, etc.) of the agent.", Paragraph 165, "In some embodiments, the natural language input sentences are received as speech segments and include an agent identifier... In the event the input sentence contains specific data, including basic information, the conversational AI server updates the saved information. In some embodiments, information received by the conversational AI server is stored using the episodic process of 1923 in an individual memory associated with the agent.", Paragraph 42, "In some embodiments, a scalable artificial intelligence (AI) framework is achieved by using a machine learning model. For example, a machine learning model is used to determine the actions of each autonomous AI character based on the individual character's input, beliefs, and/or goals.", Paragraph 43, "For example, an autonomous AI non-player character (NPC) receives visual and audio input such as detecting movement of other objects in the game environment and/or recognizing voice input from other users... using a machine learning model such as a deep convolutional neural network (DCNN), a solution... is determined based on the sensory inputs, the beliefs, and the goals of the autonomous AI NPC.", Paragraph 102, "In some embodiments, solver 1021 utilizes a light-weight deep neural network (DNN) planner. In some embodiments, solver 1021 utilizes a deep convolutional neural network (DCNN).", Paragraph 105, "In some embodiments, changes to in-game environment 1031, including auditory and visual changes detected by auditory and visual sensor 1035, relevant to an AI character, are encoded into perception input 1033 and used as input to AI game component 1011 including input to AI loop 1017."
The agent identifier that the system appends to the natural language speech segments is the generated prompt. It is associated with the specific agent which is the productivity object. The server uses the identifier to select and update that agent's individual memory, which becomes the belief state the model reasons over.);
determining, using the generated prompt and the received user input, a model output associated with the … model (Paragraph 163 of Kovacs, "At 1907, an agent controller receives the natural language (NL) input of 1905, responses from the conversational artificial intelligence (AI) server of 1921, and perception input of 1919... In some embodiments, the natural language input is then fed to the conversational AI server of 1921 and the belief update of 1909.", Paragraph 164, "At 1909, a belief is updated. For example, a belief is updated based on the natural language (NL) input of 1905 (as processed by the agent controller of 1907) as well as the in-game environment of 1915... any changes to the belief set can change the artificial intelligence (AI) planning problem and the resulting actions of the agent's action list.", Paragraph 189, “Intelligent user interface (UI)—A dynamic natural language-based graphical user interface (GUI) automatically and most ergonomically adapts to a dynamically changing game context and artificial intelligence problems that constantly change based on changes in the game state… For example, in some embodiments, a player can invoke a GUI menu without touching the screen by saying the word “menu.””, Paragraph 233, “the actions of the human player character are controlled by a human user…” Paragraph 101, "Solver 1021 is used to generate a solution, such as solution plan 1023, to the provided AI planning problem.",
The deep neural network solver 1021 produces the solution plan, which is the model output. The generated prompt is the natural-language sentence the user directs to the agent, and the received user input is the perception input reflecting the user's control of the game since the actions of the human player character are controlled by a human user. Both are provided to the model and are used to update the belief set which defines the AI planning problem the model solves. The deep neural network solver then produces the solution plan, which is the model output.).
processing the model output to control functionality of the productivity application; and (Paragraph 104 of Kovacs, "During each iteration of artificial intelligence (AI) loop 1017, the first action of solution plan 1023 is extracted and encoded using first action extractor and encoder 1025. The extracted action is executed using action executor 1027. For example, the action may include a movement action, a speaking action, an action to end a conversation, etc... In various embodiments, the actions performed by NPCs influence the game environment and in turn the environment also influences the NPCs during the next iteration of AI loop 1017.”, Paragraph 248, "In the event an AI action is scheduled for execution by the AI of an agent, the associated game actions are executed... The implementation of the ReceiveActionFromAI function calls the game code (depending on the game logic)."
The solution plan output by the DNN is consumed by action executor, which invokes the application's own functions.);
as a result of processing the model output, modifying a productivity object of the productivity application. (Paragraph 158 of Kovacs, "When a player talks or inputs a sentence to an NPC, the game system recognizes the provided input and utilizes a conversational AI server to process the voice input and provide an answer… In some embodiments, the answer is provided both in text and audibly. For example, in some embodiments, an audible version of the answer is provided using a text-to-speech technique while the text is displayed on a game device display. In some embodiments, these changes also influence the flow of the game, which causes a belief update and subsequent changes to the game environment.", Paragraph 114, "In some embodiments, the action of an NPC agent can influence the objects in the game environment. For example, an NPC can move an object, carry an object, and/or consume an object, etc."
Executing the action selected by the DNN changes the state of an object maintained by the application. The object whose state is changed is the productivity object.)
Kovacs does not teach generating a prompt for priming a … model.
Keskar, in the same field of endeavor, teaches generating a prompt for priming a … model (Page 1 Introduction, “…we train a language model that is conditioned on a variety of control codes … that make desired features of generated text more explicit. With 1.63 billion parameters, our Conditional Transformer Language (CTRL) model can generate text conditioned on control codes that specify domain, style, topics, dates, entities, relationships between entities, plot points, and task-related behavior.”, Page 2 Section 3, “CTRL is a conditional language model that is always conditioned on a control code c and learns the distribution p(x|c). The distribution can still be decomposed using the chain rule of probability and trained with a loss that takes the control code into account.”
The control code that the system generates and adds to the input is the generated prompt, and the model is conditioned on that code to learn p(x|c), which is the priming.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Kovacs’ teaching of updating the game environments based on user input with Keskar’s control code conditioning in order to give explicit and predictable control over the model's output (Introduction of Keskar).
Kovacs in view of Keskar does not teach generating a prompt for … a multimodal machine learning model.
Analogous art Singh teaches a multimodal machine learning model (Claim 1 of Singh, “…accessing, by the multimodal query subsystem, a multimodal question-answering model comprising (a) a first stream of language models comprising a first set of transformer-based models concatenated with each other and (b) a second stream of language models comprising a second set of transformer-based models concatenated with each other, wherein each transformer-based model comprises a respective cross-attention layer using data generated by both the first stream of language models and the second stream of language models as input;”, Claim 6, “…the multimodal question-answering model comprises a first stream of language models configured for text content and a second stream of language models configured for image content”
Singh's multimodal question-answering model processes multiple content types (text content via the first stream of transformer-based language models and image content via the second stream) and interoperates between those content types through cross-attention layers that pass data between the two streams.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Kovacs in view of Keskar’s teachings with Singh's multimodal question-answering model in order to enable the model to process and interoperate across multiple content types such as text and image content to improve tools for processing text (Paragraph 18 and claims 1 and 6 of Singh).
Regarding claim 2,
Kovacs teaches determining the model output comprises: providing, to a multimodal generative platform, an indication of the generated prompt and the user input; and receiving, from the multimodal generative platform, the model output (See Figure 19, Paragraph 163, "At 1907, an agent controller receives the natural language (NL) input of 1905, responses from the conversational artificial intelligence (AI) server of 1921, and perception input of 1919… In some embodiments, the natural language input is then fed to the conversational AI server of 1921 and the belief update of 1909. In some embodiments, natural language speech segments are sent to the conversational AI server of 1921 with the identity (e.g., name or user identifier, etc.) of the agent. In some embodiments, the conversational AI server of 1921 generates an answer received by the agent controller for the natural language input sentences.", Paragraph 164, "At 1909, a belief is updated. For example, a belief is updated based on the natural language (NL) input of 1905 (as processed by the agent controller of 1907) as well as the in-game environment of 1915.", Paragraph 166 of Kovacs, "Based on the conversational artificial intelligence (AI) server of 1921, a natural language response is created. In some embodiments, a natural language generator creates the natural language output."
The natural language speech segments sent from agent controller to conversational AI server corresponds to the indication of the user input provided to the platform. The answer generated by conversational AI server and received back at agent controller is the model output received from the platform. Conversational AI server is a generative platform because it generates a natural language answer rather than retrieving a fixed one, and it is multimodal because it both consumes and produces speech and text.).
Regarding claim 3,
Kovacs teaches the set of operations further comprises: storing the received user input and the determined model output as part of a context (Paragraph 165, "At 1921, a conversational artificial intelligence (AI) server processes the natural language input. In some embodiments, the input is stored using the episodic process of 1923 in a memory of an agent such as a non-player character (NPC) agent… In the event the input sentence contains specific data, including basic information, the conversational AI server updates the saved information. In some embodiments, information received by the conversational AI server is stored using the episodic process of 1923 in an individual memory associated with the agent.", Paragraph 103, "In some embodiments, solution plan 1023 is stored in a solution plan repository (not shown in FIG. 10). In various embodiments, each iteration of artificial intelligence (AI) loop 1017 potentially generates a new solution that may be stored in a solution plan repository… In some embodiments, a solution plan repository is implemented using a memory graph data structure and stores at least the problem, solution, and context for each solved problem."
Kovacs teaches storing the received natural language user input and storing the solution plan (the model output), and states that the repository stores the problem, solution, and context for each solved problem.).
Regarding claim 4,
Kovacs teaches the set of operations further comprises: receiving a second user input; determining, based on the second user input and the context, a second model output associated with the multimodal machine learning model; and processing the second model output to control functionality of the productivity application to modify the productivity object (See Figure 10, Paragraph 169 of Kovacs, "In some embodiments, the flow described by FIG. 19 is performed in part by AI loop 1017 of FIG. 10, game AI platform 1101 of FIG. 11, AI processor 1201 of FIG. 12, game AIs 1333 and 1343 of FIG. 13, and/or game AI component 1401 of FIG. 14.", Paragraph 101 of Kovacs, "In some embodiments, artificial intelligence (AI) loop 1017 is run until a goal is achieved and begins by searching for an AI planning problem such as initial planning problem 1019… On subsequent iterations, changes to the game environment may be reflected by changing the problem and updating the data structures corresponding to the AI problem. For example, after an initial iteration of AI loop 1017, the problem data structure is revised to address changes to the game that influence the current problem. For each iteration of AI loop 1017, solver 1021 solves the revised problem (reflected by the changes in the problem data structures) to find a best action list for the current state.", Paragraph 103, "In various embodiments, solver 1021 may retrieve a previously saved solution from the solution plan repository when encountering a previously solved problem."
The AI loop is iterative where each subsequent pass receives an input, solves a revised problem that carries forward the stored context from the prior pass, produces a further solution plan (the second model output), and executes its first action through action executor to again change the state of the object.);
Regarding claim 5,
Kovacs teaches the received user input and the second user input are the same; and the first model output and the second model output are different (Paragraph 36, "AI actions such as chasing may depend on the NPC's internal memory. For example, each sentry can implement the same AI behavior but each sentry can perform different actions based on its unique context, which may be stored in its own personal memory store.", Paragraph 42, "Its responses are then based on solving an AI planning problem using the current game environment. The actions performed are dynamic behavior and based in time with an unpredictable, changing game environment.", Paragraph 101, "For each iteration of AI loop 1017, solver 1021 solves the revised problem (reflected by the changes in the problem data structures) to find a best action list for the current state.", Paragraph 164, "In various embodiments, any changes to the belief set can change the artificial intelligence (AI) planning problem and the resulting actions of the agent's action list."
Since the model output is determined from the belief state and not from the input alone and because the stored context is updated after every iteration, an identical input presented on a later iteration is applied against a different belief state and therefore yields a different solution plan.).
Regarding claim 8,
Kovacs teaches modifying the productivity object of the productivity application comprises one or more of: changing formatting in the productivity object; (See Figure 19, Paragraph 158, "In some embodiments, a text-based answer is converted to voice speech using a text-to-speech technique. In some embodiments, the answer is provided both in text and audibly.", Paragraph 166, "In some embodiments, grammatical and word substitutions are made to generate an NL output. For example, name and the appropriate pronouns may be substituted for one another.", Paragraph 167, "At 1933, speech output is created. In some embodiments, the natural language (NL) output of 1931 is converted to speech. In some embodiments, the converted speech is in text form. In some embodiments, the text speech is displayed on a device screen of a game device for the player to read."
The grammatical and word substitutions applied at 1931 and the conversion of the same answer between text form and audible form at 1933/1935 are changes to the format in which the object's content is presented.)
changing a transition in the productivity object; (Paragraph 105, "In some embodiments, changes to in-game environment 1031, including auditory and visual changes detected by auditory and visual sensor 1035, relevant to an AI character, are encoded into perception input 1033 and used as input to AI game component 1011 including input to AI loop 1017.", Paragraph 152, "At 1713, the in-game environment is updated. For example, based on the action executed at 1711, the in-game environment is updated to reflect the action and/or consequences of the action. For example, the non-player character (NPC) agent may have consumed an object, moved an object, moved to a new location, etc. The updated in-game environment at 1713 and/or changes to the in-game environment at 1713 are used for perception input at 1705."
Executing the action carries the object from one state to a succeeding state, and the resulting state change is fed back as a percept that drives a belief update. This state-to-state change is the transition and the executed action changes it.)
adding graphical content to the productivity object (See Figure 14, Paragraph 127, "In some embodiments, percept listener 1407 receives input events from AI game framework 1451. For example, percept listener 1407 receives inputs such as audio and visual input events from AI game framework 1451 via a communication channel… In various embodiments, the actions executed by action executor 1411 are received as input events to AI game framework 1451 and the agents of AI game framework 1451 are updated."
Actions executed by the action executor produce visual events in the environment that are subsequently perceived through the visual channel of the percept listener. The visual content is added to the object as a result of the executed action.);
adding audio content to the productivity object; (See Figure 14, Paragraph 127, "For example, percept listener 1407 receives inputs such as audio and visual input events from AI game framework 1451 via a communication channel.", Paragraph 129, "For example, an action may result in the creation of a random sound or noise that can be heard by other characters."
Kovacs teaches that executing an action creates audio content associated with the object.)
adding textual content to the productivity object; (See Figure 19, Paragraph 158, "In some embodiments, the answer is provided both in text and audibly. For example, in some embodiments, an audible version of the answer is provided using a text-to-speech technique while the text is displayed on a game device display.", Paragraph 167, "In some embodiments, the text speech is displayed on a device screen of a game device for the player to read."
The generated answer text is written to and displayed as part of the object.)
or generating a formula in the productivity object (Paragraph 158, "In some embodiments, these changes also influence the flow of the game, which causes a belief update and subsequent changes to the game environment.", Paragraph 243, "In various embodiments, an AI Action can be scheduled for execution by the AI logic (i.e., by the deep neural network (DNN) based agent-controller), if and only if its preconditions are true (satisfied) in the current belief state of the agent.", Paragraph 244, "For example, a goal can be specified using a conjunction (and-relation) of conditions that should be true. If the goal formula is satisfied, the agent will believe that it has achieved its goal."
Kovacs teaches a goal formula which is a conjunction of conditions that is associated with the object and is evaluated against it.)
Regarding claim 9,
Kovacs teaches [a] method for controlling a productivity application using a … machine learning model, the method comprising: receiving, at a productivity application, user input associated with [an application] (See Figures 10 and 19, Paragraph 162, "At 1905, natural language (NL) input is received. For example, the output of speech recognition at 1903 is a natural language input. In some embodiments, the speech is converted to natural language input so that it can be used to update the belief set of the agent and to determine possible actions.", Paragraph 106, "Using at least conversational AI server 1051, a player of game device 1001 can initiate and even interrupt a conversation with one or more autonomous AI NPCs."
The game application is executed on a game device which receives a spoken/typed sentence from the player, converts it to natural language input at element 1905, and passes it into the application's AI logic. Kovacs's game application is a productivity application because it is an application in which a user provides natural-language input to a conversational agent to control the application and act on its objects. The player directs input to an NPC agent, and the agent's resulting action modifies an object in the game environment. The received sentence corresponds to the natural language user input, and the game application executing on game device is the receiving application.);
generating a prompt for … the … model, wherein the prompt is associated with at least one of the productivity application or the document (Paragraph 163 of Kovacs, "In some embodiments, natural language speech segments are sent to the conversational AI server of 1921 with the identity (e.g., name or user identifier, etc.) of the agent.", Paragraph 165, "In some embodiments, the natural language input sentences are received as speech segments and include an agent identifier... In the event the input sentence contains specific data, including basic information, the conversational AI server updates the saved information. In some embodiments, information received by the conversational AI server is stored using the episodic process of 1923 in an individual memory associated with the agent.", Paragraph 42, "In some embodiments, a scalable artificial intelligence (AI) framework is achieved by using a machine learning model. For example, a machine learning model is used to determine the actions of each autonomous AI character based on the individual character's input, beliefs, and/or goals.", Paragraph 43, "For example, an autonomous AI non-player character (NPC) receives visual and audio input such as detecting movement of other objects in the game environment and/or recognizing voice input from other users... using a machine learning model such as a deep convolutional neural network (DCNN), a solution... is determined based on the sensory inputs, the beliefs, and the goals of the autonomous AI NPC.", Paragraph 102, "In some embodiments, solver 1021 utilizes a light-weight deep neural network (DNN) planner. In some embodiments, solver 1021 utilizes a deep convolutional neural network (DCNN).", Paragraph 105, "In some embodiments, changes to in-game environment 1031, including auditory and visual changes detected by auditory and visual sensor 1035, relevant to an AI character, are encoded into perception input 1033 and used as input to AI game component 1011 including input to AI loop 1017."
The agent identifier that the system appends to the natural language speech segments is the generated prompt. It is associated with the specific agent which is the productivity object. The server uses the identifier to select and update that agent's individual memory, which becomes the belief state the model reasons over.)
determining, based on the generated prompt and the received user input, a model output of the … model (See Figures 10 and 19, Paragraph 169, "In some embodiments, the planning performed at 1911 uses a lightweight AI framework and is trained as described by FIGS. 1-6 and used to solve AI planning problems as described by FIGS. 7 and 8.", Paragraph 35 of Kovacs, "The AI behavior modeled by the DNN is used to infer AI actions that are mapped to game actions.", Paragraph 101, "Solver 1021 is used to generate a solution, such as solution plan 1023, to the provided AI planning problem, such as initial planning problem 1019.", Paragraph 102, "In some embodiments, solver 1021 utilizes a light-weight deep neural network (DNN) planner."
The deep neural network solver 1021 produces the solution plan, which is the model output. The generated prompt is the natural-language sentence the user directs to the agent, and the received user input is the perception input reflecting the user's control of the game since the actions of the human player character are controlled by a human user. Both are provided to the model and are used to update the belief set which defines the AI planning problem the model solves. The deep neural network solver then produces the solution plan, which is the model output.);
wherein the model output does not include content to include in the document; (Paragraph 104, "During each iteration of artificial intelligence (AI) loop 1017, the first action of solution plan 1023 is extracted and encoded using first action extractor and encoder 1025. The extracted action is executed using action executor 1027.", Paragraph 88, "In some embodiments, the output of the DNN may be a vector of floating point numbers, such as doubles, between 0.0 and 1.0. The action is selected based on the maximal output element of the DNN output vector… For example, a grounded action can be: MOVE FROM-HOME TO-WORK where MOVE is the action-name and FROM-HOME and TO-WORK are the values and/or the two parameters of the action. When executed, an agent such as a non-player character (NPC) moves from home to work.", Paragraph 260, "The encoded AI action provided by the deep neural network (DNN) at 2607 is decoded into a proper, parameterized AI action. In some embodiments, the AI action is represented as an instance of an AI action object."
The DNN's output is an encoded vector that is decoded into a parameterized action instance such as MOVE FROM-HOME TO-WORK.);
processing the model output to control functionality of the productivity application; (See Figures 10 and 19, Paragraph 104, "The extracted action is executed using action executor 1027. For example, the action may include a movement action, a speaking action, an action to end a conversation, etc.", Paragraph 248, "The implementation of the ReceiveActionFromAI function calls the game code (depending on the game logic).");
and as a result of processing the model output, modifying the document. (Paragraph 158, "In some embodiments, these changes also influence the flow of the game, which causes a belief update and subsequent changes to the game environment.", Paragraph 114, "In some embodiments, the action of an NPC agent can influence the objects in the game environment. For example, an NPC can move an object, carry an object, and/or consume an object, etc.").
Kovacs does not teach generating a prompt for priming a … model.
Keskar, in the same field of endeavor, teaches generating a prompt for priming the … model (Page 1 Introduction, “…we train a language model that is conditioned on a variety of control codes … that make desired features of generated text more explicit. With 1.63 billion parameters, our Conditional Transformer Language (CTRL) model can generate text conditioned on control codes that specify domain, style, topics, dates, entities, relationships between entities, plot points, and task-related behavior.”, Page 2 Section 3, “CTRL is a conditional language model that is always conditioned on a control code c and learns the distribution p(x|c). The distribution can still be decomposed using the chain rule of probability and trained with a loss that takes the control code into account.”
The control code that the system generates and adds to the input is the generated prompt, and the model is conditioned on that code to learn p(x|c), which is the priming.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Kovacs’ teaching of updating the game environments based on user input with Keskar’s control code conditioning in order to give explicit and predictable control over the model's output (Introduction of Keskar).
Kovacs in view of Keskar does not teach generating a prompt for … a multimodal machine learning model.
Analogous art Singh teaches the multimodal machine learning model (Claim 1 of Singh, “…accessing, by the multimodal query subsystem, a multimodal question-answering model comprising (a) a first stream of language models comprising a first set of transformer-based models concatenated with each other and (b) a second stream of language models comprising a second set of transformer-based models concatenated with each other, wherein each transformer-based model comprises a respective cross-attention layer using data generated by both the first stream of language models and the second stream of language models as input;”, Claim 6, “…the multimodal question-answering model comprises a first stream of language models configured for text content and a second stream of language models configured for image content”
Singh's multimodal question-answering model processes multiple content types (text content via the first stream of transformer-based language models and image content via the second stream) and interoperates between those content types through cross-attention layers that pass data between the two streams.)
receiving, at a productivity application, user input associated with a document (Paragraph 14, “A modality of a document refers to a type of content in the document, such as text, image, chart, audio, or video… is considered when generating answers to a query. Certain embodiments described herein address these limitations by generating and training a multimodal query-answer model to generate answers to queries by taking into account multiple modalities of the source documents.”, Paragraph 17, “For each of these queries, the multimodal computing system extracts the images in the document that contains the answer to the query and calculates a relevance score of each image to the query. ”, Paragraph 62, “Although the above description focuses on English query-answer application, the modality adaptive knowledge retrieval presented herein applies to any language as long as the training datasets are in the proper language.”)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Kovacs in view of Keskar’s teachings with Singh's multimodal question-answering model in order to enable the model to process and interoperate across multiple content types such as text and image content to improve tools for processing text (Paragraph 18 and claims 1 and 6 of Singh).
Regarding claim 10,
Kovacs teaches the model output is a first model output and the method further comprises: storing the received user input and the determined model output as part of a context (Figure 19, Paragraph 165 of Kovacs, "In some embodiments, the input is stored using the episodic process of 1923 in a memory of an agent such as a non-player character (NPC) agent.", Paragraph 103 of Kovacs, "In some embodiments, a solution plan repository is implemented using a memory graph data structure and stores at least the problem, solution, and context for each solved problem.", Paragraph 106 of Kovacs, "In various embodiments, conversational artificial intelligence (AI) server 1051 is utilized by a computer-based game to store, process, and/or transmit conversations between game entities...");
receiving a second user input (Paragraph 106 of Kovacs, "Using at least conversational AI server 1051, a player of game device 1001 can initiate and even interrupt a conversation with one or more autonomous AI NPCs.", Paragraph 169, "In some embodiments, the flow described by FIG. 19 is performed in part by AI loop 1017 of FIG. 10...", Paragraph 101, "In some embodiments, artificial intelligence (AI) loop 1017 is run until a goal is achieved...");
determining, based on the second user input and the context, a second model output associated with the multimodal machine learning model (Paragraph 101 of Kovacs, "On subsequent iterations, changes to the game environment may be reflected by changing the problem and updating the data structures corresponding to the AI problem. For example, after an initial iteration of AI loop 1017, the problem data structure is revised to address changes to the game that influence the current problem. For each iteration of AI loop 1017, solver 1021 solves the revised problem (reflected by the changes in the problem data structures) to find a best action list for the current state.", Paragraph 103 of Kovacs, "In various embodiments, solver 1021 may retrieve a previously saved solution from the solution plan repository when encountering a previously solved problem.", Paragraph 169 of Kovacs, "In some embodiments, the flow described by FIG. 19 is performed in part by AI loop 1017 of FIG. 10...");
and processing the second model output to modify the document. (Figure 10, Paragraph 104, "In various embodiments, the actions performed by NPCs influence the game environment and in turn the environment also influences the NPCs during the next iteration of AI loop 1017.").
Claim 11 recites similar limitations to claim 5. Therefore, claim 11 is rejected using the same rationale as claim 5.
Regarding claim 12,
Kovacs teaches the model output includes a set of programmatic steps associated with the functionality of the productivity application (Paragraph 65 of Kovacs, "In various embodiments, each solution plan is created and stored as a solution plan file. In some embodiments, the solution plan includes action plans.", Paragraph 101, "In various embodiments, solver 1021 generates a plan and corresponding actions to achieve one or more goals defined by the problem.", Paragraph 114, "In various embodiments, the result of an agent reasoner such as agent reasoners 1141, 1143, and 1149 is a list of actions (not shown). In some embodiments, the list of actions is stored in an action pool (not shown), which is used by runners (not shown) to run the action."
The solution plan output by the DNN is an ordered action list. Each action in that list maps through the AI API to a sequence of game actions that are implemented in the application's own code. The action list corresponds to the set of programmatic steps associated with the functionality of the productivity application.);
and processing the model output comprises executing the set of programmatic steps to control functionality of the productivity application (Paragraph 36, "In some embodiments, a game developer uses an AI application programming interface (API) of the AI tool to interface with the trained DNN. The API may be used to convert inference results such as AI actions into game actions.", Paragraph 37, "In various embodiments, the AI tool is integrated into a development environment such as an application or software development environment.", Paragraph 248, "In the event an AI action is scheduled for execution by the AI of an agent, the associated game actions are executed. In some embodiments, the game actions are organized as a sequence of game actions.", Paragraph 158, "In some embodiments, these changes also influence the flow of the game, which causes a belief update and subsequent changes to the game environment."
The action executor 1027 executes the program steps to control that functionality.).
Regarding claim 13,
Kovacs teaches [a] method for controlling a productivity application using a … machine learning model, the method comprising (Paragraph 33, "The invention can be implemented in numerous ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer readable storage medium; and/or a processor, such as a processor configured to execute instructions stored on and/or provided by a memory coupled to the processor.");
receiving, at a productivity application, a natural language user input; (See Figures 10 and 19, Paragraph 162, "At 1905, natural language (NL) input is received. For example, the output of speech recognition at 1903 is a natural language input.", Paragraph 106, "Using at least conversational AI server 1051, a player of game device 1001 can initiate and even interrupt a conversation with one or more autonomous AI NPCs."
The game application is executed on a game device which receives a spoken/typed sentence from the player, converts it to natural language input at element 1905, and passes it into the application's AI logic. Kovacs's game application is a productivity application because it is an application in which a user provides natural-language input to a conversational agent to control the application and act on its objects. The player directs input to an NPC agent, and the agent's resulting action modifies an object in the game environment. The received sentence corresponds to the natural language user input, and the game application executing on game device is the receiving application.);
generating a prompt for … the … model, wherein the prompt is associated with at least one of the productivity application or a document of the productivity application (Paragraph 163 of Kovacs, "In some embodiments, natural language speech segments are sent to the conversational AI server of 1921 with the identity (e.g., name or user identifier, etc.) of the agent.", Paragraph 165, "In some embodiments, the natural language input sentences are received as speech segments and include an agent identifier... In the event the input sentence contains specific data, including basic information, the conversational AI server updates the saved information. In some embodiments, information received by the conversational AI server is stored using the episodic process of 1923 in an individual memory associated with the agent.", Paragraph 42, "In some embodiments, a scalable artificial intelligence (AI) framework is achieved by using a machine learning model. For example, a machine learning model is used to determine the actions of each autonomous AI character based on the individual character's input, beliefs, and/or goals.", Paragraph 43, "For example, an autonomous AI non-player character (NPC) receives visual and audio input such as detecting movement of other objects in the game environment and/or recognizing voice input from other users... using a machine learning model such as a deep convolutional neural network (DCNN), a solution... is determined based on the sensory inputs, the beliefs, and the goals of the autonomous AI NPC.", Paragraph 102, "In some embodiments, solver 1021 utilizes a light-weight deep neural network (DNN) planner. In some embodiments, solver 1021 utilizes a deep convolutional neural network (DCNN).", Paragraph 105, "In some embodiments, changes to in-game environment 1031, including auditory and visual changes detected by auditory and visual sensor 1035, relevant to an AI character, are encoded into perception input 1033 and used as input to AI game component 1011 including input to AI loop 1017."
The agent identifier that the system appends to the natural language speech segments is the generated prompt. It is associated with the specific agent which is the productivity object. The server uses the identifier to select and update that agent's individual memory, which becomes the belief state the model reasons over.)
determining, based on the generated prompt and the received user input, a model output associated with the … model; (Figures 10 and 19, Paragraph 169, "In some embodiments, the planning performed at 1911 uses a lightweight AI framework and is trained as described by FIGS. 1-6...", Paragraph 35, "The AI behavior modeled by the DNN is used to infer AI actions that are mapped to game actions.", Paragraph 102, "In some embodiments, solver 1021 utilizes a light-weight deep neural network (DNN) planner."
The deep neural network solver 1021 produces the solution plan, which is the model output. The generated prompt is the natural-language sentence the user directs to the agent, and the received user input is the perception input reflecting the user's control of the game since the actions of the human player character are controlled by a human user. Both are provided to the model and are used to update the belief set which defines the AI planning problem the model solves. The deep neural network solver then produces the solution plan, which is the model output);
processing the model output to control functionality of the productivity application; and (Figures 10 and 19, Paragraph 104 of Kovacs, "The extracted action is executed using action executor 1027. For example, the action may include a movement action, a speaking action, an action to end a conversation, etc."
The solution plan output by the DNN is consumed by action executor, which invokes the application's own functions);
as a result of processing the model output, modifying a document of the productivity application. (Paragraph 158 of Kovacs, "In some embodiments, these changes also influence the flow of the game, which causes a belief update and subsequent changes to the game environment."
Executing the action selected by the DNN changes the state of an object maintained by the application. The object whose state is changed is the productivity object).
Kovacs does not teach generating a prompt for priming a … model.
Keskar, in the same field of endeavor, teaches generating a prompt for priming the … model (Page 1 Introduction, “…we train a language model that is conditioned on a variety of control codes … that make desired features of generated text more explicit. With 1.63 billion parameters, our Conditional Transformer Language (CTRL) model can generate text conditioned on control codes that specify domain, style, topics, dates, entities, relationships between entities, plot points, and task-related behavior.”, Page 2 Section 3, “CTRL is a conditional language model that is always conditioned on a control code c and learns the distribution p(x|c). The distribution can still be decomposed using the chain rule of probability and trained with a loss that takes the control code into account.”
The control code that the system generates and adds to the input is the generated prompt, and the model is conditioned on that code to learn p(x|c), which is the priming.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Kovacs’ teaching of updating the game environments based on user input with Keskar’s control code conditioning in order to give explicit and predictable control over the model's output (Introduction of Keskar).
Kovacs in view of Keskar does not teach generating a prompt for … a multimodal machine learning model.
Analogous art Singh teaches the multimodal machine learning model (Claim 1 of Singh, “…accessing, by the multimodal query subsystem, a multimodal question-answering model comprising (a) a first stream of language models comprising a first set of transformer-based models concatenated with each other and (b) a second stream of language models comprising a second set of transformer-based models concatenated with each other, wherein each transformer-based model comprises a respective cross-attention layer using data generated by both the first stream of language models and the second stream of language models as input;”, Claim 6, “…the multimodal question-answering model comprises a first stream of language models configured for text content and a second stream of language models configured for image content”
Singh's multimodal question-answering model processes multiple content types (text content via the first stream of transformer-based language models and image content via the second stream) and interoperates between those content types through cross-attention layers that pass data between the two streams.)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Kovacs in view of Keskar’s teachings with Singh's multimodal question-answering model in order to enable the model to process and interoperate across multiple content types such as text and image content to improve tools for processing text (Paragraph 18 and claims 1 and 6 of Singh).
Claim 14 recites similar limitations to claim 2. Therefore, claim 14 is rejected using the same rationale as claim 2.
Claim 15 recites similar limitations to claim 3. Therefore, claim 15 is rejected using the same rationale as claim 3.
Claim 16 recites similar limitations to claim 4. Therefore, claim 16 is rejected using the same rationale as claim 4.
Claim 17 recites similar limitations to claim 5. Therefore, claim 17 is rejected using the same rationale as claim 5.
Claim 20 recites similar limitations to claim 8. Therefore, claim 20 is rejected using the same rationale as claim 8.
Regarding claim 21,
Kovacs teaches the set of programmatic steps comprises a call to an application programming interface (API) of the productivity application (Paragraph 262 of Kovacs, "At 2613, the decoded, parameterized AI action is mapped to one or more game actions. For example, game actions corresponding to the AI action are received via an AI application programming interface (API) and translated into function calls (or similar game code) in the game. The function calls result in game world execution...", Paragraph 248, "In the event an AI action is scheduled for execution by the AI of an agent, the associated game actions are executed... The implementation of the ReceiveActionFromAI function calls the game code (depending on the game logic).", Paragraph 36, "In some embodiments, a game developer uses an AI application programming interface (API) of the AI tool to interface with the trained DNN. The API may be used to convert inference results such as AI actions into game actions."
The programmatic steps is the ordered game actions mapped from the AI action. Each game action is executed by calling into the productivity application's code through the AI API, which corresponds to a call to an API of the productivity application.).
Regarding claim 22,
Kovacs teaches the model output comprises a set of programmatic steps; and modifying the productivity object comprises executing the set of programmatic steps to modify the productivity object of the productivity application (Paragraph 65 of Kovacs, "In various embodiments, each solution plan is created and stored as a solution plan file. In some embodiments, the solution plan includes action plans.", Paragraph 101, "In various embodiments, solver 1021 generates a plan and corresponding actions to achieve one or more goals defined by the problem.", Paragraph 114, "In various embodiments, the result of an agent reasoner such as agent reasoners 1141, 1143, and 1149 is a list of actions (not shown). In some embodiments, the list of actions is stored in an action pool (not shown), which is used by runners (not shown) to run the action... In some embodiments, the action of an NPC agent can influence the objects in the game environment. For example, an NPC can move an object, carry an object, and/or consume an object, etc.", Paragraph 104, "During each iteration of artificial intelligence (AI) loop 1017, the first action of solution plan 1023 is extracted and encoded using first action extractor and encoder 1025. The extracted action is executed using action executor 1027."
The solution plan output by the DNN comprises an ordered action list, which corresponds to the set of programmatic steps. The action executor executes those actions, and executing an action influences and changes the state of the object, which corresponds to modifying the productivity object.).
Claims 7 and 19 are rejected under 35 U.S.C 103 as being unpatentable over Kovacs (US 20190197402 A1) in view of Keskar (“CTRL: A CONDITIONAL TRANSFORMER LANGUAGE
MODEL FOR CONTROLLABLE GENERATION”, 2019), in view of Singh (US 20220230061 A1), and in further view of Liao (US 20210097133 A1).
Regarding claim 7,
Kovacs does not teach the system wherein the model output is further determined based on: a prompt associated with the productivity application; and a context associated with a different productivity application.
Liao, in the same field of endeavor, teaches wherein the model output is further determined based on: a context associated with a different productivity application (Paragraph 37, "In another example, the personalization module 308 determines that a context of the user activities based on the user activities data. For example, the personalization module 308 determines that the context of the user 132 indicates that the user 132 is deeply focused on operating the enterprise application and does not need interruptions. The personalization module 308 determines the context from signals provided by the enterprise application (e.g., user 132 enabled full screen mode of the enterprise application), and the operating system of client device 106 (e.g., user 132 has set the client device 106 or the programmatic client 108 to silence all notifications).", Paragraph 41, "The individual context module 402 determines that a context of the user activities based on the user activities data. For example, the individual context module 402 determines that the user 132 is deeply focused on operating the enterprise application and is avoiding interruptions. For example, the individual context module 402 detects that the user 132 has enabled a full screen feature mode of the enterprise application, and that the user 132 has set the client device 106 or the programmatic client 108 to silence all notifications.", Paragraph 39, "The personalization module 308 issues a request to the personalized proactive trigger module 306 to either trigger a display of the suggestion in a user interface element of the enterprise application 124 or programmatic client 108 based on a combination of the user activities data, the context of the user activities data, and the cohort profile corresponding to the user 132.", Paragraph 32, "The enterprise application 124 includes, for example, a communication application 202 (e.g., Microsoft Outlook™), a content creation application 204 (e.g., Microsoft Word™, Microsoft Power Point™, and a collaboration application 206 (e.g., Microsoft SharePoint™).", Paragraph 55, "At block 804, the individual context module 402 identifies personal activities of the user inside and outside the enterprise application. For example, the individual context module 402 determines that the user has minimized all other applications operating on the client device 106 except for the enterprise application... In another example, the individual context module 402 determines that the user has silenced notifications from the communication application 202.", Paragraph 60, "At block 906, the individual context module 402 identifies user personal activities. For example, the user application activities monitor module 302 determines that the user 132 has both the communication application 202 and the content creation application 204 opened on the client device 106."
Liao's individual context module gathers context signals from applications other than the one in which the suggestion will be surfaced. The notification state of communication application and the communication application is concurrently open alongside with the content creation application 204. The cross-application context is combined with the in-application signals and fed to the learning engine which determines the resulting suggestion.);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Kovacs in view of Keskar and in further view of Singh's teachings with Liao's teaching of conditioning the model output on an application prompt and on context drawn from a different application in order to optimize the model for different scenarios and use cases and to avoid disrupting the user's attention while operating the application (Paragraphs 2 and 19 of Liao).
Regarding claim 19,
Kovacs does not teach the system wherein the model output is further determined based on: a prompt associated with the productivity application; and a context associated with a different productivity application.
Liao, in the same field of endeavor, teaches wherein the model output is further determined based on: a prompt associated with the productivity application (Figures 3 and 4, Paragraph 34 of Liao, "Other examples of user activities data include: a frequency of whether the user has chosen the suggestion offered in the past; the kind of content already present in the user document; the quality of suggestion offered in the pane prompt (as determined by the user); the types of actions the user was taking prior to the pane prompt (to determine whether the user is deeply focused on operating the application—this also referred to as “focus mode”); and the user's frequency of engagement with the product (e.g., the application).", Paragraph 35, "The suggestion module 304 determines a suggestion based on the user activities data and a content being provided by the user 132 in the web client 112 or the programmatic client 108. For example, the user 132 operates the programmatic client 108 to create a slide and provides text content describing a title for a team meeting: “how to increase productivity.” The user application activities monitor module 302 provides the text content to the suggestion module 304. The suggestion module 304 generates a suggestion based on the user activities data and the text content. For example, the suggestion includes an image of a chart showing an upward trend to match the “increase productivity” text content.”, Paragraph 42, "In another example, the individual context module 402 determines a context of the user activities based on the following signals: a frequency of whether the user has chosen the suggestion offered in the past; the kind of content already present in the user document; the quality of suggestion offered in the pane prompt (as determined by the user); the types of actions the user was taking prior to the pane prompt...", Paragraph 43, "The learning engine 406 trains a machine learning model based on the above signals to determine whether to cause a display of a suggestion pop up user interface element of in the enterprise application. In one example embodiments, the learning engine 406 analyzes events in the enterprise application 124 or programmatic client 108 to identify trends... Based on the machine learning model, the learning engine 406 can, in one embodiment, suggest to the personalized proactive trigger module 306 whether to trigger and cause a display of the suggestion or prevent the suggestion from being displayed."
The pane prompt and the text content that the user entered in the enterprise application is fed to the suggestion module and learning engine and conditions the suggestion the model produces.);
and a context associated with a different productivity application (Paragraph 37, "In another example, the personalization module 308 determines that a context of the user activities based on the user activities data. For example, the personalization module 308 determines that the context of the user 132 indicates that the user 132 is deeply focused on operating the enterprise application and does not need interruptions. The personalization module 308 determines the context from signals provided by the enterprise application (e.g., user 132 enabled full screen mode of the enterprise application), and the operating system of client device 106 (e.g., user 132 has set the client device 106 or the programmatic client 108 to silence all notifications).", Paragraph 41, "The individual context module 402 determines that a context of the user activities based on the user activities data. For example, the individual context module 402 determines that the user 132 is deeply focused on operating the enterprise application and is avoiding interruptions. For example, the individual context module 402 detects that the user 132 has enabled a full screen feature mode of the enterprise application, and that the user 132 has set the client device 106 or the programmatic client 108 to silence all notifications.", Paragraph 39, "The personalization module 308 issues a request to the personalized proactive trigger module 306 to either trigger a display of the suggestion in a user interface element of the enterprise application 124 or programmatic client 108 based on a combination of the user activities data, the context of the user activities data, and the cohort profile corresponding to the user 132.", Paragraph 32, "The enterprise application 124 includes, for example, a communication application 202 (e.g., Microsoft Outlook™), a content creation application 204 (e.g., Microsoft Word™, Microsoft Power Point™, and a collaboration application 206 (e.g., Microsoft SharePoint™).", Paragraph 55, "At block 804, the individual context module 402 identifies personal activities of the user inside and outside the enterprise application. For example, the individual context module 402 determines that the user has minimized all other applications operating on the client device 106 except for the enterprise application... In another example, the individual context module 402 determines that the user has silenced notifications from the communication application 202.", Paragraph 60, "At block 906, the individual context module 402 identifies user personal activities. For example, the user application activities monitor module 302 determines that the user 132 has both the communication application 202 and the content creation application 204 opened on the client device 106."
Liao's individual context module gathers context signals from applications other than the one in which the suggestion will be surfaced. The notification state of communication application and the communication application is concurrently open alongside with the content creation application 204. The cross-application context is combined with the in-application signals and fed to the learning engine which determines the resulting suggestion.);
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to combine Kovacs in view of Keskar and in further view of Singh's teachings with Liao's teaching of conditioning the model output on an application prompt and on context drawn from a different application in order to optimize the model for different scenarios and use cases and to avoid disrupting the user's attention while operating the application (Paragraphs 2 and 19 of Liao).
Claim 19 recites similar limitations to claim 7. Therefore, claim 19 is rejected using the same rationale as claim 7.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAJD MAHER HADDAD whose telephone number is (571)272-2265. The examiner can normally be reached Mon-Friday 8-5 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar, can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.M.H./Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125