Prosecution Insights
Last updated: September 17, 2026
Application No. 18/828,450

MULTI-MODAL DEVELOPMENT INTERFACE FOR LARGE LANGUAGE MODEL APPLICATIONS

Non-Final OA §101§102§103§112
Filed
Sep 09, 2024
Priority
May 20, 2024 — provisional 63/649,642
Examiner
O'CONNOR-EMANUEL, LAWRENCE SCOTT
Art Unit
Tech Center
Assignee
Gopal Datt Joshi
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
6 currently pending
Career history
10
Total Applications
across all art units

Statute-Specific Performance

§101
32.0%
-8.0% vs TC avg
§103
41.3%
+1.3% vs TC avg
§102
10.7%
-29.3% vs TC avg
§112
14.7%
-25.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §102 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This Office Action is in response to the application filed on 09/09/2024. Claims 1-20 are pending in this application. Claims 1, 8,12 and 16 are independent claims. Claim Objections Claim 16 objected to because of the following informalities: The claim recites the term “LLM” in the first line. The examiner suggests that the applicant spell the term fully as done in independent claims 1 and 8 as “large language model (LLM)”. Appropriate correction is required. Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: element 728 tangible storage devices at [Spec 0073]. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: Fig.7 elements 726 and 730 . Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. The drawings are objected to because: Fig. 2, Element 208 is labeled “Modified LLM Inputs”, however the specification refers to Element 208 as “modified inputs”. Fig. 3, Element 312 is labeled “Continuous Output”, however the specification refers to Element 312 as “continuation output”. Fig. 4, Element 426 is labeled “User LLM Input Editor”, however the specification refers to Element 426 as “User LLM Editor”. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: multi-modal user input interface in claims 1 and 8 user input encoder in claims 1 and 8 user review interface in claims 1 and 8 LLM interface in claims 1 and 8 LLM engine in claim 1 recommendation module in claim 4 handling module in claim 5 output module in claim 6 Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim limitations: • multi-modal user input interface in claims 1 and 8 • user input encoder in claims 1 and 8 • user review interface in claims 1 and 8 • LLM interface in claims 1 and 8 • LLM engine in claim 1 invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The disclosure is devoid of sufficient structure that is explicitly stated as performing the functions in the claims. Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph. Applicant may: (a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph; (b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)). If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either: (a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or (b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181. Similarly, claims 2-7 and 9-11 are also rejected as being dependent on the rejected base claims 1 and 8 above. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention recites a judicial exception, is directed to that judicial exception, an abstract idea, it has not been integrated into practical application, and the claims further do not recite significantly more than the judicial exception. Examiner has evaluated the claims under the framework provided in the 2019 Patent Eligibility Guidance published in the Federal Register 01/07/2019 and has provided such analysis below. Regarding claim 1, the limitations, “to encode the acquired multi-modal inputs and to generate LLM inputs for the LLM engine”, and “and ”to modify the generated LLM inputs based upon user review inputs” as drafted, are functions that, under its broadest reasonable interpretation, recite the abstract idea of a mental process. These limitations encompass a human mind carrying out these functions through observation, evaluation judgment and /or opinion, or even with the aid of pen and paper. For example, the “encoding” limitation can be carried out by a user listening to audio and writing down the corresponding text. Thus, these limitations recite and fall within the “Mental Processes” grouping of abstract ideas under Prong 1. Under Prong 2, the judicial exception is not integrated into a practical application. The additional elements, “to acquire a plurality of multi-modal inputs from a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs”, ”to present the generated LLM inputs to the user” and “ to provide the modified inputs to the LLM engine” do nothing more than add insignificant extra solution activity to the judicial exception of mere data gathering and outputting. The additional element “to process the modified inputs to generate a desired output”, merely recites instructions to implement an abstract idea on a generic computer, or merely uses a generic computer or computer components as a tool to perform the abstract idea, thus is not a practical application under Prong 2. Accordingly, the additional elements do not integrate the recited judicial exception into a practical application, and the claim is therefore directed to the judicial exception. See MPEP 2106.05 (g) and 2106.05 (f) respectively. Under Step 2B, the claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As stated above in prong 2, the additional element, “to process the modified inputs to generate a desired output”, merely recites instructions to implement an abstract idea on a generic computer, or merely uses a generic computer or computer components as a tool to perform the abstract idea. The additional elements, “to acquire a plurality of multi-modal inputs from a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs”, ”to present the generated LLM inputs to the user” and “ to provide the modified inputs to the LLM engine”, is merely gathering/outputting data which the courts have identified as well-understood, routine conventional activity. Claim 1, further recites, “for a large language model (LLM) engine, wherein the multi-modal development interface system”, “a multi-modal user input interface” configured”, “a user input encoder configured”, “a user review interface configured”, “a LLM interface configured”, and “ wherein the LLM engine is configured”. These elements are recited at a high level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer and/or generic computer components. See MPEP 2106.05(f). Therefore, the additional elements recited in claim 1 do not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B; thus, cannot provide an inventive concept. Accordingly, the claims are not patent eligible under 35 USC 101. Regarding claim 2, the limitation, “wherein the non-textual inputs comprise inputs acquired via multi-modal interactions of the user with the system and wherein the multi-modal interactions comprise drawing, annotating, gestures, facial expressions, voice notes, video, images, or combinations thereof” , which can be classified as mere data gathering and outputting, does not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B for the reasons provided in the rejection of claim 1. Regarding claim 3, the limitation, “acquire the plurality of multi-modal inputs from a plurality of input sources, wherein the input sources comprise a database, a user-interaction digital device, repository of files, uniform resource locator (URLs), or combinations thereof” , which can be classified as mere data gathering and outputting, does not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B for the reasons provided in the rejection of claim 1. Regarding claim 4, the limitation “configured to provide one or more recommendations to modify the generated LLM inputs” , recites additional mental process under Prong 1. Regarding claim 5, the limitation “detect errors/faults in the inputs to the LLM”, recites additional mental process under Prong 1. The additional element, “ generate warning messages upon such detection”, can be classified as mere data gathering and outputting which does not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B for the reasons provided in the rejection of claim 1. Regarding claim 6, the limitation, “configured to present an exportable output from the LLM to the user”, which can be classified as mere data gathering and outputting, does not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B for the reasons provided in the rejection of claim 1. Claims 4-6, further recite, “recommendation module” in claim 4, “handling module” in claim 5, and “output module” in claim 6. These elements are recited at a high level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer and/or generic computer components. See MPEP 2106.05(f). Therefore, the additional elements recited in claims 4-6 do not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B; thus, cannot provide an inventive concept. Accordingly, the claims are not patent eligible under 35 USC 101. Regarding claim 7, the limitation, “process one or more of documents (text), images, video, URLs, and audio to generate inputs for the LLM engine”, merely recites instructions to implement an abstract idea on a generic computer, or merely uses a generic computer or computer components as a tool to perform the abstract idea; which does not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B for the reasons provided in the rejection of claim 1. Regarding claim sets 8-11, 12-15 and 16-19, the limitations recited in the claims are similar to those of claims 1-7 and thus are rejected for similar reasons as stated in the rejection of claims 1-7 above. Claim 8, further recites, “plurality of interconnected agents”, “an application” and “a scheduler” in claim 8. These elements are recited at a high level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer and/or generic computer components. See MPEP 2106.05(f). Therefore, the additional elements recited in claim 8 do not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B. Claim 12, further recites, “a memory storing one or more processor-executable routines”, and “a processor” in claim 12. These elements are recited at a high level of generality such that they amount to no more than mere instructions to apply the exception using a generic computer and/or generic computer components. See MPEP 2106.05(f). Therefore, the additional elements recited in claim 12 do not integrate the judicial exception into a practical application under Prong 2, nor amount to significantly more under Step 2B; thus, cannot provide an inventive concept. Accordingly, the claims are not patent eligible under 35 USC 101. Regarding claim 20, the limitation “wherein modifying the generated LLM inputs comprises substantially aligning the generated LLM inputs with user intent”, recites additional mental process under Prong 1. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim 1-3, 6-7, and 12-20 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by (US 20250094132 A1) hereafter Cappetta. Regarding claim 1, Cappetta teaches: A multi-modal development interface system for a large language model (LLM) engine, wherein the multi-modal development interface system comprises: a multi-modal user input interface configured to acquire a plurality of multi-modal inputs from a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs; ([0034] ”Initially, in subprocess 205, a natural-language user request may be received. A natural-language user request is a request to generate an integration process 170 that is expressed in natural language….Subprocess 205 may comprise the user typing the user request into a textbox of graphical user interface 150. Alternatively, or additionally, subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”) The graphical user interface 150 accepts input via text and voice which corresponds to, a multi-modal user input interface configured to acquire a plurality of multi-modal inputs from a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs. a user input encoder configured to encode the acquired multi-modal inputs ([0101] “Baseband system 360 also receives analog audio signals from a microphone. These analog audio signals are converted to digital signals and encoded by baseband system 360.”) The Baseband system 360 encoding the audio signals/inputs received from the microphone corresponds to, a user input encoder configured to encode the acquired multi-modal inputs. and to generate LLM inputs for the LLM engine; ([0038] “In subprocess 210, the original user request, received in subprocess 205, may be appended to or otherwise combined with the contextual wrapper, or a user request that is derived from the original user request (e.g., modified to fit a specific format) may be appended to or otherwise combined with the contextual wrapper. In an embodiment, the resulting request-level prompt represents a request to provide a set of objectives for generating the integration process 170. The request-level prompt may be designed to steer a generative AI model towards a specific set of contextual formats and objectives required to construct the integration process 170.” [0039]” In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM).”) The request-level prompt generated based on the users input which is then used as input to the LLM corresponds to, generate LLM inputs for the LLM engine. and a user review interface configured to present the generated LLM inputs to the user and to modify the generated LLM inputs based upon user review inputs; ([0084]” In an embodiment, during execution of process 200, feedback may be provided to the user via graphical user interface 150 or an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like.” ) Graphical user interface 150 provides the generated prompts and allows the user to modify them which corresponds to, a user review interface configured to present the generated LLM inputs to the user and to modify the generated LLM inputs based upon user review inputs. a LLM interface configured to provide the modified inputs to the LLM engine, wherein the LLM engine is configured to process the modified inputs to generate a desired output. ([0039] “In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM). [0084] “In an embodiment, during execution of process 200, feedback may be provided to the user via graphical user interface 150 or an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like.”) See Fig 2, steps 210 and 215. The graphical user interface 150 allows the user to modify the generated prompts at step 210 prior to step 215 which inputs the modified prompts to the LLM to produce a set of objectives. Furthermore, user interface 150, may be used for every stage of process 200 including step 215, which corresponds to a LLM interface configured to provide the modified inputs to the LLM engine, wherein the LLM engine is configured to process the modified inputs to generate a desired output. Regarding claim 2, Cappetta teaches the system of claim 1 above, Cappetta further teaches: wherein the non-textual inputs comprise inputs acquired via multi-modal interactions of the user with the system and wherein the multi-modal interactions comprise drawing, annotating, gestures, facial expressions, voice notes, video, images, or combinations thereof.([0034] “subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”)See Fig. 2 subprocess 205 for receiving the natural-language user request. The user submitting the request via a microphone of the system corresponds to, wherein the non-textual inputs comprise inputs acquired via multi-modal interactions of the user with the system and wherein the multi-modal interactions comprise [] voice notes [] or combinations thereof. Regarding claim 3, Cappetta teaches the system of claim 1 above, Cappetta further teaches: wherein the multi-modal user input interface is configured to acquire the plurality of multi-modal inputs from a plurality of input sources, wherein the input sources comprise a database, a user-interaction digital device, repository of files, uniform resource locator (URLs), or combinations thereof.([0026] “User system(s) 130 may comprise any type or types of computing devices capable of wired and/or wireless communication, including without limitation, desktop computers, laptop computers, tablet computers, smart phones or other mobile phones, servers, game consoles, televisions, set-top boxes, electronic kiosks, point-of-sale terminals, and/or the like. [0030] “It should be understood that multiple users, on multiple user systems 130, may manage the same integration process(es) 170 and/or different integration processes 170 in this manner, according to the permissions or roles of their associated user accounts.” [0034]”in subprocess 205, a natural-language user request may be received… to generate an integration process 170 … subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”) The plurality of computing devices may be used to provide the user request inputs from multiple users to the system which corresponds to, wherein the multi-modal user input interface is configured to acquire the plurality of multi-modal inputs from a plurality of input sources, wherein the input sources comprise [] a user-interaction digital device, [] or combinations thereof. Regarding claim 6, Cappetta teaches the system of claim 1 above, Cappetta further teaches: wherein the system further comprises an output module configured to present an exportable output from the LLM to the user. ([0081] ”In subprocess 270, the integration process 170 may be generated from the process definition output by the most recent iteration of subprocess 260. For example, the process definition may be provided to the application programming interface of a process-generation service of server application 112 or another application… the process-generation service may generate a software instance of the integration process 170, according to the process definition, and return a reference to the software instance of the integration process 170. For example, the reference may comprise a unique process identifier for the software instance of the integration process 170, a URI of the software instance of the integration process 170, and/or the like. Alternatively, the process-generation service could return a data structure representing the software instance itself.”)The output comprising a software instance corresponds to an exportable output because the user can use the file/reference to deploy the integration process. Regarding claim 7, Cappetta teaches the system of claim 1 above, Cappetta further teaches: wherein the user input encoder is configured to process one or more of documents (text), images, video, URLs, and audio to generate inputs for the LLM engine.([0034-0040] “subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112. … user system 130 or server application 112 may convert the user's speech into text, representing the user's request, …In subprocess 210, a request-level prompt is generated based on the user request…. In subprocess 215, the request-level prompt, output by subprocess 210, is input into a … a large language model (LLM)” [0101] “Baseband system 360 also receives analog audio signals from a microphone. These analog audio signals are converted to digital signals and encoded by baseband system 360.”) The Baseband system 360 encoding the audio signals/inputs received from the microphone, which are used to generate the request-level prompt that is input to the LLM corresponds to, wherein the user input encoder is configured to process [] audio to generate inputs for the LLM engine. Regarding claim 12, Cappetta teaches: A multi-modal development interface system for a large language model (LLM) engine, wherein the multi-modal development interface system comprises: a memory storing one or more processor-executable routines; and a processor communicatively coupled to the memory, the processor configured to execute the one or more processor-executable routines to: ([0088-0090] “System 300 may comprise one or more processors 310. … System 300 may comprise main memory 315. Main memory 315 provides storage of instructions and data for programs executing on processor 310, such as any of the software discussed herein.”) System 300 comprises an equivalent memory and processor. receive a plurality of multi-modal inputs from a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs; ([0034] ”Initially, in subprocess 205, a natural-language user request may be received. A natural-language user request is a request to generate an integration process 170 that is expressed in natural language….Subprocess 205 may comprise the user typing the user request into a textbox of graphical user interface 150. Alternatively, or additionally, subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”) The graphical user interface 150 accepts input via text and voice which corresponding to, receive a plurality of multi-modal inputs from a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs; process the acquired multi-modal inputs to generate LLM inputs for the LLM engine; ([0038] “In subprocess 210, the original user request, received in subprocess 205, may be appended to or otherwise combined with the contextual wrapper, or a user request that is derived from the original user request (e.g., modified to fit a specific format) may be appended to or otherwise combined with the contextual wrapper.”[0039]” In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM).”) The original user request is modified to create a request-level prompt which is input to the LLM, which corresponds to, request-level prompt generated based on the users input which is then used as input to the LLM corresponds to, process the acquired multi-modal inputs to generate LLM inputs for the LLM engine; receive user review inputs from the user on the generated LLM inputs and modify the generated LLM inputs based upon the received inputs; ([0084]” In an embodiment, during execution of process 200, feedback may be provided to the user via graphical user interface 150 or an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like.” ) Graphical user interface 150 provides the generated prompts and allows the user to modify them which corresponds to, receive user review inputs from the user on the generated LLM inputs and modify the generated LLM inputs based upon the received inputs. provide the modified inputs to the LLM engine, wherein the LLM engine is configured to process the modified inputs to generate a desired output. ([0039] “In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM). [0084] “In an embodiment, during execution of process 200, feedback may be provided to the user via graphical user interface 150 or an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like.”) See Fig 2, steps 210 and 215. The graphical user interface 150 allows the user to modify the generated prompts at step 210 prior to step 215 which inputs the modified prompts to the LLM to produce a set of objectives which corresponds to, provide the modified inputs to the LLM engine, wherein the LLM engine is configured to process the modified inputs to generate a desired output. Regarding claim 13, Cappetta teaches the system of claim 12 above. Cappetta further teaches: wherein the LLM engine is a generative AI based engine, or an autoregressive language model.([Abstract] “Embodiments wrap the user request with a contextual wrapper to generate a prompt, submit the prompt to a generative AI model (e.g., large language model) to produce an actionable set of objectives.”)The generative AI model (e.g., large language model) corresponds to, wherein the LLM engine is a generative AI based engine []. Regarding claim 14, Cappetta teaches the system of claim 12 above. Cappetta further teaches: wherein the processor is configured to process one or more of documents (text), images, video, URLs, and audio to generate inputs for the LLM engine. ([0034-0040] “subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112. … user system 130 or server application 112 may convert the user's speech into text, representing the user's request, …In subprocess 210, a request-level prompt is generated based on the user request…. In subprocess 215, the request-level prompt, output by subprocess 210, is input into a … a large language model (LLM)”) Audio inputs from the user are converted to text, which is used to generate the request-level prompt that is input to the LLM which corresponds to, wherein the processor is configured to process one or more of [] audio to generate inputs for the LLM engine. Regarding claim 15, Cappetta teaches the system of claim 12 above. Cappetta further teaches: wherein the processor is further configured to receive user review inputs via non textual inputs.([0084-0085] “In an embodiment, during execution of process 200, feedback may be provided to the user via … an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like… In summary, process 200 enables a user to engage with a … audio-based chatbot”) Regarding claim 16, Cappetta teaches: A method of generating LLM inputs for or a LLM engine, the method comprising: acquiring a plurality of multi-modal inputs provided by a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs; ([0034] ”Initially, in subprocess 205, a natural-language user request may be received. A natural-language user request is a request to generate an integration process 170 that is expressed in natural language….Subprocess 205 may comprise the user typing the user request into a textbox of graphical user interface 150. Alternatively, or additionally, subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”) The graphical user interface 150 accepts input via text and voice which corresponds to, acquiring a plurality of multi-modal inputs provided by a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs. converting the acquired multi-modal inputs and to generate LLM inputs for the LLM engine; ([0034] “Initially, in subprocess 205, a natural-language user request may be received… subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112. In this latter case, user system 130 or server application 112 may convert the user's speech into text, representing the user's request, using a standard speech-to-text engine. ”[0039]” In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM).”) When the original user request is audio , the system converts the user's speech into text to create a request-level prompt which is input to the LLM, which corresponds to, converting the acquired multi-modal inputs and to generate LLM inputs for the LLM engine. receiving user review inputs on the generated LLM inputs; and modifying the generated LLM inputs based on the user review inputs to generate modified inputs. ([0084]” In an embodiment, during execution of process 200, feedback may be provided to the user via graphical user interface 150 or an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like.” ) Graphical user interface 150 provides the generated prompts and allows the user to modify them which corresponds to, receiving user review inputs on the generated LLM inputs; and modifying the generated LLM inputs based on the user review inputs to generate modified inputs. Regarding claim 17, Cappetta teaches the method of claim 16 above. Cappetta further teaches: wherein the method further comprises processing the generated LLM inputs via the LLM engine and generating a desired output based on the LLM inputs. ([0038-0039] “The request-level prompt may be designed to steer a generative AI model towards a specific set of contextual formats and objectives required to construct the integration process 170… In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM).”) The LLM uses the request level prompt to produce specific set of contextual formats and objectives required to construct the integration process, which corresponds to, wherein the method further comprises processing the generated LLM inputs via the LLM engine and generating a desired output based on the LLM inputs. Regarding claim 18, Cappetta teaches the method of claim 16 above. Cappetta further teaches: wherein the method further comprises acquiring the multi-modal inputs via multi-modal interactions of the user.([0034] ”Subprocess 205 may comprise the user typing the user request into a textbox of graphical user interface 150. Alternatively, or additionally, subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”) The user can provide input to the system via typing or speaking into a microphone, which corresponds to wherein the method further comprises acquiring the multi-modal inputs via multi-modal interactions of the user. Regarding claim 19, Cappetta teaches the method of claim 18 above. Cappetta further teaches: wherein the method further comprises acquiring the multi-modal inputs via drawing, annotating, gestures, facial expressions, voice notes, video, images, or combinations thereof.([0034] “subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”) See Fig. 2 subprocess 205 for receiving the natural-language user request. The user submitting the request via a microphone of the system corresponds to, wherein the method further comprises acquiring the multi-modal inputs [] voice notes [] or combinations thereof. Regarding claim 20, Cappetta teaches the method of claim 16 above. Cappetta further teaches: wherein modifying the generated LLM inputs comprises substantially aligning the generated LLM inputs with user intent. ([0084]” In an embodiment, during execution of process 200, feedback may be provided to the user via graphical user interface 150 or an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like.”) The user may modify the LLM prompt inputs in any manner they desire which corresponds to, wherein modifying the generated LLM inputs comprises substantially aligning the generated LLM inputs with user intent; consistent with applicants’ statements at [Specification 0035]. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 4 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over (US 20250094132 A1) hereafter Cappetta in view of (US 20240256762 A1) hereafter Beauchamp. Regarding claim 4, Cappetta teaches the system of claim 1 above, Cappetta does not explicitly teach: wherein the user review interface further comprises a recommendation module configured to provide one or more recommendations to modify the generated LLM inputs. However, Beauchamp teaches: wherein the user review interface further comprises a recommendation module configured to provide one or more recommendations to modify the generated LLM inputs. ( [0155-00156] “a text-editing UI (e.g., the UI 600) was provided via the user device to enable input of text-editing instruction(s…Consider the example where a text-editing instruction inadvertently included a spelling error (e.g., a user inputs the instruction to replace “fat” with “obees”, where “obees” is a spelling error of “obese”). The LLM may recognize that the text-editing instruction contains a spelling error and thus the generated revised text may include a preamble that includes request for clarification, such as “I think you meant “obese” instead of “obees”.)The LLM outputs a request for clarification that contains a recommended change in response to the user input which corresponds to, wherein the user review interface further comprises a recommendation module configured to provide one or more recommendations to modify the generated LLM inputs. Therefore, it would have been obvious to one of ordinary skill in the art to which said subject matter pertains before the effective filing date of the claimed invention to combine Cappetta with, wherein the user review interface further comprises a recommendation module configured to provide one or more recommendations to modify the generated LLM inputs, as taught by Beauchamp, in order to help the user refine the prompt to address “any errors, inconsistencies or lack of clarity…prior to being provided in a prompt to the LLM”. [Beauchamp, 0156] Regarding claim 5, Cappetta in view of Beauchamp teaches the system of claim 4 above, Cappetta does not explicitly teach: wherein the user review interface further comprises a handling module configured to detect errors/faults in the inputs to the LLM and to generate warning messages upon such detection. However, Beauchamp teaches: wherein the user review interface further comprises a handling module configured to detect errors/faults in the inputs to the LLM and to generate warning messages upon such detection.([0155-0156] “a text-editing UI (e.g., the UI 600) was provided via the user device to enable input of text-editing instruction(s)…Consider the example where a text-editing instruction inadvertently included a spelling error (e.g., a user inputs the instruction to replace “fat” with “obees”, where “obees” is a spelling error of “obese”). The LLM may recognize that the text-editing instruction contains a spelling error and thus the generated revised text may include a preamble that includes request for clarification, such as “I think you meant “obese” instead of “obees”.) The LLM outputs a request for clarification in response to detecting a spelling error in the user input which corresponds to, wherein the user review interface further comprises a handling module configured to detect errors/faults in the inputs to the LLM and to generate warning messages upon such detection. Therefore, it would have been obvious to one of ordinary skill in the art to which said subject matter pertains before the effective filing date of the claimed invention to combine Cappetta with, wherein the user review interface further comprises a handling module configured to detect errors/faults in the inputs to the LLM and to generate warning messages upon such detection, as taught by Beauchamp, in order to help the user refine the prompt to address “any errors, inconsistencies or lack of clarity…prior to being provided in a prompt to the LLM”. [Beauchamp, 0156] Claim 8-11 are rejected under 35 U.S.C. 103 as being unpatentable over (US 20250094132 A1) hereafter Cappetta in view of Enumulapally et al. (US20230325232A1) hereafter Enumulapally. Regarding claim 8, Cappetta teaches: A system of interconnected multi-modal interfaces integrated with a large language model (LLM), wherein the LLM system comprises: a plurality of interconnected agents, each of the plurality of agents configured to receive multi-modal inputs and to process the multi-modal inputs via an LLM engine to produce an output. ([0024-0027] “Platform 110 may execute a server application 112, which may comprise one or more software modules implementing one or more of the disclosed processes… the backend functionality of server application 112 may include a process for constructing integration processes 170 using natural language.”[0034] “ Subprocess 205 may comprise the user typing the user request into a textbox of graphical user interface 150. Alternatively, or additionally, subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112. … Whether the user inputs the user request through a graphical user interface or audio interface, the user may interact with a chatbot service implemented by server application 112.”[0037-0040]” In subprocess 210, a request-level prompt is generated based on the user request… In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM). “) The graphical user interface and audio interface implemented by server application 112 which accept input used to generate the request level prompt input to the LLM corresponds to, a system of interconnected multi-modal interfaces integrated with a large language model (LLM). The functions of server application 112 for producing the integration process (Fig. 2) can be implemented through modules which correspond to agents. The agents as claimed comprise functions that are directly copied from claim 1 and thus fully anticipated by Cappetta above, which corresponds to, wherein the LLM system comprises: a plurality of interconnected agents, each of the plurality of agents configured to receive multi-modal inputs and to process the multi-modal inputs via an LLM engine to produce an output. wherein the plurality of agents are further configured to interact with each other to generate a desired system output ([0038]”In subprocess 210, the original user request, received in subprocess 205, may be appended to or otherwise combined with the contextual wrapper, or a user request that is derived from the original user request (e.g., modified to fit a specific format) may be appended to or otherwise combined with the contextual wrapper.” [0039]” In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM).” [0062-0063]” In subprocess 220, any components, required to complete the set of objectives, are extracted from the set of objectives, output by subprocess 215…a generative AI model is used to extract the components from the set of objectives… In particular, an extraction prompt may be input to the generative AI model to produce a set of components from the set of objectives.”)As stated previously the server application 112 implements the subprocesses of Fig 2 using modules which correspond to agents, thus the passing of data generated by the first AI model in stage 215 to the next AI model at stage 220 which is orchestrated by the modules corresponds to interaction between agents to generate a desired system output. and wherein each of the plurality of interconnected agents further comprises: a multi-modal user input interface configured to acquire the multi-modal inputs from a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs; ([0034]” Initially, in subprocess 205, a natural-language user request may be received. A natural-language user request is a request to generate an integration process 170 that is expressed in natural language... Subprocess 205 may comprise the user typing the user request into a textbox of graphical user interface 150. Alternatively, or additionally, subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”) The graphical user interface 150 accepts input via text and voice which corresponds to, a multi-modal user input interface configured to acquire the multi-modal inputs from a user, wherein the multi-modal inputs comprise textual and/or non-textual inputs; a user input encoder configured to encode the acquired multi-modal inputs ([0101] “Baseband system 360 also receives analog audio signals from a microphone. These analog audio signals are converted to digital signals and encoded by baseband system 360.”) The Baseband system 360 encoding the audio signals/inputs received from the microphone corresponds to, a user input encoder configured to encode the acquired multi-modal inputs . and to generate LLM inputs for the respective LLM engine []; ([0038] “In subprocess 210, the original user request, received in subprocess 205, may be appended to or otherwise combined with the contextual wrapper, or a user request that is derived from the original user request (e.g., modified to fit a specific format) may be appended to or otherwise combined with the contextual wrapper. In an embodiment, the resulting request-level prompt represents a request to provide a set of objectives for generating the integration process 170. The request-level prompt may be designed to steer a generative AI model towards a specific set of contextual formats and objectives required to construct the integration process 170.” [0039]” In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM).”) The request-level prompt generated based on the users input which is then used as input to the LLM corresponds to, and to generate LLM inputs for the respective LLM engine []; and a user review interface configured to present the generated LLM inputs to the user and to modify inputs based upon user review inputs; ([0084]” In an embodiment, during execution of process 200, feedback may be provided to the user via graphical user interface 150 or an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like.” ) Graphical user interface 150 provides the generated prompts and allows the user to modify them which corresponds to, a user review interface configured to present the generated LLM inputs to the user and to modify the generated LLM inputs based upon user review inputs. and a LLM interface configured to provide the modified inputs to the LLM engine, wherein the LLM engine is configured to process the modified inputs to generate the respective output; ([0039] “In subprocess 215, the request-level prompt, output by subprocess 210, is input into a generative AI model to produce the requested set of objectives. In an embodiment, the model comprises a large language model (LLM). [0084] “In an embodiment, during execution of process 200, feedback may be provided to the user via graphical user interface 150 or an audio interface between the user and server application 112. Thus, the user may monitor each subprocess in process 200, for example, to review the prompts that are generated, the outputs of the generative AI model, and/or the like. The user may also be provided with one or more inputs to pause or otherwise disrupt process 200, modify the generated prompts, modify the outputs of the generative AI model, and/or the like.”) See Fig 2, steps 210 and 215. The graphical user interface 150 allows the user to modify the generated prompts at step 210 prior to step 215 which inputs the modified prompts to the LLM to produce a set of objectives. Furthermore, user interface 150, may be used for every stage of process 200 including step 215, which corresponds to, a LLM interface configured to provide the modified inputs to the LLM engine, wherein the LLM engine is configured to process the modified inputs to generate the respective output; an application configured to receive the system output resulting from the interactions of the plurality of interconnected agents, wherein the application is configured to generate a continuation output [] ([0027] "Server application 112 may manage an integration environment 160. In particular, server application 112 may provide a graphical user interface 150 and backend functionality, including one or more of the processes disclosed herein, to enable users, via user systems 130, to construct, develop, modify, save, delete, test, deploy, undeploy, and/or otherwise manage integration processes 170 within integration environment 160."[0081]"In subprocess 270, the integration process 170 may be generated from the process definition output by the most recent iteration of subprocess 260. For example, the process definition may be provided to the application programming interface of a process-generation service of server application 112. ... In response, the process-generation service may generate a software instance of the integration process 170, according to the process definition, and return a reference to the software instance of the integration process 170." Application 112 receives the final process definition which is the output of the full generation process between the multiple AI models as described above, which corresponds to an application configured to receive the system output resulting from the interactions of the plurality of interconnected agents. The Integration pipeline that is deployed by the application corresponds to, wherein the application is configured to generate a continuation output. Cappetta does not explicitly teach: through a scheduler However, Enumulapally suggest that application deployments can be scheduled ([0051] “Business applications (e.g., 112) are kept synchronized within an integration platform 110 using processes 116 as integration flows that are executed on a real-time or scheduled basis.” ) When combined with the integration process of Cappetta corresponds to deployment through a scheduler. Therefore, it would have been obvious to one of ordinary skill in the art to which said subject matter pertains before the effective filing date of the claimed invention to combine Cappetta with, through a scheduler to allow the user to deploy the integration process on a desired schedule. Regarding claim 9, Cappetta in view of Enumulapally teaches the system of claim 8 above, Cappetta further teaches: wherein the application comprises a computer application configured to achieve a high-order task using the system output. ([0081] ”In subprocess 270, the integration process 170 may be generated from the process definition output by the most recent iteration of subprocess 260. For example, the process definition may be provided to the application programming interface of a process-generation service of server application 112 or another application… the process-generation service may generate a software instance of the integration process 170, according to the process definition, and return a reference to the software instance of the integration process 170. For example, the reference may comprise a unique process identifier for the software instance of the integration process 170, a URI of the software instance of the integration process 170, and/or the like. Alternatively, the process-generation service could return a data structure representing the software instance itself.”) See Figure 2 in its entirety. The original user request undergoes multiple steps of processing using multiple AI models to generate a software instance representative of the user’s initial natural language request. The software instance that is output thus corresponds to a high-order task using the system output; consistent with applicants’ definition at [Specification 0052 & 0061]. Regarding claim 10, Cappetta in view of Enumulapally teaches the system of claim 8 above, Cappetta further teaches: wherein the plurality of interconnected agents are configured to receive the multi-modal inputs from a plurality of data systems, input acquisition interfaces, or combinations thereof. ([0086]”In summary, process 200 enables a user to engage with a screen-based or audio-based chatbot or other service (e.g., provided by server application 112).”)Process 200 encompasses the full integration process pipeline including all AI models and generative processes used to achieve the final output. The initial input to the system may be through a screen-based or audio-based chatbot which corresponds to, wherein the plurality of interconnected agents are configured to receive the multi-modal inputs from a plurality of [] input acquisition interfaces, or combinations thereof. Regarding claim 11, Cappetta teaches: wherein the non-textual inputs comprise inputs acquired via multi-modal interactions of the user with the system and wherein the multi-modal interactions comprise drawing, annotating, gestures, facial expressions, voice notes, video, images, or combinations thereof.([0034] “subprocess 205 may comprise the user speaking into a microphone of user system 130 with an audio interface of server application 112.”) See Fig. 2 subprocess 205 for receiving the natural-language user request. The user submitting the request via a microphone of the system corresponds to, wherein the non-textual inputs comprise inputs acquired via multi-modal interactions of the user with the system and wherein the multi-modal interactions comprise [] voice notes [] or combinations thereof. Prior Art Made of Record The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: (Google Bard and ChatGPT: Battle of the AI Wordsmiths Unleashed! ,Published: 07/3/2024). Bard and ChatGPT are multimodal LLM interfaces. Relevant to claims 1-20. (20250315629 A1) Discloses a LLM user interface and suggestion module to help users refine inputs. Relevant to claims 1,4-6. (20250258996 A1) Discloses a gesture-based user system interacting with an LLM to refine inputs. Relevant to claims 1 and 2. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to LAWRENCE O'CONNOR EMANUEL whose telephone number is (571)272-8975. The examiner can normally be reached M-F 9:00 - 5:00 pm . Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chat Do can be reached at (571) 272-3721. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /L.S.O./Examiner, Art Unit 2193 /Chat C Do/Supervisory Patent Examiner, Art Unit 2193
Read full office action

Prosecution Timeline

Sep 09, 2024
Application Filed
Aug 13, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month