Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim’s 1, 8-10, 18-20 are rejected are rejected under U.S.C. 103 as being unpatentable over Talebirad (NPL: Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents) (“Talebirad”) in view of Hughes (NPL: A Pythonista’s Intro to Semantic Kernel) (“Hughes”)
Regarding Claim 1, Talebirad teaches invoke, using the manager module, the plurality of workers comprising a first worker and a second worker ([Page 4-Section 3.2] teaches a agent dynamically adding additional agents to the system to distribute their tasks, the agent here is being interpreted as a manager module and the additional agents that are being created are being interpreted as a plurality of workers which would include a first and second worker which would teach this limitation)
invoke, using the first worker, a generic large language model ([Page 1: Lines 20-23] teaches the use of multiple LLMs in their multi-agent system with the models they are focusing on being GPT-4 and GPT-3.5-turbo, those models are generic LLMs which teaches the generic LLM, and the agents here are interpreted as a worker which would teach the worker limitation)
and invoke, using the second worker, a secondary large language model, wherein the secondary large language model is configured to process structured data ([Page 1: Lines 20-23] teaches the use of multiple LLMs in their multi-agent system with the models they are focusing on being GPT-4 and GPT-3.5-turbo, it can be reasonably interpreted that those models can process structured data and the agents here are being interpreted as a worker which would teach the worker limitation
Talebirad does not teach teaches A system for processing input data, the system comprising: a memory, a communication interface, and a processor operatively coupled to the memory and the communication interface; an application stored in the memory and executable by the processor, and the application comprising a semantic kernel, a manager module, and a plurality of workers;
the processor configured to: receive an input via the semantic kernel;
However, Hughes does teach A system for processing input data, the system comprising: a memory, a communication interface, and a processor operatively coupled to the memory and the communication interface; an application stored in the memory and executable by the processor, and the application comprising a semantic kernel, a manager module, and a plurality of workers; ([Page 4-The kernel] teaches the kernel, [Page 54-orchestrating workflows with a planner], teaches planner objects are objects used with semantic kernels to chain functions together for one goal, and connectors teaches using one or more AI models which teaches the plurality of workers [Page 5-connectors])
the processor configured to: receive an input via the semantic kernel; ([Page 8-prompt functions] teaches how a semantic kernel is able to interact with a LLM and that is through prompt functions which takes in natural language as an input which teaches this limitation)
Hughes and Talebirad are analogous art because they both deal with handling multiple LLMs for task completion
It would have been obvious to a person skilled in the art before the effective filling date of the
claimed invention to combine Talebirad with the semantic kernel technology of Hughes. Doing so would lead to additional functionality for the LLMs for completing tasks ([Hughes: lines 7-10])
Regarding Claim 8, Hughes and Talebirad teaches all the limitations of Claim 1
Talebirad also teaches wherein the processor is further configured to: determine, using the manager module, a goal derived from the input ([3.1 Agent-Agent Connections] teaches agents working in collaboration towards a common goal, teaching this limitation)
determine, using the manager module, that the plurality of workers is associated with the goal ([3.2: Lines 32-39] teaches creating additional agents to delegate tasks to while assigning a goal to those new agents)
and determine, using the plurality of workers, a plurality of prompts organized in a hierarchy to send to generic large language model and to the secondary large language model. ([4.1.2 Model] teaches Auto-GPT being used in their framework which is a model capable of creating other agents which is capable of sending tasks to different models. It can be reasonably interpreted the models being used and created are capable of processing structured data along with also functioning as generic LLMs. The task assignments that can be assigned to these created models are being interpreted as prompts. The hierarchy is present with their being a main agent and agents stemming from that main agent teaching this limitation)
Regarding Claim 9, Talebirad and Hughes teaches all the limitations of Claim 1
Talebirad also teaches wherein the plurality of workers further comprises a third worker, and the processor is further configured to: invoke, using the third worker, the generic large language model and the secondary large language model ([4.1.1 Model] teaches Auto-GPT which has a main agent capable of creating other models for delegating tasks to perform tasks. The main agent is being interpreted as a third worker teaching this limitation)
Regarding Claim 10, Hughes and Talebirad teaches all the limitations of Claim 1
Talebirad also teaches wherein the processor is further configured to: receive, via the manager module, a first intermediate result from the first worker and a second intermediate result from the second worker; ([3.2: Lines 32-43] teaches an agent creating new agents with assigned goals all working towards a common goal with each of those new agents getting assigned their own goals. The agent dynamically creating the new agents is being interpreted as a manager model, and with the additional agents are being interpreted as first and second workers. It can be reasonably interpreted that the original agent is receiving those results from the created agents)
Talebirad does not teach transmit the first intermediate result and the second intermediate result from the manager module to the semantic kernel;
merge, using the semantic kernel, the first intermediate result and the second intermediate result to generate a reply, wherein the reply comprises unstructured output data;
and output the reply using the semantic kernel.
However, Hughes does teach transmit the first intermediate result and the second intermediate result from the manager module to the semantic kernel; ([Page 57- Orchestrating workflows with a Planner] teaches the planner object being used with semantic kernels taking in results from two different functions which teaches this limitation)
merge, using the semantic kernel, the first intermediate result and the second intermediate result to generate a reply, wherein the reply comprises unstructured output data; ([Page 57-Orchestrating workflows with a Planner] teaches a poem object that is taking the results from the planner object which teaches this limitation)
and output the reply using the semantic kernel. ([Page 57-58: Orchestrating workflows with a Planner] teaches printing that planner object which teaches this limitation)
Regarding Claim 11, see the analysis of Claim 1
Regarding Claim 18, see the analysis of Claim 8
Regarding Claim 19, See the analysis of Claim 10
Regarding Claim 20, See the analysis of Claim 1
Claims 2-3 and 12-13 are rejected under U.S.C. 103 as being unpatentable over Talebirad (NPL: Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents) (“Talebirad”) in view of Hughes (NPL: A Pythonista’s Intro to Semantic Kernel) (“Hughes”) and Lee et al (NPL: A Unified Industrial Large Knowledge Model Framework in Smart Manufacturing) (“Lee”)
Regarding Claim 2, Talebirad and Hughes teaches all the limitations of Claim 1
Talebirad does not teach wherein the secondary large language model is trained using specific domain knowledge.
However, Lee does teach wherein the secondary large language model is trained using specific domain knowledge ([4.3] teaches training their LLM with domain-specific data teaching this limitation)
Talebirad, Hughes and Lee are analogous art because they all deal with LLMs for task completion
It would have been obvious to a person skilled in the art before the effective filling date of the
claimed invention to combine Talebirad with the semantic kernel processing of Hughes, and the private data handling of Lee. Doing so would lead to more efficient LLMs being used in the industry ([Lee-Abstract])
Regarding Claim 3, Talebirad, Hughes and Lee teaches all the limitations of Claim 2
Hughes also teaches wherein the generic language model is public ([Page 5-Connectors] teaches different models being used with semantic kernels including models from OpenAI, which provides public generic LLMs which teaches this limitation)
Talebirad does not teach wherein the secondary language model is private to an organization, and the specific domain knowledge comprises training data labeled as private to the organization;
and wherein the input comprises structured input data labeled as private to the organization and unstructured natural language.
However, Lee does teach wherein the secondary language model is private to an organization, and the specific domain knowledge comprises training data labeled as private to the organization; ([Page 4-Table 1] teaches the ILKM containing domain specific data, that data being from a private closed source with the data privacy and security step detailing that the data can be hosed with the company’s environment which teaches this limitation)
and wherein the input comprises structured input data labeled as private to the organization and unstructured natural language. ([Page 4-Table 1] teaches the data being human-interpretable data and structured machine-generated data the human interpretable data is being interpreted as unstructured natural language which teaches this limitation)
Regarding Claim 12, See the analysis of Claim 2
Regarding Claim 13, See the analysis of Claim 3
Claims 4 and 14 are rejected under U.S.C. 103 as being unpatentable over Talebirad (NPL: Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents) (“Talebirad”) in view of Hughes (NPL: A Pythonista’s Intro to Semantic Kernel) (“Hughes”) and Dr. Varashita Sher (NPL: A Gentle Intro to Chaining LLMs, Agents, and utils via LangChain) (“Sher”)
Regarding Claim 4, Talebirad and Hughes teaches all the limitations of Claim 1
Talebirad does not teach wherein the manager module invokes the plurality of workers in a stepwise sequence, including invoking the first worker first, and after determining the first worker has completed a first process, the manager module invokes the second worker.
However, Sher does teach wherein the manager module invokes the plurality of workers in a stepwise sequence, including invoking the first worker first, and after determining the first worker has completed a first process, the manager module invokes the second worker ([Page 14] teaches a tool invoking multiple LLM chains with the goal of recommending a podcast, with the first LLMChain used for making the API call, and the second LLMChain summarizing that response to answer the question. The tool is interpreted as a manager module and the two LLMChains being the first and second worker)
Talebirad, Hughes, and Sher are analogous art because they all deal with LLMs for task completion
It would have been obvious to a person skilled in the art before the effective filling date of the
claimed invention to combine Talebirad with the semantic kernel processing of Hughes, and the LLM chaining of Sher. Doing so would lead to a more efficient chatbots ([Sher-Introduction])
Regarding Claim 14, See the analysis of Claim 4
Claims 5-6 and 15-16 are rejected under U.S.C. 103 as being unpatentable over Talebirad (NPL: Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents) (“Talebirad”) in view of Hughes (NPL: A Pythonista’s Intro to Semantic Kernel) (“Hughes”) and Block et al (NPL: Summary Cycles: Exploring the Impact of Prompt Engineering on Large Language Models’ Interaction with Interaction Log Information) (“Block”)
Regarding Claim 5, Talebirad and Hughes teaches all the limitations of Claim 1
Talebirad does not teach wherein the input comprises unstructured input data and structured input data, and the first worker invokes the generic large language module by at least generating a first prompt based on the unstructured data input and sending the first prompt to the generic large language model.
However, Block does teach wherein the input comprises unstructured input data and structured input data, and the first worker invokes the generic large language module by at least generating a first prompt based on the unstructured data input and sending the first prompt to the generic large language model. ([3.6: prompt design] teaches their prompting method following previous work which shows clear and specific instructions which functions as unstructured input data. The prompt sequences come from an interactions log dataset [3.4] which functions as structured data and those prompts are being used with ChatGPT which teaches this limitation)
Talebirad, Hughes, and Block are analogous art because they all deal with LLMs for task completion
It would have been obvious to a person skilled in the art before the effective filling date of the
claimed invention to combine Talebirad with the semantic kernel processing of Hughes, and the prompt construction of Block. Doing so would improve work efficiency ([Block-Abstract])
Regarding Claim 6, Talebirad, Hughes, and Block teaches all the limitation of Claim 5
Block also teaches wherein the second worker invokes the secondary large language module by at least generating a second prompt based on the structured input data and sending the second prompt to the secondary large language model ([Page 5: 15-25] teaches prompting ChatGPT by linking the input of different requests. The follow up prompt is a second prompt and ChatGPT is a model that can process structured data which teaches the secondary large language model)
Regarding Claim 15, see the analysis of Claim 5
Regarding Claim 16, see the analysis of Claim 6
Claim 7 and 17 are rejected Talebirad (NPL: Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents) (“Talebirad”) in view of Hughes (NPL: A Pythonista’s Intro to Semantic Kernel) (“Hughes”) and Mark et al (NPL: Multiple Model Support for Semantic Functions) (“Mark”)
Regarding Claim 7, Talebirad and Hughes teaches all the limitations of Claim 1
Hughes also teaches wherein the application further comprises a plurality of connectors that are in data communication with the semantic kernel, the plurality of connectors comprising a first connector configured to communicate with the generic large language model and a second connector configured to communicate with the secondary large language model; ([Page 5-Connectors] teaches connectors used with kernels with those connectors allowing the user to add one or more ai models, with some examples of those models being models from OpenAI, Azure OpenAI, and Hugging face. One of the models they teach is the azure_gpt_35_text_completion is being interpreted as a generic model. [Page 5-Page 6: Creating a custom connector] teaches multiple connectors in their code snippet. Since semantic kernel has access to different models from OpenAI, Azure OpenAI, and Hugging face it can be reasonably interpreted that one of those models are able to process structured data which teaches the secondary LLM which teaches this limitation)
Hughes does not teach and wherein the first worker generates a first prompt that is transmitted via the semantic kernel and the first connector to the generic large language model;
and wherein the second worker generates a second prompt that is transmitted via the semantic kernel and the second connector to the secondary large language model.
However, Mark does teach and wherein the first worker generates a first prompt that is transmitted via the semantic kernel and the first connector to the generic large language model; ([Pages 1-2: Descriptions of the Use Cases] teaches four different LLMs being used in conjunction with a kernel object. With the kernel taking in a prompt and executing a function. In order for a kernel to interact with a LLM a connector has to be present teaching the first connector to the LLM limitation. With the first model in the ordered list being used which is being interpreted as a first worker which teaches this limitation)
and wherein the second worker generates a second prompt that is transmitted via the semantic kernel and the second connector to the secondary large language model. ([Pages 1-2: Descriptions of the Use Cases] teaches four different LLMs being used in conjunction with a kernel object. With the kernel taking in a prompt and executing a function. In order for a kernel to interact with a LLM a connector has to be present teaching the first connector to the LLM limitation. The second model in the ordered list being present that would be used when the max tokens are used on the first model would teach the limitation of the secondary large language model)
Talebirad, Hughes, and Mark are analogous art with handling multiple LLMs for task completion
It would have been obvious to a person skilled in the art before the effective filling date of the
claimed invention to combine Talebirad with the semantic kernel technology of Hughes and the multiple LLM handling of Mark. Doing so would for cost reduction in the use of the LLMs ([Mark-Context and Problem Statement])
Regarding Claim 17, see the analysis of Claim 7
Conclusion
The prior arts are made of record and relied upon is considered to applicant’s disclosure
Hari et al NPL: Tryage: Real-time, Intelligent Routing of User Prompts to Large Language Models (08-23-2023) ([Abstract] “The introduction of the transformer architecture and the self-attention mechanism has led to an explosive production of language models trained on specific downstream tasks and data domains. With over 200,000 models in the Hugging Face ecosystem, users grapple with selecting and optimizing models to suit multifaceted workflows and data domains while addressing computational, security, and recency concerns. There is an urgent need for machine learning frameworks that can eliminate the burden of model selection and customization and unleash the incredible power of the vast emerging model library for end users. Here, we propose a context aware routing system, Tryage, that leverages a language model router for optimal selection of expert models from a model library based on analysis of individual input prompts. Inspired by the thalamic router in the brain, Tryage employs a perceptive router to predict down-stream model performance on prompts and, then, makes a routing decision using an objective function that integrates performance predictions with user goals and constraints that are incorporated through flags (e.g., model size, model recency). Tryage allows users to explore a Pareto front and automatically trade-off between task accuracy and secondary goals including minimization of model size, recency, security, verbosity, and readability. Across heterogeneous data sets that include code, text, clinical data, and patents, the Tryage framework surpasses Gorilla and GPT3.5 turbo in dynamic model selection identifying the optimal model with an accuracy of 50.9%, compared to 23.6% by GPT 3.5 Turbo and 10.8% by Gorilla. Conceptually, Tryage demonstrates how routing models can be applied to program and control the behavior of multi-model LLM systems to maximize efficient use of the expanding and evolving language model ecosystem.”
Tu et al US 20250095039 A1 (2023-12-29) ([Abstract] “The present disclosure relates to systems, software, and computer-implemented methods for autonomous conversational ordering using AI technologies. An example method includes receiving one or more input messages from a user via a dialogue user interface (UI). The method further includes generating, by a first large language model (LLM), one or more output messages based on the one or more input messages and information of a provider. The method further includes transmitting the one or more output messages to the user via the dialogue UI. The method further includes determining, by a second LLM, that the user has submitted a request associated with an order with the provider. The method further includes generating a description of the request in a format in compliance with an order processing system of the provider and transmitting the description to the order processing system.”)
Arne et al EP 4564260 A1 (2023-12-01) ([Abstract] “The present invention relates to a method for automated review of documents (200) to be checked for compliance by using a large language model. The method comprises the steps of providing a collection of legal questions (202) to an electronic question generator system (300); receiving, with the question generator system (300), electronic document data representing at least a part of a document (200) to be reviewed; generating (140), with the question generator system (300), at least one question message (240) for the large language model based on at least one part of the document data and at least one of the legal question (202) from the collection of legal questions (202); providing the question message (240) from the question generator system (300) to the large language model for processing the question message (240) in view of reference material (230); and receiving, with the question generator system (300), a response of the large language model to the question message (240). The invention further relates to a system for performing such method.”)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to URIAH V MOORE whose telephone number is (571)384-8341. The examiner can normally be reached Monday-Friday 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela Reyes can be reached at (571)270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/URIAH VENDELL MOORE/Examiner, Art Unit 2142 /Mariela Reyes/Supervisory Patent Examiner, Art Unit 2142