Prosecution Insights
Last updated: October 01, 2026
Application No. 18/795,841

ORCHESTRATING QUERIES BASED ON COMPUTING RESOURCE TYPE

Non-Final OA §103
Filed
Aug 06, 2024
Examiner
MACKALL, LARRY T
Art Unit
2139
Tech Center
2100 — Computer Architecture & Software
Assignee
Cisco Technology Inc.
OA Round
1 (Non-Final)
85%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
680 granted / 798 resolved
+30.2% vs TC avg
Moderate +8% lift
Without
With
+8.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
13 currently pending
Career history
823
Total Applications
across all art units

Statute-Specific Performance

§101
7.6%
-32.4% vs TC avg
§103
54.7%
+14.7% vs TC avg
§102
23.7%
-16.3% vs TC avg
§112
8.0%
-32.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 798 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Information Disclosure Statement The Information Disclosure Statement filed on 6 August 2024 has been considered by the examiner. Claim Interpretation No claims are interpreted under 35 U.S.C. 112(f). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-2, 4-9, and 11-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al. (Pub. No. US 2024/0362468) in view of Hamlin et al. (Pub. No. US 2024/0111610). Claim 1: Lee et al. disclose a method for load balancing user queries for artificial intelligence (AI) processing in a network, the method comprising: receiving, at a network component configured to pre-process queries for AI processing, data indicating a user query for AI processing [pars. 0038-0040 – “As illustrated, the edge inferencing system 110 includes a plurality of prompt-generating models 210 associated with the peripheral devices 112. To allow for the prompt-generating models 210 (and/or the one or more generative models 114) to act as a proxy for the generative model 134 executing on the cloud inferencing system 130, the prompt-generating models 210 may act as one or more prompt-generating models to pre-process user inputs into the edge inferencing system 110 and generate a prompt for processing based on the received input and contextual information associated with the input.” … “To initiate processing of a query, the edge inferencing system 110 receives an input through the one or more peripheral devices 112. In some aspects, the low-power model 212, which may execute continuously (e.g., as a background process, daemon, service, or the like) can ingest signals and other data generated by the peripheral devices 112 to determine when the ingested data is associated with a query. For example, the low-power model 212 may execute continuously to identify, from spoken utterances captured by a microphone or other audio capture device connected with or integral to the edge inferencing system 110, the presence of specific key words indicative of a user presenting a query for processing. These specific keywords may be identified contemporaneously with ingesting the query or prior to ingesting the query. It should be recognized that the foregoing is merely an example of a low-power model detecting that a user is inputting a query for processing, and other techniques by which the low-power model 212 can detect that ingested data is associated with a query may be contemplated.”]; identifying metadata associated with the user query [pars. 0038-0040 – Contextual information is received along with the input. (“As illustrated, the edge inferencing system 110 includes a plurality of prompt-generating models 210 associated with the peripheral devices 112. To allow for the prompt-generating models 210 (and/or the one or more generative models 114) to act as a proxy for the generative model 134 executing on the cloud inferencing system 130, the prompt-generating models 210 may act as one or more prompt-generating models to pre-process user inputs into the edge inferencing system 110 and generate a prompt for processing based on the received input and contextual information associated with the input.”)]; determining, based on at least one of the user query or the metadata, a processing requirement associated with the user query [fig. 1; par. 0029 – Complexity is determined. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.”)]; selecting, from among a first computing resource type and a second computing resource type, the first computing resource type as being more suitable for processing the user query than the second computing resource type based at least in part on the processing requirement [fig. 1; 0029-0031 – Orchestrator selects resources according to complexity of the query. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.” … “In some aspects, the orchestrator 116 can determine that the ingested query is of a sufficient level of complexity or implicates information that is hosted at either the local inferencing system 120 or the cloud inferencing system 130. In such a case, the orchestrator 116 can offload the ingested query to the local inferencing system 120 (e.g., for processing using a generative model 124 hosted at the local inferencing system 120) and/or the cloud inferencing system 130 (e.g., for processing using a generative model 134 hosted at the cloud inferencing system 130). Subsequently, the orchestrator 116 may receive a response from the system to which the ingested query is offloaded and output the received response to a user of the edge inferencing system 110 (e.g., by rendering the response on a display communicatively coupled with or integral to the edge inferencing system 110, transmitting one or more electronic messages including the response to a user of the edge inferencing system 110, outputting the received response as an audio output (e.g., as spoken output generated by a text-to-voice system) to a user of the edge inferencing system 120 etc.).”)]; and sending the user query to the first computing resource type based at least in part on the selecting [fig. 1; 0029-0031 – Orchestrator dispatches the query to processing resources according to complexity of the query. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.” … “In some aspects, the orchestrator 116 can determine that the ingested query is of a sufficient level of complexity or implicates information that is hosted at either the local inferencing system 120 or the cloud inferencing system 130. In such a case, the orchestrator 116 can offload the ingested query to the local inferencing system 120 (e.g., for processing using a generative model 124 hosted at the local inferencing system 120) and/or the cloud inferencing system 130 (e.g., for processing using a generative model 134 hosted at the cloud inferencing system 130). Subsequently, the orchestrator 116 may receive a response from the system to which the ingested query is offloaded and output the received response to a user of the edge inferencing system 110 (e.g., by rendering the response on a display communicatively coupled with or integral to the edge inferencing system 110, transmitting one or more electronic messages including the response to a user of the edge inferencing system 110, outputting the received response as an audio output (e.g., as spoken output generated by a text-to-voice system) to a user of the edge inferencing system 120 etc.).”)]. However, Lee et al. do not specifically disclose, wherein the first computing resource type is an AI computing resource and the second computing resource type is a non-AI computing resource [pars. 0022, 0027, 0099-0106 – Lee et al. disclose that different models may be sized for different compute capabilities and that edge devices have the most restricted capabilities, but do not specifically disclose that they do not have ai accelerator hardware.]; In the same field of endeavor, Hamlin et al. disclose, wherein the first computing resource type is an AI computing resource and the second computing resource type is a non-AI computing resource [par. 0200 – CPU and NPU versions. The combination provides that the edge devices of Lee may use CPU and the devices with more compute capability may use NPU. (“In a case where an AI model trained to perform semantic segmentation of background video is invoked by application(s) 412-414, for instance, orchestrator 501A may select one of a plurality of instances of such an AI model, each instance trained for the same use case on a different device (e.g., an NPU version of the AI model has better accuracy but higher latency, whereas a CPU version of the same AI model has worse accuracy but lower latency). These instances may be identified, for example, in one or more AI model instance tables contained in or referenced to in polic(ies) 602.”)]; It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Lee et al. to use different versions of models, as taught by Hamlin et al., in order to balance accuracy and latency. Claim 2 (as applied to claim 1 above): Lee et al. disclose, wherein the data is first data indicating a first user query, the processing requirement is a first processing requirement, and the metadata is first metadata, the method further comprising: receiving, at the network component, second data indicating a second user query for AI processing [pars. 0038-0040 – “As illustrated, the edge inferencing system 110 includes a plurality of prompt-generating models 210 associated with the peripheral devices 112. To allow for the prompt-generating models 210 (and/or the one or more generative models 114) to act as a proxy for the generative model 134 executing on the cloud inferencing system 130, the prompt-generating models 210 may act as one or more prompt-generating models to pre-process user inputs into the edge inferencing system 110 and generate a prompt for processing based on the received input and contextual information associated with the input.” … “To initiate processing of a query, the edge inferencing system 110 receives an input through the one or more peripheral devices 112. In some aspects, the low-power model 212, which may execute continuously (e.g., as a background process, daemon, service, or the like) can ingest signals and other data generated by the peripheral devices 112 to determine when the ingested data is associated with a query. For example, the low-power model 212 may execute continuously to identify, from spoken utterances captured by a microphone or other audio capture device connected with or integral to the edge inferencing system 110, the presence of specific key words indicative of a user presenting a query for processing. These specific keywords may be identified contemporaneously with ingesting the query or prior to ingesting the query. It should be recognized that the foregoing is merely an example of a low-power model detecting that a user is inputting a query for processing, and other techniques by which the low-power model 212 can detect that ingested data is associated with a query may be contemplated.”]; identifying second metadata associated with the second user query [pars. 0038-0040 – Contextual information is received along with the input. (“As illustrated, the edge inferencing system 110 includes a plurality of prompt-generating models 210 associated with the peripheral devices 112. To allow for the prompt-generating models 210 (and/or the one or more generative models 114) to act as a proxy for the generative model 134 executing on the cloud inferencing system 130, the prompt-generating models 210 may act as one or more prompt-generating models to pre-process user inputs into the edge inferencing system 110 and generate a prompt for processing based on the received input and contextual information associated with the input.”)]; determining, based on at least one of the second user query or the second metadata, a second processing requirement associated with the second user query [fig. 1; par. 0029 – Complexity is determined. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.”)]; selecting, from among the first computing resource type and the second computing resource type, the second computing resource type as being more suitable for processing the user query than the first computing resource type based at least in part on the second processing requirement [fig. 1; 0029-0031 – Orchestrator selects resources according to complexity of the query. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.” … “In some aspects, the orchestrator 116 can determine that the ingested query is of a sufficient level of complexity or implicates information that is hosted at either the local inferencing system 120 or the cloud inferencing system 130. In such a case, the orchestrator 116 can offload the ingested query to the local inferencing system 120 (e.g., for processing using a generative model 124 hosted at the local inferencing system 120) and/or the cloud inferencing system 130 (e.g., for processing using a generative model 134 hosted at the cloud inferencing system 130). Subsequently, the orchestrator 116 may receive a response from the system to which the ingested query is offloaded and output the received response to a user of the edge inferencing system 110 (e.g., by rendering the response on a display communicatively coupled with or integral to the edge inferencing system 110, transmitting one or more electronic messages including the response to a user of the edge inferencing system 110, outputting the received response as an audio output (e.g., as spoken output generated by a text-to-voice system) to a user of the edge inferencing system 120 etc.).”)]; and sending the second user query to the second computing resource type based at least in part on the selecting [fig. 1; 0029-0031 – Orchestrator dispatches the query to processing resources according to complexity of the query. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.” … “In some aspects, the orchestrator 116 can determine that the ingested query is of a sufficient level of complexity or implicates information that is hosted at either the local inferencing system 120 or the cloud inferencing system 130. In such a case, the orchestrator 116 can offload the ingested query to the local inferencing system 120 (e.g., for processing using a generative model 124 hosted at the local inferencing system 120) and/or the cloud inferencing system 130 (e.g., for processing using a generative model 134 hosted at the cloud inferencing system 130). Subsequently, the orchestrator 116 may receive a response from the system to which the ingested query is offloaded and output the received response to a user of the edge inferencing system 110 (e.g., by rendering the response on a display communicatively coupled with or integral to the edge inferencing system 110, transmitting one or more electronic messages including the response to a user of the edge inferencing system 110, outputting the received response as an audio output (e.g., as spoken output generated by a text-to-voice system) to a user of the edge inferencing system 120 etc.).”)]. Claim 4 (as applied to claim 1 above): Lee et al. disclose, wherein the data is first data indicating a first user query, the processing requirement is a first processing requirement, and the metadata is first metadata, the method further comprising: receiving, at the network component, user input data, wherein the user input data is responsive to an output associated with the first user query and the first computing resource type [pars. 0038-0040 – A query may be received based on the response of a previous query. (“By using different-sized models, generative models can be used to generate responses to queries on a variety of devices. However, generally speaking, the size of a model may be related to the ability of the model to generate accurate responses to input queries. For example, more compact models may be able to generate accurate responses to a smaller range of queries than larger models, but as discussed, may be deployed on devices which may not be able to execute operations using these larger models due to a lack of available computing resources.” … “As illustrated, the edge inferencing system 110 includes a plurality of prompt-generating models 210 associated with the peripheral devices 112. To allow for the prompt-generating models 210 (and/or the one or more generative models 114) to act as a proxy for the generative model 134 executing on the cloud inferencing system 130, the prompt-generating models 210 may act as one or more prompt-generating models to pre-process user inputs into the edge inferencing system 110 and generate a prompt for processing based on the received input and contextual information associated with the input.” … “To initiate processing of a query, the edge inferencing system 110 receives an input through the one or more peripheral devices 112. In some aspects, the low-power model 212, which may execute continuously (e.g., as a background process, daemon, service, or the like) can ingest signals and other data generated by the peripheral devices 112 to determine when the ingested data is associated with a query. For example, the low-power model 212 may execute continuously to identify, from spoken utterances captured by a microphone or other audio capture device connected with or integral to the edge inferencing system 110, the presence of specific key words indicative of a user presenting a query for processing. These specific keywords may be identified contemporaneously with ingesting the query or prior to ingesting the query. It should be recognized that the foregoing is merely an example of a low-power model detecting that a user is inputting a query for processing, and other techniques by which the low-power model 212 can detect that ingested data is associated with a query may be contemplated.”)]; receiving, at the network component, second data indicating a second user query for AI processing [pars. 0038-0040 – Input and contextual information is received. (“As illustrated, the edge inferencing system 110 includes a plurality of prompt-generating models 210 associated with the peripheral devices 112. To allow for the prompt-generating models 210 (and/or the one or more generative models 114) to act as a proxy for the generative model 134 executing on the cloud inferencing system 130, the prompt-generating models 210 may act as one or more prompt-generating models to pre-process user inputs into the edge inferencing system 110 and generate a prompt for processing based on the received input and contextual information associated with the input.”)]; identifying second metadata associated with the second user query [pars. 0038-0040 – Contextual information is received along with the input. (“As illustrated, the edge inferencing system 110 includes a plurality of prompt-generating models 210 associated with the peripheral devices 112. To allow for the prompt-generating models 210 (and/or the one or more generative models 114) to act as a proxy for the generative model 134 executing on the cloud inferencing system 130, the prompt-generating models 210 may act as one or more prompt-generating models to pre-process user inputs into the edge inferencing system 110 and generate a prompt for processing based on the received input and contextual information associated with the input.”)]; determining, based on at least one of the second user query or the second metadata, a second processing requirement associated with the second user query [fig. 1; par. 0029 – Complexity is determined. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.”)]; selecting, from among the first computing resource type and the second computing resource type, the second computing resource type as being more suitable for processing the second user query than the first computing resource type based at least in part on the second processing requirement and the user input data [fig. 1; 0029-0031 – Orchestrator selects resources according to complexity of the query. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.” … “In some aspects, the orchestrator 116 can determine that the ingested query is of a sufficient level of complexity or implicates information that is hosted at either the local inferencing system 120 or the cloud inferencing system 130. In such a case, the orchestrator 116 can offload the ingested query to the local inferencing system 120 (e.g., for processing using a generative model 124 hosted at the local inferencing system 120) and/or the cloud inferencing system 130 (e.g., for processing using a generative model 134 hosted at the cloud inferencing system 130). Subsequently, the orchestrator 116 may receive a response from the system to which the ingested query is offloaded and output the received response to a user of the edge inferencing system 110 (e.g., by rendering the response on a display communicatively coupled with or integral to the edge inferencing system 110, transmitting one or more electronic messages including the response to a user of the edge inferencing system 110, outputting the received response as an audio output (e.g., as spoken output generated by a text-to-voice system) to a user of the edge inferencing system 120 etc.).”)]; and sending the second user query to the second computing resource type based at least in part on the selecting [fig. 1; 0029-0031 – Orchestrator dispatches the query to processing resources according to complexity of the query. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.” … “In some aspects, the orchestrator 116 can determine that the ingested query is of a sufficient level of complexity or implicates information that is hosted at either the local inferencing system 120 or the cloud inferencing system 130. In such a case, the orchestrator 116 can offload the ingested query to the local inferencing system 120 (e.g., for processing using a generative model 124 hosted at the local inferencing system 120) and/or the cloud inferencing system 130 (e.g., for processing using a generative model 134 hosted at the cloud inferencing system 130). Subsequently, the orchestrator 116 may receive a response from the system to which the ingested query is offloaded and output the received response to a user of the edge inferencing system 110 (e.g., by rendering the response on a display communicatively coupled with or integral to the edge inferencing system 110, transmitting one or more electronic messages including the response to a user of the edge inferencing system 110, outputting the received response as an audio output (e.g., as spoken output generated by a text-to-voice system) to a user of the edge inferencing system 120 etc.).”)]. Claim 5 (as applied to claim 1 above): Lee et al. disclose, wherein the metadata includes an indication of: a feature associated with a file included with the user query; a file extension associated with the file; or a feature associated with a user prompt included with the user query [pars. 0038-0040 – Contextual information is received along with the input. (“As illustrated, the edge inferencing system 110 includes a plurality of prompt-generating models 210 associated with the peripheral devices 112. To allow for the prompt-generating models 210 (and/or the one or more generative models 114) to act as a proxy for the generative model 134 executing on the cloud inferencing system 130, the prompt-generating models 210 may act as one or more prompt-generating models to pre-process user inputs into the edge inferencing system 110 and generate a prompt for processing based on the received input and contextual information associated with the input.”)]. Claim 6 (as applied to claim 1 above): Lee et al. disclose the method, further comprising: receiving, at the network component, configuration data indicating a configuration associated with the network [par. 0023 – Nodes in the system receive information in order to facilitate coordination between nodes. (“Aspects of the present disclosure provide techniques for orchestrating query processing by generative artificial intelligence models in a hybrid computing environment. In orchestrating or otherwise coordinating query processing across different devices in a hybrid computing environment including edge devices and cloud computing environments, queries can be executed on specific devices based on the properties of the query. In some aspects, information available at these devices can be used to augment responses generated by generative artificial intelligence models in the hybrid computing environment. Thus, queries can be routed for execution by the device(s) in the hybrid computing environment which can generate an accurate response while allowing computing resources on other devices to remain available for processing other queries using generative artificial intelligence models.”)]; and determining, based at least in part on the processing requirement and the configuration data, the first computing resource type as being more suitable for processing the user query [fig. 1; 0029-0031 – Orchestrator dispatches the query to processing resources according to complexity of the query. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.” … “In some aspects, the orchestrator 116 can determine that the ingested query is of a sufficient level of complexity or implicates information that is hosted at either the local inferencing system 120 or the cloud inferencing system 130. In such a case, the orchestrator 116 can offload the ingested query to the local inferencing system 120 (e.g., for processing using a generative model 124 hosted at the local inferencing system 120) and/or the cloud inferencing system 130 (e.g., for processing using a generative model 134 hosted at the cloud inferencing system 130). Subsequently, the orchestrator 116 may receive a response from the system to which the ingested query is offloaded and output the received response to a user of the edge inferencing system 110 (e.g., by rendering the response on a display communicatively coupled with or integral to the edge inferencing system 110, transmitting one or more electronic messages including the response to a user of the edge inferencing system 110, outputting the received response as an audio output (e.g., as spoken output generated by a text-to-voice system) to a user of the edge inferencing system 120 etc.).”)]. Claim 7 (as applied to claim 6 above): Lee et al. disclose, wherein the configuration includes: a threshold usage associated with the AI computing resource; a threshold time associated with the AI computing resource; a priority associated with a user; a priority associated with the user query; or computing resources available in the network [fig. 1; 0029-0031 – Orchestrator dispatches the query to processing resources according to complexity of the query. (“The orchestrator 116 generally identifies which system in the hybrid computing environment 100 is to process the ingested query (or parts thereof) and routes the ingested query to the identified system for processing. In some aspects, to identify the system that is to process the ingested query, the orchestrator 116 can, in some aspects, examine information about the topic of the query and estimate the complexity of the query (e.g., a complexity metric) based on the topic of the query. For topics with a complexity metric below a defined threshold or topics included in a defined set of topics that can be addressed using the generative model 114 at the edge inferencing system 110, the orchestrator 116 can dispatch the ingested query (and, in some aspects, the contextual information derived from the peripheral devices 112 and/or knowledge from the personal knowledge repository 118) to the generative model 114 for processing. In some aspects, the information based on which the generative model 114 generates a response to the ingested query may be further supplemented by data retrieved by the orchestrator 116 from one or more external resources (e.g., external tools 136 hosted at the cloud inferencing system 130) and/or one or more internal resources.” … “In some aspects, the orchestrator 116 can determine that the ingested query is of a sufficient level of complexity or implicates information that is hosted at either the local inferencing system 120 or the cloud inferencing system 130. In such a case, the orchestrator 116 can offload the ingested query to the local inferencing system 120 (e.g., for processing using a generative model 124 hosted at the local inferencing system 120) and/or the cloud inferencing system 130 (e.g., for processing using a generative model 134 hosted at the cloud inferencing system 130). Subsequently, the orchestrator 116 may receive a response from the system to which the ingested query is offloaded and output the received response to a user of the edge inferencing system 110 (e.g., by rendering the response on a display communicatively coupled with or integral to the edge inferencing system 110, transmitting one or more electronic messages including the response to a user of the edge inferencing system 110, outputting the received response as an audio output (e.g., as spoken output generated by a text-to-voice system) to a user of the edge inferencing system 120 etc.).”)]. Claim 8: Claim 8, directed to a system [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 1 above, mutatis mutandis. Claim 9 (as applied to claim 8 above): Claim 9, directed to a system [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 2 above, mutatis mutandis. Claim 11 (as applied to claim 8 above): Claim 11, directed to a system [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 4 above, mutatis mutandis. Claim 12 (as applied to claim 8 above): Claim 12, directed to a system [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 5 above, mutatis mutandis. Claim 13 (as applied to claim 8 above): Claim 13, directed to a system [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 6 above, mutatis mutandis. Claim 14 (as applied to claim 13 above): Claim 14, directed to a system [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 7 above, mutatis mutandis. Claim 15: Claim 15, directed to one or more non-transitory computer-readable media [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 1 above, mutatis mutandis. Claim 16 (as applied to claim 15 above): Claim 16, directed to one or more non-transitory computer-readable media [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 2 above, mutatis mutandis. Claim 17 (as applied to claim 15 above): Claim 17, directed to one or more non-transitory computer-readable media [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 4 above, mutatis mutandis. Claim 18 (as applied to claim 15 above): Claim 18, directed to one or more non-transitory computer-readable media [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 5 above, mutatis mutandis. Claim 19 (as applied to claim 15 above): Claim 19, directed to one or more non-transitory computer-readable media [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 6 above, mutatis mutandis. Claim 20 (as applied to claim 19 above): Claim 20, directed to one or more non-transitory computer-readable media [Lee et al. – par. 0007], is rejected for the same reasons set forth in the rejection of claim 7 above, mutatis mutandis. Allowable Subject Matter Claims 3 and 10 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to LARRY T MACKALL whose telephone number is (571)270-1172. The examiner can normally be reached Monday - Friday, 9am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Reginald G Bragdon can be reached at (571) 272-4204. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. LARRY T. MACKALL Primary Examiner Art Unit 2131 17 August 2026 /LARRY T MACKALL/Primary Examiner, Art Unit 2139
Read full office action

Prosecution Timeline

Aug 06, 2024
Application Filed
Aug 19, 2026
Non-Final Rejection mailed — §103
Sep 01, 2026
Interview Requested
Sep 17, 2026
Applicant Interview (Telephonic)
Sep 19, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743234
STORAGE DEVICE AND METHOD OF OPERATING THE SAME
2y 5m to grant Granted Sep 22, 2026
Patent 12743207
MULTITENANCY SSD CONFIGURATION
2y 5m to grant Granted Sep 22, 2026
Patent 12717666
APPLICATION PROGRAMMING INTERFACE TO INDICATE ALLOCATION OF OPERATIONS
3y 1m to grant Granted Aug 25, 2026
Patent 12710896
SEPARATE COMMAND ADDRESS (SCA) BASED MEMORY CONTROLLER
1y 6m to grant Granted Aug 18, 2026
Patent 12704986
READ REFRESH TECHNIQUES FOR NONVOLATILE MEMORY DEVICES
2y 2m to grant Granted Aug 11, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
85%
Grant Probability
93%
With Interview (+8.0%)
2y 7m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 798 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month