Prosecution Insights
Last updated: October 02, 2026
Application No. 18/754,313

HIERARCHICAL AND PEER PRUNING STRATEGIES FOR GENERATIVE ARTIFICIAL INTELLIGENCE MODELS IN TELECOMMUNICATIONS NETWORKS

Non-Final OA §101§103§112
Filed
Jun 26, 2024
Examiner
LEVEL, BARBARA HENRY
Art Unit
2142
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
72%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
250 granted / 348 resolved
+16.8% vs TC avg
Strong +28% interview lift
Without
With
+28.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
15 currently pending
Career history
359
Total Applications
across all art units

Statute-Specific Performance

§101
16.8%
-23.2% vs TC avg
§103
48.4%
+8.4% vs TC avg
§102
10.0%
-30.0% vs TC avg
§112
16.9%
-23.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 348 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This correspondence is responsive to the application filed on June 26, 2024. Claims 1-24 are pending in the case with claims 1, 9 and 17 in independent form. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Summary of Detailed Action Claims 5, 13, and 21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite. Claims 1-24 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite. Claims 1-24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Claims 1-2, 7, 9-10, 15, 17-18 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Abdollahian Noghabi et al. in view of Surat Teerapittayanon et al. and Kim et al. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 5, 13, and 21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 5, 13 and 21 recite the limitation wherein LLM-specific caching at “the edge” is performed. There is insufficient antecedent basis for this limitation in the claim. Claims 1-24 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “large language model” in independent claims 1, 9 and 17 is a relative term which renders the claim indefinite. The term large language model is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Dependent claims 2-8, 10-16 and 18-24 depend directly or indirectly from independent claims 1, 9 and 17 respectively, and are rejected for the same reasons discussed above with respect to claims 1, 9 and 17. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-24 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim(s) recite(s) subject matter at a general, high-level of a method for hierarchical inference, classify a request as appropriate for the pruned model and directing inference generation to the pruned model or to another model at a higher tier, which are mental processes or concepts that can be performed in the human mind, including observation, evaluation, judgment or opinion, or by a human using pen and paper. MPEP 210604(a)(2)(III). This judicial exception is not integrated into a practical application and the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Claims 1-24 recite one of the four statutory categories of patent able subject matter and belong to the statutory class(es) of a process (method claims 1-8), a machine (system/apparatus claims 9-16), and an article of manufacture (non-transitory computer readable media claims 17-24). Claim 1 recites a method, thus a process and one of the four statutory categories of patentable subject matter. However, claim 1 further recites for hierarchical inference, classify a request as appropriate for the pruned model and directing inference generation to the pruned model or to another model at a higher tier, which are mental processes or concepts that can be performed in the human mind, including observation, evaluation, judgment or opinion, or by a human using pen and paper. MPEP 210604(a)(2)(III). The claim does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: utilizing a large language model (LLM) (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). training, at a central location, a helper model and a pruned model for each layer of a hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). wherein the helper model is trained to (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). wherein the pruned model is generated from a reduction process of the LLM (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). distributing the helper model and pruned model to different levels of the hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). utilizing the helper model at each level of the hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). Thus, the claim is directed to the abstract idea. Further, the additional elements, alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more, and generally linking the use of the judicial exception to a particular technological field of use does not meaningfully limit the claims (MPEP 2106.04(d)) and the combination of additional elements does not provide an inventive concept. Thus, the claim is ineligible. Claim 2, dependent on claim 1, does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: using helper models in the hierarchy to accelerate processing at an edge computing node.(This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). Claim 3, dependent on claim 2, does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: wherein different helper models and differently pruned models are used at each layer of the hierarchy. (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). Claim 4, dependent on claim 3, does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: wherein a combination of helper models (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). with caching of relevant documents is performed. (An additional element of extra-solution activity that courts have identified is well understood, routine and conventional activity for receiving or transmitting data over a network, e.g., using the internet to gather data. See also, MPEP 2106.05(d)(II), MPEP 2106.05(g), 2019 Guidance, 84 FR 50 at 55, 2019 Guidance, 84 FR 50, footnote 31.). Claim 5, dependent on claim 4, does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: wherein LLM-specific (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). caching at the edge is performed. (An additional element of extra-solution activity that courts have identified is well understood, routine and conventional activity for receiving or transmitting data over a network, e.g., using the internet to gather data. See also, MPEP 2106.05(d)(II), MPEP 2106.05(g), 2019 Guidance, 84 FR 50 at 55, 2019 Guidance, 84 FR 50, footnote 31.). Claim 6, dependent on claim 1, does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: wherein peer nodes are used for inference prior to a hierarchically higher node. (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). Claim 7, dependent on claim 1, does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: wherein a caching strategy for local documents is performed. (An additional element of extra-solution activity that courts have identified is well understood, routine and conventional activity for receiving or transmitting data over a network, e.g., using the internet to gather data. See also, MPEP 2106.05(d)(II), MPEP 2106.05(g), 2019 Guidance, 84 FR 50 at 55, 2019 Guidance, 84 FR 50, footnote 31.). Claim 8, dependent on claim 7, does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: wherein documents with higher relevance scores to a request are cached at different locations. (An additional element of extra-solution activity that courts have identified is well understood, routine and conventional activity for receiving or transmitting data over a network, e.g., using the internet to gather data. See also, MPEP 2106.05(d)(II), MPEP 2106.05(g), 2019 Guidance, 84 FR 50 at 55, 2019 Guidance, 84 FR 50, footnote 31.). Claim 9 recites a system, thus a machine and one of the four statutory categories of patentable subject matter. However, claim 9 further recites for hierarchical inference, classify a request as appropriate for the pruned model and directing inference generation to the pruned model or to another model at a higher tier, which are mental processes or concepts that can be performed in the human mind, including observation, evaluation, judgment or opinion, or by a human using pen and paper. MPEP 210604(a)(2)(III). The claim does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: utilizing a large language model (LLM), (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). the system comprising: a memory; and a processor coupled to the memory, wherein the processor performs operations, the operations comprising (an additional element merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See also, MPEP 2106.05(f), MPEP 2106.04(d), 2019 Guidance, 84 FR 50 at 55, footnote 30.). training, at a central location, a helper model and a pruned model for each layer of a hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). wherein the helper model is trained to (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). wherein the pruned model is generated from a reduction process of the LLM (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). distributing the helper model and pruned model to different levels of the hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). utilizing the helper model at each level of the hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). Thus, the claim is directed to the abstract idea. Further, the additional elements, alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more, and generally linking the use of the judicial exception to a particular technological field of use does not meaningfully limit the claims (MPEP 2106.04(d)) and the combination of additional elements does not provide an inventive concept. Thus, the claim is ineligible. Claim 17 recites a computer program product, thus an article of manufacture and one of the four statutory categories of patentable subject matter. However, claim 17 further recites for hierarchical inference, classify a request as appropriate for the pruned model and directing inference generation to the pruned model or to another model at a higher tier, which are mental processes or concepts that can be performed in the human mind, including observation, evaluation, judgment or opinion, or by a human using pen and paper. MPEP 210604(a)(2)(III). The claim does not include any additional elements which integrate the abstract idea into a practical application since the additional elements consist of: A computer program product (an additional element merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See also, MPEP 2106.05(f), MPEP 2106.04(d), 2019 Guidance, 84 FR 50 at 55, footnote 30.). utilizing a large language model (LLM) (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). the computer program product comprising a computer readable storage medium, wherein code stored in the computer readable storage medium when executed by a processor performs operations, the operations comprising (an additional element merely recites the words “apply it” (or an equivalent) with the judicial exception, or merely includes instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See also, MPEP 2106.05(f), MPEP 2106.04(d), 2019 Guidance, 84 FR 50 at 55, footnote 30.). training, at a central location, a helper model and a pruned model for each layer of a hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). wherein the helper model is trained to (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). wherein the pruned model is generated from a reduction process of the LLM (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). distributing the helper model and pruned model to different levels of the hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). utilizing the helper model at each level of the hierarchy (This additional element amounts to merely the words to “apply it” (or an equivalent) or are mere instructions to implement an abstract idea or other exception on a computer. MPEP 2106.05(f).) Also, this additional element amounts to no more than generally linking the use of the judicial exception to a particular technologic environment or field of use - The application or use of the judicial exception in this manner does not meaningfully limit the claim by going beyond generally linking the use of the judicial exception to a particular technological environment. MPEP 2106.05(h)). Thus, the claim is directed to the abstract idea. Further, the additional elements, alone or in combination, do not provide significantly more than the abstract idea itself, because implementation on a computer (MPEP 2106.05(f)) cannot provide significantly more, and generally linking the use of the judicial exception to a particular technological field of use does not meaningfully limit the claims (MPEP 2106.04(d)) and the combination of additional elements does not provide an inventive concept. Thus, the claim is ineligible. Dependent claims 10-16 and 18-24 are comparably rejected as set forth above with respect to dependent claims 2-8. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-2, 7, 9-10, 15, 17-18 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Abdollahian Noghabi et al. (Pub. No. US 2025/0348349 A1, filed May 7, 2024) hereinafter Noghabi in view Surat Teerapittayanon et al., (“Distributed Deep Neural Networks over the Cloud, the Edge and End Devices”, DOI 10.1109/ICDCS.2017.226, September 2017) hereinafter Teerapittayanon and Kim et al. (Pub. No. US 2024/0311405 A1, filed June 19, 2023) hereinafter Kim. The Examiner notes that Teerapittayanon is cited on Applicant’s Information Disclosure Statement filed September 12, 2025. Regarding claim 1, Noghabi teaches: A method for hierarchical inference utilizing a large language model (LLM), the method comprising: training, at a central location, a helper model and a pruned model for each layer of a hierarchy, wherein the helper model is trained to classify a request as appropriate for the pruned model, wherein the pruned model is generated from a reduction process of the LLM; Noghabi teaches that, The systems and methods include a model selector that dynamically selects a language model to use in a tier of the hierarchical edge architecture in response to a query received by a user. In some implementations, the model selector uses query features and a context in selecting a language model (a small language model at the frontline edge, a medium language model at the back office edge, or an LLM at the cloud) to use to provide response to the query (hierarchical inference utilizing a )). Noghabi, Figs 1-3, para 17. Noghabi teaches and illustrates in Figure 1 fine-tuning training at a central cloud location a model and fine-tuning training a model at each frontline edge, backoffice edge and cloud tier layer of the hierarchy. Noghabi teaches that, The systems and methods include a model selector (helper) that dynamically selects a language model (a model) to use in a tier of the hierarchical edge architecture (a model for each layer of a hierarchy) in response to a query received by a user. In some implementations, the model selector ((a helper ) uses query features and a context in selecting a language model (a small language model at the frontline edge, a medium language model at the back office edge, or an LLM at the cloud) to use to provide response to the query (and a ). For example, the context includes available network connectivity at a user device that receives the query from the user. Another example of context includes device parameters of the user device that receives the query (e.g., a current load of the device). Another example of context includes user context with information specific to the user. Another example of context includes industry specific context (e.g., agriculture context if the query relates to farming or oil context if the query relates to drilling for oil). The selected language model provides a response to the query and the response is output to the user. Noghabi, Figs 1-3, para 17, 16, 19, 44- 49. Thus, Noghabi teaches hierarchical inference utilizing a language model, a helper and a language model for each layer of a hierarchy, wherein the helper is to classify a query request as appropriate for the language model at a layer of the hierarchy. Noghabi does not specifically disclose training at a central location training, a helper model and model for each layer of a hierarchy. However, Teerapittayanon teaches in the field related to distributed machine learning computing hierarchies. Teerapittayanon, Abstract, Introduction, page 328-329. Teerapittayanon is analogous to the claimed invention because Teerapittayanon is directed to distributed machine learning computing hierarchies. While DDNN inference is distributed over the distributed computing hierarchy, the DDNN system can be trained on a single powerful server or in the cloud (training at a central location training, a model). Teerapittayanon, Section III. C.DDNN Training, Figure 2(e), pages 331, 332. DDNN maps a trained DNN onto heterogeneous physical devices distributed locally, at the edge, and in the cloud (training at a central location training, a model for each layer of a hierarchy). Teerapittayanon, Section III. A. DDNN Architecture, Figure 2(e), pages 330-331, 332. Teerapittayanon in Section III C. and Figure 2 (e) teaches and illustrates how the trained model is distributed among the local devices, the edge, and the cloud (different levels of the hierarchy). Each subsection of the model may be considered a pruned version of the model obtained from reducing the overall model (a pruned model, centrally trained, for each layer of a hierarchy). Teerapittayanon, Section III. C.DDNN Training, Figure 2(e), pages 331, 332. Inference in DDNN is performed in several stages using multiple preconfigured exit thresholds T (one element T at each exit point) as a measure of confidence in the prediction of the sample (Teerapittayanon describes a helper model at each layer in the hierarchy as the DDNN model itself along with a threshold T that determines whether to exit early at that stage of the hierarchy (centrally trained helper model at each layer in the hierarchy). One way to define T is by searching over the ranges of T on a validation set and pick the one with the best accuracy. … At each exit point, η is computed and compared against T in order to determine if the sample should exit at that point. At a given exit point, if the predictor is not confident in the result (i.e., η>T), the system falls back to a higher exit point in the hierarchy until the last exit is reached which always performs classification. Teerapittayanon, Section III. D. Inference, Figure 2(e), pages 332. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement the hierarchical inference utilizing a language model, a helper and a language model for each layer of a hierarchy, wherein the helper is to classify a query request as appropriate for the language model at a layer of the hierarchy of Noghabi using the training at a central location a helper model and pruned model for each layer of a hierarchy of Teerapittayanon, with a reasonable expectation of success, in order to provide a simpler and more principled approach. For distributed computing hierarchies, consisting of the cloud, the edge (fog) and geographically distributed end devices. Teerapittayanon, Section I. Introduction, pages 329, 328. This would have provided the advantages of centralized training with distributed computing hierarchies. Thus, Noghabi in view of Teerapittayanon teaches a method for hierarchical inference utilizing a However, Kim teaches in the field related to large language models. Kim, Abstract, para 1-3. Kim, which is analogous to the claimed invention because Kim is directed to directed to selecting, in response to receiving a request and from among multiple candidate generative models with differing computational efficiencies, a particular generative model to utilize in generating a response to the request, teaches that, As a non-limiting working example, assume that the candidate generative models include a larger LLM that includes over 200 billion parameters and a smaller LLM that includes twenty, thirty, forty, fifty, or other percent less parameters than the larger LLM. For instance, the smaller LLM can include less than 100 billion parameters. In some implementations, the smaller LLM can be a quantized and/or pruned version of the larger LLM (pruned model, wherein the pruned model is generated from a reduction process of the LLM). LLM. Kim, Fig 1, para 5, 4. In some implementations and/or for some requests, a trained machine learning (ML) model (e.g., a neural network model) can be used in selecting from among multiple candidate generative models (e.g. from between at least the smaller LLM and the larger LLM). Kim, Fig. 1, para 12, 97, 45, 4-5. In some implementations, selecting, based on the request features of the request, the particular LLM, includes processing the request features using a trained machine learning model to generate output that indicates, for each of the candidate LLMs, a corresponding probability (e.g., a corresponding probability of generating a correct response) and selecting the particular LLM based on the particular probability, for the particular LLM, indicated by the output The corresponding probabilities indicated by the output include a particular probability for the particular LLM and an additional probability for the additional LLM. Further, the trained machine learning model is more computationally efficient than the additional LLM and, optionally, is more computationally efficient than the particular LLM. Kim, Fig. 1, para 97, 12, 45, 5,4. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement the hierarchical inference utilizing a language model, a helper and a language model for each layer of a hierarchy, wherein the helper is to classify a query request as appropriate for the language model at a layer of the hierarchy of Noghabi using the training at a central location a helper model and a pruned model for each layer of a hierarchy of Teerapittayanon and the pruned model, wherein the pruned model is generated from a reduction process of the LLM of Kim, with a reasonable expectation of success, in order to provide a simpler and more principled approach. For distributed computing hierarchies, consisting of the cloud, the edge (fog) and geographically distributed end devices and in order to provide selecting from among multiple candidate generative models with differing computational efficiencies, a particular generative model to utilize in generating a response to the request and to reduce latency and/or conserve computational resource(s) and to mitigate occurrences of a generated response being inaccurate. Teerapittayanon, Section I. Introduction, pages 329, 328. Kim, para 4. This would have provided the advantages of centralized training with distributed computing hierarchies and of providing model selection from efficient generative large language models. Thus, Noghabi in view of Teerapittayanon and Kim teaches a method for hierarchical inference utilizing a large language model (LLM), the method comprising: training, at a central location, a helper model and a pruned model for each layer of a hierarchy, wherein the helper model is trained to classify a request as appropriate for the pruned model, wherein the pruned model is generated from a reduction process of the LLM. distributing the helper model and pruned model to different levels of the hierarchy; As similarly discussed above, Noghabi teaches hierarchical inference utilizing a language model. Noghabi, Figs 1-3, para 17, 16, 19, 44-49. As similarly discussed above, Noghabi does not specifically disclose distributing the helper model and pruned model to different levels of the hierarchy. However, Teerapittayanon teaches in the field related to distributed machine learning computing hierarchies. Teerapittayanon, Abstract, Introduction, page 328-329. Teerapittayanon is analogous to the claimed invention because Teerapittayanon is directed to distributed machine learning computing hierarchies. While DDNN inference is distributed over the distributed computing hierarchy, the DDNN system can be trained on a single powerful server or in the cloud (distributing the helper model (helper model, see also section III.D) and pruned model (pruned model, see also section III.C) to different levels of the hierarchy (see also III. A and Figure 2(e))). Teerapittayanon, Section III. C.DDNN Training, Figure 2(e), pages 331, 332. DDNN maps a trained DNN onto heterogeneous physical devices distributed locally, at the edge, and in the cloud (training at a central location training, a model for each layer of a hierarchy, distributing the helper model and pruned model to different levels of the hierarchy (see also III. A and Figure 2(e))).). Teerapittayanon, Section III. A. DDNN Architecture, Figure 2(e), pages 330-331, 332. Teerapittayanon in Section III C. and Figure 2 (e) teaches and illustrates how the trained model is distributed among the local devices, the edge, and the cloud (distributing the helper model (helper model, see also section III.D) and pruned model (pruned model, see also section III.C) to different levels of the hierarchy (see also III. A and Figure 2(e))).). Each subsection of the model may be considered a pruned version of the model obtained from reducing the overall model (a pruned model, centrally trained for each layer of a hierarchy). Teerapittayanon, Section III. C.DDNN Training, Figure 2(e), pages 331, 332. Inference in DDNN is performed in several stages using multiple preconfigured exit thresholds T (one element T at each exit point) as a measure of confidence in the prediction of the sample (Teerapittayanon describes a helper model distributed at each layer in the hierarchy as the DDNN model itself along with a threshold T that determines whether to exit early at that stage of the hierarchy (centrally trained helper model distributed to each layer in the hierarchy). One way to define T is by searching over the ranges of T on a validation set and pick the one with the best accuracy. … At each exit point, η is computed and compared against T in order to determine if the sample should exit at that point. At a given exit point, if the predictor is not confident in the result (i.e., η>T), the system falls back to a higher exit point in the hierarchy until the last exit is reached which always performs classification. Teerapittayanon, Section III. D. Inference, Figure 2(e), pages 332. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement the hierarchical inference utilizing a language model, a helper and a language model for each layer of a hierarchy, wherein the helper is to classify a query request as appropriate for the language model at a layer of the hierarchy of Noghabi using the training at a central location a helper model and pruned model for each layer of a hierarchy and distributing the helper model and pruned model to different levels of the hierarchy of Teerapittayanon and the pruned model, wherein the pruned model is generated from a reduction process of the LLM of Kim, with a reasonable expectation of success, in order to provide a simpler and more principled approach. For distributed computing hierarchies, consisting of the cloud, the edge (fog) and geographically distributed end devices and in order to provide selecting from among multiple candidate generative models with differing computational efficiencies, a particular generative model to utilize in generating a response to the request and to reduce latency and/or conserve computational resource(s) and to mitigate occurrences of a generated response being inaccurate. Teerapittayanon, Section I. Introduction, pages 329, 328. Kim, para 4. This would have provided the advantages of centralized training with distributed computing hierarchies and of providing model selection from efficient generative large language models. Thus, Noghabi in view of Teerapittayanon and Kim teaches a method for hierarchical inference utilizing a large language model (LLM), the method comprising: training, at a central location, a helper model and a pruned model for each layer of a hierarchy, wherein the helper model is trained to classify a request as appropriate for the pruned model, wherein the pruned model is generated from a reduction process of the LLM, and distributing the helper model and pruned model to different levels of the hierarchy. directing, utilizing the helper model at each level of the hierarchy, inference generation to the pruned model or to another model at a higher tier. As similarly discussed above, Noghabi teaches hierarchical inference utilizing a language model. Noghabi, Figs 1-3, para 17, 16, 19, 44- 49. As similarly discussed above, Noghabi does not specifically disclose directing, utilizing the helper model at each level of the hierarchy, inference generation to the pruned model or to another model at a higher tier. However, Teerapittayanon teaches in the field related to distributed machine learning computing hierarchies. Teerapittayanon, Abstract, Introduction, page 328-329. D. DDNN Inference … Inference in DDNN is performed in several stages using multiple preconfigured exit thresholds T (one element T at each exit point) as a measure of confidence in the prediction of the sample. … This normalized entropy η has values between 0 and 1 which allows easier interpretation and searching of its corresponding threshold T. For example, η close to 0 means that the DDNN is confident about the prediction of the sample; η close to 1 means it is not confident. At each exit point, η is computed and compared against T in order to determine if the sample should exit at that point. At a given exit point, if the predictor is not confident in the result (i.e., η>T), the system falls back to a higher exit point in the hierarchy until the last exit is reached which always performs classification. We now provide an example of the inference procedure for a DDNN which has multiple end devices and three exit points (configuration (e) in Figure 2) (As disclosed above in section III.D, Teerapittayanon describes how the entropy of the outputs of the model at a hierarchy level is compared to a threshold (a helper model at each level is utilized) to determine whether to output an inference from that level of the hierarchy or to forward information to a higher level (inference generation to the pruned model or to another model at a higher tier (pruned model, see also section III.C)). Teerapittayanon, Section III. D. DDNN Inference, pages 332, Section III. C. DDNN Training, page 331. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement the hierarchical inference utilizing a language model, a helper and a language model for each layer of a hierarchy, wherein the helper is to classify a query request as appropriate for the language model at a layer of the hierarchy of Noghabi using the training at a central location a helper model and pruned model for each layer of a hierarchy and distributing the helper model and pruned model to different levels of the hierarchy and directing, utilizing the helper model at each level of the hierarchy, inference generation to the pruned model or to another model at a higher tier of Teerapittayanon and the pruned model, wherein the pruned model is generated from a reduction process of the LLM of Kim, with a reasonable expectation of success, in order to provide a simpler and more principled approach. For distributed computing hierarchies, consisting of the cloud, the edge (fog) and geographically distributed end devices and in order to provide selecting from among multiple candidate generative models with differing computational efficiencies, a particular generative model to utilize in generating a response to the request and to reduce latency and/or conserve computational resource(s) and to mitigate occurrences of a generated response being inaccurate. Teerapittayanon, Section I. Introduction, pages 329, 328. Kim, para 4. This would have provided the advantages of centralized training with distributed computing hierarchies and of providing model selection from efficient generative large language models. Regarding claim 2, which depends from claim 1 and further recites: using helper models in the hierarchy to accelerate processing at an edge computing node. Noghabi in view of Teerapittayanon and Kim teaches the method of claim 1 from which claim 2 depends, including using the helper models in the hierarchy. the helper model and pruned model. Kim, Fig 1, para 4-5, 45. Noghabi teaches that, Another technical advantage of the systems and methods of the present disclosure is fast responses. The systems and methods use the hierarchical structure (frontline edge, back office edge, cloud) with use of several tiers of compute to enable fast and cost-effective responses (using helper frontline edge, back office edge computing node)). … The systems and methods use fine-tuned specialized models ensuring accurate and relevant insights tailored to each user's specific requirements, optimizing the overall efficiency to the systems and methods. Another technical advantage of the systems and methods of the present disclosure is dynamic model selection for each tier (using helper frontline edge, back office edge computing node)). The systems and methods dynamically select a language model to use (e.g., a local language model, a language model in the back office, or a language model in the cloud) in response to query parameters and available network connectivity. Noghabi, para 19, 16-17, 49. As similarly discussed above, Noghabi does not specifically disclose using helper models in the hierarchy. However, Teerapittayanon teaches in the field related to distributed machine learning computing hierarchies. Teerapittayanon, Abstract, Introduction, page 328-329. Teerapittayanon is analogous to the claimed invention because Teerapittayanon is directed to distributed machine learning computing hierarchies. While DDNN inference is distributed over the distributed computing hierarchy, the DDNN system can be trained on a single powerful server or in the cloud. Teerapittayanon, Section III. C.DDNN Training, Figure 2(e), pages 331, 332. DDNN maps a trained DNN onto heterogeneous physical devices distributed locally, at the edge, and in the cloud. Teerapittayanon, Section III. A. DDNN Architecture, Figure 2(e), pages 330-331, 332. Teerapittayanon in Section III C. and Figure 2 (e) teaches and illustrates how the trained model is distributed among the local devices, the edge, and the cloud (different levels of the hierarchy). Each subsection of the model may be considered a pruned version of the model obtained from reducing the overall model. Teerapittayanon, Section III. C.DDNN Training, Figure 2(e), pages 331, 332. Inference in DDNN is performed in several stages using multiple preconfigured exit thresholds T (one element T at each exit point) as a measure of confidence in the prediction of the sample (Teerapittayanon describes using helper models in the hierarchy and a helper model at each layer in the hierarchy as the DDNN model itself along with a threshold T that determines whether to exit early at that stage of the hierarchy (centrally trained helper model at each layer in the hierarchy (using helper models in the hierarchy)). One way to define T is by searching over the ranges of T on a validation set and pick the one with the best accuracy. … At each exit point, η is computed and compared against T in order to determine if the sample should exit at that point. At a given exit point, if the predictor is not confident in the result (i.e., η>T), the system falls back to a higher exit point in the hierarchy until the last exit is reached which always performs classification. Teerapittayanon, Section III. D. Inference, Figure 2(e), pages 332. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement the hierarchical inference utilizing a language model, a helper and a language model for each layer of a hierarchy, wherein the helper is to classify a query request as appropriate for the language model at a layer of the hierarchy and using a helper in the hierarchy to accelerate faster processing at an edge computing node and using helper models in the hierarchy of Noghabi using the training at a central location a helper model and pruned model for each layer of a hierarchy and distributing the helper model and pruned model to different levels of the hierarchy and directing, utilizing the helper model at each level of the hierarchy, inference generation to the pruned model or to another model at a higher tier of Teerapittayanon and the pruned model, wherein the pruned model is generated from a reduction process of the LLM of Kim, with a reasonable expectation of success, in order to provide a simpler and more principled approach. For distributed computing hierarchies, consisting of the cloud, the edge (fog) and geographically distributed end devices and in order to provide selecting from among multiple candidate generative models with differing computational efficiencies, a particular generative model to utilize in generating a response to the request and to reduce latency and/or conserve computational resource(s) and to mitigate occurrences of a generated response being inaccurate. Teerapittayanon, Section I. Introduction, pages 329, 328. Kim, para 4. This would have provided the advantages of centralized training with distributed computing hierarchies and of providing model selection from efficient generative large language models. Regarding claim 7, which depends from claim 1 and recites: wherein a caching strategy for local documents is performed. Noghabi in view of Teerapittayanon and Kim teaches the method of claim 1 from which claim 7 depends, including a method for hierarchical inference utilizing a large language model (LLM), the method comprising: training, at a central location, a helper model and a pruned model for each layer of a hierarchy, wherein the helper model is trained to classify a request as appropriate for the pruned model, wherein the pruned model is generated from a reduction process of the LLM, and distributing the helper model and pruned model to different levels of the hierarchy and directing, utilizing the helper model at each level of the hierarchy, inference generation to the pruned model or to another model at a higher tier. Noghabi teaches that, Another technical advantage of the systems and methods of the present disclosure is offline preprocessing of data for improving model results. The systems and methods integrate with various data sources and run compute pipelines in the background to update data sources stored on the edge. By using the offline processing of data, the systems and methods leverage the load between a user device at the frontline edge and the devices at the back office edge and the cloud. The systems and methods determine which data to bring to edge and when to update the data at each tier, so the data is available for use by language models at each tier of the edge architecture. The systems and methods of the present disclosure optimize the model selection and data processing for the user. At the edge, a single tenant user is providing the queries. The systems and methods may cache personalized data and documents for the user and may optimize the latency of the responses to the queries (wherein a caching strategy for local documents is performed). Noghabi, para 20, 29, 32. Claims 9-10 and 15 recite systems that parallel the methods of claims 1-2 and 7. Therefore, the analysis discussed above with respect to claims 1-2 and 7 also applies to claims 9-10 and 15, respectively. Accordingly, claims 9-10 and 15 are rejected based on substantially the same rationale as set forth above with respect to claims 1-2 and 7, respectively. More specifically regarding A system, the system comprising: a memory; and a processor coupled to the memory, wherein the processor performs operations, the operations comprising (i.e., Noghabi, Fig 4, para 65-73). Claims 17-18 and 23 recite computer program products that parallel the methods of claims 1-2 and 7. Therefore, the analysis discussed above with respect to claims 1-2 and 7 also applies to claims 17-18 and 23, respectively. Accordingly, claims 17-18 and 23 are rejected based on substantially the same rationale as set forth above with respect to claims 1-2 and 7, respectively. More specifically regarding A computer program product, the computer program product comprising a computer readable storage medium, wherein code stored in the computer readable storage medium when executed by a processor performs operations, the operations comprising (i.e., Noghabi, Fig 4, para 74-76, 65-73). Allowable Subject Matter Claims 3-6, 8, 11-14, 16, 19-22 and 24 would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and if the rejections as being indefinite are overcome and if the rejections as being directed to an abstract idea without significantly more are overcome. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US-20250265467-A1, US-20250371288-A1, US-20190140913-A1, US-20210357752-A1, US-20240249182-A1, US-20250103809-A1, US-20240320421-A1, US-20260004148-A1, US-12236193-B1, US-20250061334-A1, US-20250259047-A1, US-20240144051-A1. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BARBARA LEVEL whose telephone number is (303)297-4748. The examiner can normally be reached Monday through Friday 8:00 AM - 5:00 PM MT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BARBARA M LEVEL/ Examiner, Art Unit 2142
Read full office action

Prosecution Timeline

Jun 26, 2024
Application Filed
Sep 16, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731070
APPARATUS AND METHOD OF TRAINING MACHINE LEARNING MODEL, AND APPARATUS AND METHOD FOR SUMMARIZING DOCUMENT USING THE SAME
3y 10m to grant Granted Sep 08, 2026
Patent 12718116
METHOD AND SYSTEM FOR PRODUCING A SEMANTIC MAPPING OF SENSOR DATA
3y 5m to grant Granted Aug 25, 2026
Patent 12717875
AUTOMATED EXPLORATORY DATA ANALYSIS (EDA)
3y 6m to grant Granted Aug 25, 2026
Patent 12718110
LEARNING DEVICE, LEARNING METHOD, AND LEARNING PROGRAM
3y 1m to grant Granted Aug 25, 2026
Patent 12694954
METHOD FOR GENERATING SMALL MOLECULE BASED ON PHARMACOPHORE MODEL, DEVICE, AND MEDIUM
2y 11m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
72%
Grant Probability
99%
With Interview (+28.0%)
2y 8m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 348 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month