Prosecution Insights
Last updated: October 02, 2026
Application No. 19/048,018

HYBRID INFERENCING USING AN ARTIFICIAL INTELLIGENCE OFFLOAD DIE IN A SYSTEM-IN-A-PACKAGE

Non-Final OA §103
Filed
Feb 07, 2025
Examiner
PRIFTI, AUREL
Art Unit
2175
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
83%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
527 granted / 635 resolved
+28.0% vs TC avg
Strong +22% interview lift
Without
With
+22.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
27 currently pending
Career history
655
Total Applications
across all art units

Statute-Specific Performance

§101
9.5%
-30.5% vs TC avg
§103
59.2%
+19.2% vs TC avg
§102
14.5%
-25.5% vs TC avg
§112
13.9%
-26.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 635 resolved cases

Office Action

§103
DETAILED ACTION Claims 1-20 are presented for examination. The present application is being examined under the AIA (America Invents Act) First Inventor to File. This Office Action is Non-Final. Claims 1, 7 and 14 are independent claims. Claims 2-6, 8-13, 15-20 are dependent claims. This action is responsive to the following communication: corresponding claims filed on 02-07-2025. Domestic Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, or 365(c) (International) is acknowledged. Information Disclosure Statement The information disclosure statement (IDS) submitted on 05-13-2026 is in compliance with the provisions of 37 CFR 1.97 Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3, 5, 7, 9, 11-12, 14-15, 17, 19 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Publication No. 2021/0019634 (hereinafter, “Pudipeddi”) in view of U.S. Publication No. 2026/0086846 (hereinafter, “O’Brien”). As per claim 1,Pudipeddi a method implemented in an artificial intelligence (Al) offload die that comprises a network controller, and that is communicatively coupled with a compute die in a system-in-a-package, the method comprising, during a hybrid inferencing of an Al model by a remote computing system and the compute die: identifying a portion of the Al model for use by the compute die; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k) using the network controller, ( network interface/modem for server 102/target device ) fetching the portion of the Al model from the remote computing system; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k) communicating the portion of the Al model to the compute die; and (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded using suitable interfaces, for example, “networks” (¶ [0048]) om parameter server 102 for processing to a target device 134a-k) synchronizing Al model inferencing state between the compute die and the remote computing system. (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 ) Pudipeddi does not distinctly disclose wherein an AI model workload is executed by using a network controller. However, O’Brien explicitly discloses wherein processing workload is offloaded to a network controller. (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs). It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi and O’Brien because both references are in the same field of endeavor. O’Brien’s teaching of offloading certain processing to SNICs would enhance Pudipeddi 's system to be more efficient during peak processing times, thus enhancing processing operations. As per claim 7, Pudipeddi discloses an artificial intelligence (Al) offload die comprising: a network controller; and (network interface/modem for server 102/target device ) an Al offloading engine configured, during a hybrid inferencing of an Al model by a remote computing system and a compute die, to: (parameter server 102 may include AI model manager 116 configured to manage AI model 106 during inference or training of AI model 106. AI model manager 116 includes computer program logic such as data manager 118, batch manager 120, transmitter 122 and output data manager 124 for managing AI model 106. Output data manager 124 is configured to receive and manage output data, among other data, from target devices 134a-134k, for use in the management of AI model 106; ¶ [0047]) identify a portion of the Al model for use by the compute die; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k) using the network controller, fetch the portion of the Al model from the remote computing system; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k) communicate the portion of the Al model to the compute die; and (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded using suitable interfaces, for example, “networks” (¶ [0048]) om parameter server 102 for processing to a target device 134a-k) synchronize Al model inferencing state between the compute die and the remote computing system. (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 ) Pudipeddi does not distinctly disclose wherein an AI model workload is executed by using a network controller. However, O’Brien explicitly discloses wherein processing workload is offloaded to a network controller. (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs). It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi and O’Brien because both references are in the same field of endeavor. O’Brien’s teaching of offloading certain processing to SNICs would enhance Pudipeddi 's system to be more efficient during peak processing times, thus enhancing processing operations. As per claim 14, Pudipeddi discloses a system-in-a-package, comprising: a compute die comprising a processor system and an artificial intelligence (AI) accelerator; (parameter server may comprise processor 132 and FPGA; ¶ [0046]) a memory; and (memory 104; Fig 1) an Al offload die, comprising: (model manager 116 having logic/scheduler…etc; ¶ [0047]) a network controller; and (network interface/modem for server 102/target device ) an Al offloading engine configured, during a hybrid inferencing of an Al model by a remote computing system and the compute die, to: (parameter server 102 may include AI model manager 116 configured to manage AI model 106 during inference or training of AI model 106. AI model manager 116 includes computer program logic such as data manager 118, batch manager 120, transmitter 122 and output data manager 124 for managing AI model 106. Output data manager 124 is configured to receive and manage output data, among other data, from target devices 134a-134k, for use in the management of AI model 106; ¶ [0046]) identify a portion of the Al model for use by the compute die; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k) using the network controller, fetch the portion of the Al model from the remote computing system; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k) communicate the portion of the Al model to the compute die; and (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded using suitable interfaces, for example, “networks” (¶ [0048]) om parameter server 102 for processing to a target device 134a-k) synchronize Al model inferencing state between the compute die and the remote computing system. (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 ) Pudipeddi does not distinctly disclose wherein an AI model workload is executed by using a network controller. However, O’Brien explicitly discloses wherein processing workload is offloaded to a network controller. (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs). It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi and O’Brien because both references are in the same field of endeavor. O’Brien’s teaching of offloading certain processing to SNICs would enhance Pudipeddi 's system to be more efficient during peak processing times, thus enhancing processing operations. As per claim(s) 3, 9, 17, Pudipeddi as modified discloses a method wherein the network controller fetches the portion of the Al model into at least one of a first memory in the Al offload die, a second memory in the system-in-a- package that is separate from the Al offload die, or a third memory in the compute die. (Pudipeddi: parameter server having memory 104 and each of the target devices having memory 142; Fig 1) As per claims 5, 12, 19, Pudipeddi as modified discloses a method wherein: the Al offload die is communicatively coupled with a plurality of compute dies in the system-in-a- package; and the method comprises(Pudipeddi: Fig 1/parameter server 102, target devices 134a-134k, parameter server 904 and target devices 906a-906n, and flowcharts 200, 300, 600-800, and/or 1100-1600 may be implemented together in a SoC. ; ¶ [0122]) identifying a plurality of portions of the Al model, each portion corresponding to one the plurality of compute dies; (Pudipeddi: a batch manager configured to determine a microbatch size suitable for the target device; ¶[008]) using the network controller, fetching the plurality of portions of the Al model from the remote computing system;communicating each portion of the Al model to its corresponding compute die of the plurality of compute dies; and (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 ) & (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs).) using the network controller, synchronizing the Al model inferencing state between the plurality of compute dies and the remote computing system. (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 ) & (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs).) As per claim 11, Pudipeddi as modified discloses wherein the Al offload die is in a system-in-a-package that also comprises the compute die. (Pudipeddi: SOC ¶ [0122]) As per claim 15, Pudipeddi as modified discloses wherein the Al accelerator in the compute die is one of a neural processing unit (NPU), a tensor processing unit (TPU), or a graphics processing unit (GPU). (Pudipeddi: GPU ¶ [0038]) Claim(s) 2, 4, 8, 10, 16, 18 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Publication No. 2021/0019634 (hereinafter, “Pudipeddi”) in view of U.S. Publication No. 2026/0086846 (hereinafter, “O’Brien”) and further view of U.S. Publication No. 2020/0334195 (hereinafter, “Chen”). As per claims 2, 8, 16, Pudipeddi as modified does not distinctly disclose a method wherein the network controller communicates with the remote computing system using remote direct access memory (RDMA). However, Chen explicitly discloses a method wherein the network controller communicates with the remote computing system using remote direct access memory (RDMA). It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi as modified and Chen because all references are in the same field of endeavor. Chen’s teaching of RDMA system would enhance Pudipeddi 's as modified system by moving data between memories more efficiently, thus enhancing data speed. As per claims 4, 10, 18, Pudipeddi as modified discloses a method wherein, the network controller is a first network controller; the system-in-a-package or the compute die comprises a second network controller; (Pudipeddi: parameter server having communication interface/modem and each of the target devices also having communication interface/modem) & (Chen: NIC 113 and the NIC 123 may establish an RDMA connection with each other via the plurality of network paths 140, so as to enable multi-path RDMA operations between the host 110 and the host 120 ¶ [0033] ) the first network controller is inaccessible by an operating system (OS) executing at the compute die; and (Chen: Remote Direct Memory Access (RDMA) implements the entire transport logic in a network interface card (NIC) and allows a direct access to a remote memory without involvement of a Central Processing Unit (CPU) or an operation system; ¶ [0001) the second network controller is accessible by the OS executing at the compute die. (Pudipeddi: accessing data without using RDMA, thus to a PHOSITA accessing data via OS) It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi as modified and Chen because all references are in the same field of endeavor. Chen’s teaching of RDMA system would enhance Pudipeddi 's as modified system by moving data between memories more efficiently, thus enhancing data speed. Claim(s) 6, 13, 20 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Publication No. 2021/0019634 (hereinafter, “Pudipeddi”) in view of U.S. Publication No. 2026/0086846 (hereinafter, “O’Brien”) and further view of U.S. Publication No. 2024/0184997 (hereinafter, “SHOU”). As per claims 6, 13, 20 Pudipeddi as modified discloses wherein synchronizing the Al model inferencing state includes one or more of, aligning [parameters ] between the compute die and the remote computing system, or initiating a roll-back of a speculative inference at the compute die based on a token misalignment at the remote computing system. (Pudipedd: synchronization between different processing entities ¶ [0040], [0074] ) Pudipeddi as modified does not distinctly disclose synchronizing includes aligning tokens. However, SHOU explicitly discloses synchronizing includes aligning tokens. ¶ [0024] It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi as modified and SHOU because all references are in the same field of endeavor. SHOU’s teaching of aligning tokens would enhance Pudipeddi 's as modified system by improving processing efficiency. Conclusion With respect to any newly added or amended claims, applicant should show support in the original disclosure for the new or amended claims. See MPEP §714.02 and § 2163.06. For example, when responding to this office action, applicants are advised to provide the examiner with the line numbers and page numbers in the application and/or references cited to assist the examiner in locating appropriate paragraphs. Any inquiry concerning this communication or earlier communications from the examiner should be directed to AUREL PRIFTI whose telephone number is (571)270-1743. The examiner can normally be reached on M-F 8 a.m.- 6 p.m.. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew J. Jung can be reached on 571-270-3779. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AUREL PRIFTI/Primary Examiner, Art Unit 2175 Aurel Prifti Primary Examiner Art Unit 2175 Tel. (571) 270-1743 Fax (571) 270-2743 aurel.prifti@uspto.gov
Read full office action

Prosecution Timeline

Feb 07, 2025
Application Filed
Aug 12, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743147
SLEEP MODE MANAGEMENT OF HARD DISK DRIVE
1y 10m to grant Granted Sep 22, 2026
Patent 12737469
FIRMWARE-BASED SECURE TENANCY TRANSFER
2y 5m to grant Granted Sep 15, 2026
Patent 12730473
PROCESSOR CLOCK SCALING TECHNIQUE
2y 6m to grant Granted Sep 08, 2026
Patent 12730652
INFORMATION PROCESSING APPARATUS AND CONTROL METHOD
2y 1m to grant Granted Sep 08, 2026
Patent 12724471
CONVERTOR CIRCUIT AND FAILURE REPORTING METHOD
2y 5m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
83%
Grant Probability
99%
With Interview (+22.1%)
2y 6m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 635 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month