DETAILED ACTION
Claims 1-20 are presented for examination.
The present application is being examined under the AIA (America Invents Act) First Inventor to File.
This Office Action is Non-Final.
Claims 1, 7 and 14 are independent claims. Claims 2-6, 8-13, 15-20 are dependent claims.
This action is responsive to the following communication: corresponding claims filed on 02-07-2025.
Domestic Priority
Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, or 365(c) (International) is acknowledged.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05-13-2026 is in compliance with the provisions of 37 CFR 1.97
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3, 5, 7, 9, 11-12, 14-15, 17, 19 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Publication No. 2021/0019634 (hereinafter, “Pudipeddi”) in view of U.S. Publication No. 2026/0086846 (hereinafter, “O’Brien”).
As per claim 1,Pudipeddi a method implemented in an artificial intelligence (Al) offload die that comprises a network controller, and that is communicatively coupled with a compute die in a system-in-a-package, the method comprising, during a hybrid inferencing of an Al model by a remote computing system and the compute die:
identifying a portion of the Al model for use by the compute die; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k)
using the network controller, ( network interface/modem for server 102/target device ) fetching the portion of the Al model from the remote computing system; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k)
communicating the portion of the Al model to the compute die; and (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded using suitable interfaces, for example, “networks” (¶ [0048]) om parameter server 102 for processing to a target device 134a-k)
synchronizing Al model inferencing state between the compute die and the remote computing system. (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 )
Pudipeddi does not distinctly disclose wherein an AI model workload is executed by using a network controller.
However, O’Brien explicitly discloses wherein processing workload is offloaded to a network controller. (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs).
It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi and O’Brien because both references are in the same field of endeavor. O’Brien’s teaching of offloading certain processing to SNICs would enhance Pudipeddi 's system to be more efficient during peak processing times, thus enhancing processing operations.
As per claim 7, Pudipeddi discloses an artificial intelligence (Al) offload die comprising:
a network controller; and (network interface/modem for server 102/target device )
an Al offloading engine configured, during a hybrid inferencing of an Al model by a remote computing system and a compute die, to: (parameter server 102 may include AI model manager 116 configured to manage AI model 106 during inference or training of AI model 106. AI model manager 116 includes computer program logic such as data manager 118, batch manager 120, transmitter 122 and output data manager 124 for managing AI model 106. Output data manager 124 is configured to receive and manage output data, among other data, from target devices 134a-134k, for use in the management of AI model 106; ¶ [0047])
identify a portion of the Al model for use by the compute die; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k)
using the network controller, fetch the portion of the Al model from the remote computing system; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k)
communicate the portion of the Al model to the compute die; and (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded using suitable interfaces, for example, “networks” (¶ [0048]) om parameter server 102 for processing to a target device 134a-k)
synchronize Al model inferencing state between the compute die and the remote computing system. (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 )
Pudipeddi does not distinctly disclose wherein an AI model workload is executed by using a network controller.
However, O’Brien explicitly discloses wherein processing workload is offloaded to a network controller. (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs).
It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi and O’Brien because both references are in the same field of endeavor. O’Brien’s teaching of offloading certain processing to SNICs would enhance Pudipeddi 's system to be more efficient during peak processing times, thus enhancing processing operations.
As per claim 14, Pudipeddi discloses a system-in-a-package, comprising:
a compute die comprising a processor system and an artificial intelligence (AI) accelerator; (parameter server may comprise processor 132 and FPGA; ¶ [0046])
a memory; and (memory 104; Fig 1)
an Al offload die, comprising: (model manager 116 having logic/scheduler…etc; ¶ [0047])
a network controller; and (network interface/modem for server 102/target device )
an Al offloading engine configured, during a hybrid inferencing of an Al model by a remote computing system and the compute die, to: (parameter server 102 may include AI model manager 116 configured to manage AI model 106 during inference or training of AI model 106. AI model manager 116 includes computer program logic such as data manager 118, batch manager 120, transmitter 122 and output data manager 124 for managing AI model 106. Output data manager 124 is configured to receive and manage output data, among other data, from target devices 134a-134k, for use in the management of AI model 106; ¶ [0046])
identify a portion of the Al model for use by the compute die; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k)
using the network controller, fetch the portion of the Al model from the remote computing system; (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded from parameter server 102 for processing to a target device 134a-k)
communicate the portion of the Al model to the compute die; and (Fig’s 1-2 illustrates how a portion of an artificial intelligence (AI) model is downloaded using suitable interfaces, for example, “networks” (¶ [0048]) om parameter server 102 for processing to a target device 134a-k)
synchronize Al model inferencing state between the compute die and the remote computing system. (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 )
Pudipeddi does not distinctly disclose wherein an AI model workload is executed by using a network controller.
However, O’Brien explicitly discloses wherein processing workload is offloaded to a network controller. (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs).
It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi and O’Brien because both references are in the same field of endeavor. O’Brien’s teaching of offloading certain processing to SNICs would enhance Pudipeddi 's system to be more efficient during peak processing times, thus enhancing processing operations.
As per claim(s) 3, 9, 17, Pudipeddi as modified discloses a method wherein the network controller fetches the portion of the Al model into at least one of a first memory in the Al offload die, a second memory in the system-in-a- package that is separate from the Al offload die, or a third memory in the compute die. (Pudipeddi: parameter server having memory 104 and each of the target devices having memory 142; Fig 1)
As per claims 5, 12, 19, Pudipeddi as modified discloses a method wherein:
the Al offload die is communicatively coupled with a plurality of compute dies in the system-in-a- package; and the method comprises(Pudipeddi: Fig 1/parameter server 102, target devices 134a-134k, parameter server 904 and target devices 906a-906n, and flowcharts 200, 300, 600-800, and/or 1100-1600 may be implemented together in a SoC. ; ¶ [0122])
identifying a plurality of portions of the Al model, each portion corresponding to one the plurality of compute dies; (Pudipeddi: a batch manager configured to determine a microbatch size suitable for the target device; ¶[008])
using the network controller, fetching the plurality of portions of the Al model from the remote computing system;communicating each portion of the Al model to its corresponding compute die of the plurality of compute dies; and (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 ) & (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs).)
using the network controller, synchronizing the Al model inferencing state between the plurality of compute dies and the remote computing system. (¶ [0074] discloses how a transmitter 122 may transmit a portion of AI model 106 to target device 134a after target device 134a finishes the current portion to avoid the need to buffer the portion. In this alternate example, target device 134a may perform synchronization after execution of the current portion before receiving the portion. Moreover, the execution portion of the AI model may set of microbatches that are used for inference; claims 9-10 ) & (offloading certain operations of an application to one or more Smart Network Interface Controllers (SNICs).)
As per claim 11, Pudipeddi as modified discloses wherein the Al offload die is in a system-in-a-package that also comprises the compute die. (Pudipeddi: SOC ¶ [0122])
As per claim 15, Pudipeddi as modified discloses wherein the Al accelerator in the compute die is one of a neural processing unit (NPU), a tensor processing unit (TPU), or a graphics processing unit (GPU). (Pudipeddi: GPU ¶ [0038])
Claim(s) 2, 4, 8, 10, 16, 18 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Publication No. 2021/0019634 (hereinafter, “Pudipeddi”) in view of U.S. Publication No. 2026/0086846 (hereinafter, “O’Brien”) and further view of U.S. Publication No. 2020/0334195 (hereinafter, “Chen”).
As per claims 2, 8, 16, Pudipeddi as modified does not distinctly disclose a method wherein the network controller communicates with the remote computing system using remote direct access memory (RDMA).
However, Chen explicitly discloses a method wherein the network controller communicates with the remote computing system using remote direct access memory (RDMA).
It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi as modified and Chen because all references are in the same field of endeavor. Chen’s teaching of RDMA system would enhance Pudipeddi 's as modified system by moving data between memories more efficiently, thus enhancing data speed.
As per claims 4, 10, 18, Pudipeddi as modified discloses a method wherein, the network controller is a first network controller; the system-in-a-package or the compute die comprises a second network controller; (Pudipeddi: parameter server having communication interface/modem and each of the target devices also having communication interface/modem) & (Chen: NIC 113 and the NIC 123 may establish an RDMA connection with each other via the plurality of network paths 140, so as to enable multi-path RDMA operations between the host 110 and the host 120 ¶ [0033] )
the first network controller is inaccessible by an operating system (OS) executing at the compute die; and (Chen: Remote Direct Memory Access (RDMA) implements the entire transport logic in a network interface card (NIC) and allows a direct access to a remote memory without involvement of a Central Processing Unit (CPU) or an operation system; ¶ [0001)
the second network controller is accessible by the OS executing at the compute die. (Pudipeddi: accessing data without using RDMA, thus to a PHOSITA accessing data via OS)
It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi as modified and Chen because all references are in the same field of endeavor. Chen’s teaching of RDMA system would enhance Pudipeddi 's as modified system by moving data between memories more efficiently, thus enhancing data speed.
Claim(s) 6, 13, 20 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Publication No. 2021/0019634 (hereinafter, “Pudipeddi”) in view of U.S. Publication No. 2026/0086846 (hereinafter, “O’Brien”) and further view of U.S. Publication No. 2024/0184997 (hereinafter, “SHOU”).
As per claims 6, 13, 20 Pudipeddi as modified discloses wherein synchronizing the Al model inferencing state includes one or more of, aligning [parameters ] between the compute die and the remote computing system, or initiating a roll-back of a speculative inference at the compute die based on a token misalignment at the remote computing system. (Pudipedd: synchronization between different processing entities ¶ [0040], [0074] )
Pudipeddi as modified does not distinctly disclose synchronizing includes aligning tokens.
However, SHOU explicitly discloses synchronizing includes aligning tokens. ¶ [0024]
It would have been obvious before the effective filing date of the claimed invention to modify the teachings of Pudipeddi as modified and SHOU because all references are in the same field of endeavor. SHOU’s teaching of aligning tokens would enhance Pudipeddi 's as modified system by improving processing efficiency.
Conclusion
With respect to any newly added or amended claims, applicant should show support in the original disclosure for the new or amended claims. See MPEP §714.02 and § 2163.06. For example, when responding to this office action, applicants are advised to provide the examiner with the line numbers and page numbers in the application and/or references cited to assist the examiner in locating appropriate paragraphs.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AUREL PRIFTI whose telephone number is (571)270-1743. The examiner can normally be reached on M-F 8 a.m.- 6 p.m..
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew J. Jung can be reached on 571-270-3779. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AUREL PRIFTI/Primary Examiner, Art Unit 2175
Aurel Prifti
Primary Examiner
Art Unit 2175
Tel. (571) 270-1743
Fax (571) 270-2743
aurel.prifti@uspto.gov