DETAILED ACTION
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is responsive to the amendment filed 06/26/2026.
Claims 1-28 are pending in this application.
Claim Rejections - 35 USC § 102
2. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-28 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Cai (US 20210073625). The reference was cited by Applicant in the IDS filed 02/06/2025.
It is noted that any citations to specific, pages, columns, paragraphs, lines, or figures in the prior art references and any interpretation of the reference should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. See MPEP 2123.
As to claim 1:
Cai teaches a processor-implemented method (Abstract: a method for adapting a computation graph of a machine learning model), comprising:
partitioning a graph representing a machine learning model into a plurality of subgraphs, each subgraph representing a portion of the machine learning model ( Figs.2-5, [0017]: provide a method and apparatus for adapting a computational graph, which can allow partitioning control dependency edges of a neural network model in an efficient way; [0043]: graph partitioner 320 is configured to partition a computation graph into a plurality of subgraphs... graph partitioner 320 can be configured to map the plurality of subgraphs onto multiple accelerators (e.g., target devices D1 to Dn in FIG. 2)...the computation graph to be divided by the graph partitioner 320 can be fed by the graph generator 310... the computation graph to be divided by the graph partitioner 320 can be a computation graph to which optimization techniques such as layer fusions, node clustering, etc. to maximize inference performance on accelerators have been applied; [0045]: graph partitioner 320 can partition a computation graph into multiple subgraphs that are executed on different accelerators based on the subgraph profiling information to optimize performance in executing the computation graph);
for each subgraph, simulating a plurality of execution paths based on permutations of using different processing unit types to execute portions of the subgraph and starting with each input source processing unit type selected from the different processing unit types ([0046]: graph partitioner 320 may take account of information including: 1) system and accelerator information, 2) operation profiling information per accelerator, and 3) subgraph profiling information per accelerator. The system information may include interconnect bandwidth information between accelerators or between a host unit and an accelerator. The accelerator information may include computing throughput information and memory bandwidth. The operation profiling information may include execution time or speed information and delay information of an accelerator for executing a certain operation such as a convolution, matrix multiplication, etc. The operation profiling information can be estimated by simulations or obtained by previous experiments on each of accelerators... operation profiling information for each of the accelerators can be stored for each of operations. The subgraph profiling information may include execution time or speed information and delay information for executing the subgraph on each accelerator. The subgraph profiling information can be estimated by simulations or obtained by previous experiments on each of accelerators);
for each subgraph, selecting an execution path from the plurality of execution paths having a lowest cost ([0045]: graph partitioner 320 can partition a computation graph into multiple subgraphs that are executed on different accelerators based on the subgraph profiling information to optimize performance in executing the computation graph... each subgraph can be assigned to a certain accelerator that can optimize performance of executing the subgraph); and
implementing the machine learning model based on the selected execution path for each subgraph ([0045]: graph partitioner 320 can partition a computation graph into multiple subgraphs that are executed on different accelerators based on the subgraph profiling information to optimize performance in executing the computation graph... each subgraph can be assigned to a certain accelerator that can optimize performance of executing the subgraph), wherein implementing the machine learning model comprises distributing the portions of the machine learning model over one or more processing units based on the selected execution path ([0036]: target devices D1 to Dn can be implemented as any one of CPU, GPU, FPGA. ASIC, etc. In some embodiments, at least two of the plurality of target devices D1 to Dn may have different processing speeds, power consumptions, transfer costs, etc. In some embodiments, a certain target device may be configured to be specialized to process a certain operation with high performance such as low cost and high accuracy; [0045]:a computation graph may include subgraphs that are commonly used in many machine learning models as their components. For example, the commonly used subgraphs can include MobileNets layers, ResNet layers, Region Proposal Network, etc. In some embodiments, prior history of execution, experiments, or simulations of a certain subgraph on accelerators can identify which accelerator is optimal for processing the certain subgraph. In some embodiments, each subgraph can be assigned to a certain accelerator that can optimize performance of executing the subgraph. [0048]: a first subgraph 421 and a second subgraph 422 can be mapped to different accelerators such as accelerators D1 and D2, respectively… the first subgraph 421 and the second subgraph 422 can be executed in parallel on different accelerators D1 and D2).
As to claim 2:
Cai teaches partitioning the graph representing the machine learning model into the plurality of subgraphs comprises partitioning the graph representing the machine learning model into a first set of subgraphs associated with a first processing unit type of the different processing unit types and a second set of subgraphs associated with a second processing unit type of the different processing unit types ([0044-0046]).
As to claim 3:
Cai teaches the first set of subgraphs and the second set of subgraphs comprise graphs generated based on a common fusion boundary across a first processing system associated with the first processing unit type and a second processing system associated with the second processing unit type ([0043-0044]).
As to claim 4:
Cai teaches the common fusion boundary comprises a point in the machine learning model at which a subgraph in the first set of subgraphs and a corresponding subgraph in the second set of subgraphs output a common output for ingestion into a subsequent portion of the machine learning model ([0043-0044]).
As to claim 5:
Cai teaches simulating the plurality of execution paths for each subgraph comprises simulating an execution time for executing operations identified in the subgraph including context switching time for transitions from a first processing unit type of the different processing unit types to a second processing unit type of the different processing unit types (Figs.4-5, [0036], [0046], and [0049]).
As to claim 6:
Cai teaches comprising modifying a subgraph from the plurality of subgraphs based on combining consecutive portions of a subgraph representing operations for which execution should remain with the same processing unit type, wherein the plurality of execution paths are simulated based on the modified subgraph (Figs.4-5, [0036], [0046], and [0049]).
As to claim 7:
Cai teaches selecting the execution path for each subgraph comprises selecting the execution path having the lowest cost for each subgraph based on a backwards traversal of a graph representing the simulated plurality of execution paths for the plurality of subgraphs representing the machine learning model (Figs.4-5 and [0036-0037]).
As to claim 8:
Cai teaches generating an inference using the implemented machine learning model based on an input into the implemented machine learning model and the selected execution path for each subgraph (Figs.4-5 and [0043-0044]).
As to claim 9:
Cai teaches the different processing unit types comprise two or more of a central processing unit (CPU), a graphics processing unit (GPU), or a neural processing unit (NPU) ([0028] and [0036]).
As to claims 10-18:
Refer to the discussion of claims 1-9 above, respectively, for rejections. Claims 10-18 are the same as claims 1-9, except claims 10-18 are system claims and claims 1-9 are method claims.
As to claims 19-27:
Refer to the discussion of claims 1-9 above, respectively, for rejections. Claims 19-27 are the same as claims 1-9, except claims 19-27 are system claims and claims 1-9 are method claims.
As to claim 28:
Refer to the discussion of claim 1 above for rejection. Claim 28 is the same as claim 1 except claim 28 is a non-transitory computer-readable medium claim and claim 1 is a method claim.
Response to Arguments
3. Applicants' arguments filed 06/26/2026 have been fully considered but they are not persuasive.
Claim Rejection under 35 USC § 112:
The 112 rejection has been withdrawn in view of Applicant's arguments.
Claim Rejection under 35 USC § 101:
The 101 rejection has been withdrawn in view of Applicant's arguments and amendments.
Claim Rejection under 35 USC § 102:
Applicant argues that Cai does not teach “for each subgraph, simulating a plurality of execution paths based on permutations of using different processing unit types to execute portions of the subgraph and starting with each input source processing unit type selected from the different processing unit types”.
In response, under broadest reasonable interpretation, Cai’s teaching “graph partitioner 320 may take account of information including: 1) system and accelerator information, 2) operation profiling information per accelerator, and 3) subgraph profiling information per accelerator. The system information may include interconnect bandwidth information between accelerators or between a host unit and an accelerator. The accelerator information may include computing throughput information and memory bandwidth. The operation profiling information may include execution time or speed information and delay information of an accelerator for executing a certain operation such as a convolution, matrix multiplication, etc. The operation profiling information can be estimated by simulations or obtained by previous experiments on each of accelerators... operation profiling information for each of the accelerators can be stored for each of operations. The subgraph profiling information may include execution time or speed information and delay information for executing the subgraph on each accelerator. The subgraph profiling information can be estimated by simulations or obtained by previous experiments on each of accelerators” [0046] reads-on the claimed “for each subgraph, simulating a plurality of execution paths based on permutations of using different processing unit types to execute portions of the subgraph and starting with each input source processing unit type selected from the different processing unit types”.
During patent examination, the pending claims must be “given their broadest reasonable interpretation consistent with the specification.” In re Hyatt 21 1 F.3d 1367, 1372, 54 USPQ2d 1664, 1667 (Fed. Cir. 2000).
Conclusion
3. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact Information
4. Any inquiry concerning this communication or earlier communications from the examiner should be directed to VAN H. NGUYEN whose telephone number is (571) 272-3765. The examiner can normally be reached on Monday- Friday from 9:00AM to 5:30 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, LEWIS BULLOCK, can be reached at telephone number (571) 272-3759. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from Patent Center and the Private Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from Patent Center or Private PAIR. Status information for unpublished applications is available through Patent Center or Private PAIR to authorized users only. Should you have questions about access to Patent Center or the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated- interview-request-air-form.
/VAN H NGUYEN/Primary Examiner, Art Unit 2199