Prosecution Insights
Last updated: October 02, 2026
Application No. 18/582,503

COMPILER BASED DYNAMIC SCALING POWER MANAGEMENT

Final Rejection §103§112
Filed
Feb 20, 2024
Examiner
SOLTANZADEH, AMIR
Art Unit
2191
Tech Center
2100 — Computer Architecture & Software
Assignee
Google LLC
OA Round
2 (Final)
81%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 81% — above average
81%
Career Allowance Rate
351 granted / 434 resolved
+25.9% vs TC avg
Strong +17% interview lift
Without
With
+16.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
30 currently pending
Career history
477
Total Applications
across all art units

Statute-Specific Performance

§101
16.4%
-23.6% vs TC avg
§103
66.3%
+26.3% vs TC avg
§102
2.1%
-37.9% vs TC avg
§112
9.9%
-30.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 434 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-24 are presented for examination. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a compiler configured to…” in claim 19, “the compiler is configured to” in claims 22-23. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Applicant’s argument that this interpretation should be withdrawn is addressed in the Response to Arguments section below. The 112(f) interpretation is maintained. It is noted that this interpretation does not, by itself, give rise to a rejection under 35 U.S.C. 112(b), because the specification discloses adequate corresponding structure for the recited function(s). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 6-8, 10-17, 19-24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Paul (US 2023/0273832 A1) in view of Mohammad (US 2021/0124404 A1). Regarding Claim 1, Paul (US 2023/0273832 A1) teaches A computer-implemented method, comprising: receiving uncompiled code for a computer program (Para [0016], "a compiler or performance simulator, which may be external to the VPU 100, may compile and analyze the neural network, layer-by-layer, to determine which layers are compute-intensive and which are memory-intensive") Examiner Comments: Paul teaches a compiler that receives a neural network model description (i.e., uncompiled code) and compiles and analyzes it for execution on the VPU accelerator, which corresponds to receiving uncompiled code for a computer program; compiling the uncompiled code for the computer program to generate compiled code for the computer program (Para [0020], "the layers of the neural network may be compiled by a compiler connected to, communicatively coupled to, or the like, the VPU of an ML accelerator"; see also Para [0016] and FIG. 3, Operation 300) Examiner Comments: Paul teaches compiling the neural network description into a VPU binary image (compiled code) for execution on the accelerator; retrieving, during the compiling of the uncompiled code, a respective workload for each of the plurality of processing devices defining an analysis of processor utilization for the respective portion of the uncompiled code that is to be executed by the processing device (Para [0016], "a compiler or performance simulator, which may be external to the VPU 100, may compile and analyze the neural network, layer-by-layer, to determine which layers are compute-intensive and which are memory-intensive"; Para [0017], "for any deep neural network or ML workload, it may be possible to fully estimate the execution latency and energy use or requirement of the layers of the neural network … and determine whether the performance of the layer … is limited by compute or memory performance. From the estimation, fractions of the total power budget available to the VPU 100 may be allocated to the compute and memory units"; Para [0028], "… based at least in part on the determination of Fcompute and FCMX derived during compilation of the neural network") Examiner Comments: Paul teaches that the workload analysis, i.e., the layer-by-layer determination of compute versus memory utilization (an analysis of processor utilization), is performed by the compiler during compilation of the neural network, as expressly confirmed by Para [0028] reciting that the compute and memory frequency determinations are “derived during compilation of the neural network,” thereby teaching retrieving a respective workload during the compiling of the uncompiled code; injecting a power state instruction into a respective portion of the compiled code corresponding the respective portion of the uncompiled code, that is to be executed by each of the plurality of processing devices, wherein the power state instruction identifies a voltage setting and a frequency setting that the processing device is instructed to apply when executing the respective portion of the compiled code of the computer program (Para [0011], "hooks inserted by the compilation tool to create a command image for the underlying accelerator may exploit the workload phases for better energy efficiency, enabling power-aware compilation"; Para [0028], "the local power-management unit on the VPU … may adjust the voltage and frequency ratio at the compute and memory layers … based at least in part on the determination of Fcompute and FCMX derived during compilation of the neural network. The determination of Fcompute and FCMX may be passed to the local PMU at runtime through the use of a special power-management instruction from the compiler") Examiner Comments: Paul explicitly teaches that the compiler inserts a special power-management instruction (i.e., power state instruction) into the compiled code that sets voltage and frequency for the VPU accelerator at runtime based on the workload analysis of each layer derived during compilation. Paul did not specifically teach the compiled code that is targeted for distributed execution by a plurality of processing devices in a distributed computing system, wherein the compiling includes identifying, for each processing device of the plurality of processing devices, a respective portion of the uncompiled code that is to be executed by the processing device. However, Mohammad (US 2021/0124404 A1) teaches compiled code that is targeted for distributed execution by a plurality of processing devices in a distributed computing system (Abstract, "A scheme to improve performance of power-constrained computers, comprising a heterogeneous mix of compute elements, by dynamically reacting to changes in the switching capacitance that present workload induces in each heterogeneous compute element"; Para [0010], "An example of a heterogeneous mix of compute elements is a server blade comprising a mix of CPUs and GPUs. Two other examples include a server rack comprising a mix of CPU and GPU blades and an individual SoC (system-on-chip) comprising a mix of CPU cores, GPU cores, or accelerator cores") Examiner Comments: Mohammad teaches a distributed computing system comprising multiple heterogeneous processing devices (CPUs, GPUs, accelerator cores) in server blades and server racks, where workloads are distributed across the processing devices for execution; wherein the compiling includes identifying, for each processing device of the plurality of processing devices, a respective portion of the uncompiled code that is to be executed by the processing device (Para [0014], "This MIMO control system monitors power consumption of each compute element under its control and maximizes each element’s (e.g., a CPU) frequency subject to a constraint on the overall power consumed across all compute elements under control") Examiner Comments: Mohammad’s MIMO control system independently identifies and manages the workload assigned to each of the plurality of heterogeneous compute elements, where each compute element executes a respective portion of the overall workload, teaching identification of code portions for each processing device. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul’s compiler-based power management technique with Mohammad’s distributed workload-aware power limiting across multiple heterogeneous compute elements in order to extend the energy efficiency benefits of compiler-based per-layer DVFS to distributed computing environments comprising multiple heterogeneous processing devices, thereby improving overall system performance under power constraints (Mohammad, Abstract). Both Paul and Mohammad are directed to workload-aware power management for processing devices, and applying Paul’s per-layer voltage/frequency adjustments to each device in Mohammad’s distributed system would yield the predictable result of improved energy efficiency across the distributed computing system. Regarding Claim 2, Paul and Mohammad teach The method of Claim 1. Paul further teaches wherein a first processing device and a second processing device of the plurality of processing devices are targeted to execute different portions of the compiled code with a same voltage setting and a same frequency setting according to the power state instruction (Para [0016], "a compiler or performance simulator, which may be external to the VPU 100, may compile and analyze the neural network, layer-by-layer, to determine which layers are compute-intensive and which are memory-intensive") Examiner Comments: When two processing devices execute code portions with similar workload characteristics (e.g., both compute-intensive layers), Paul’s system applies the same voltage and frequency settings to both devices. Regarding Claim 3, Paul and Mohammad teach The method of Claim 1. Paul and Mohammad further teach wherein a first processing device and a second processing device of the plurality of processing devices are targeted to execute different portions of the compiled code with different voltage and frequency settings according to the power state instruction (Paul, Para [0016] (different layers have different compute vs. memory intensity, resulting in different voltage and frequency settings); Mohammad, Para [0023], "This MIMO control system monitors power consumption of each compute element under its control and maximizes each element’s (e.g., a CPU) frequency subject to a constraint on the overall power consumed across all compute elements under control") Examiner Comments: The combination teaches that different processing devices executing portions with different workload characteristics (e.g., one compute-intensive, one memory-intensive) receive different voltage/frequency settings. Regarding Claim 6, Paul and Mohammad teach The method of Claim 1. Paul further teaches wherein a plurality of power state instructions are injected into the compiled code for execution by the plurality of processing devices in parallel (Paul, Para [0011], "hooks inserted by the compilation tool to create a command image for the underlying accelerator may exploit the workload phases") Examiner Comments: Paul’s multiple per-layer power-management instructions combined with Mohammad’s parallel execution across multiple devices teaches injecting a plurality of power state instructions for parallel execution. Regarding Claim 7, Paul and Mohammad teach The method of Claim 1. Paul further teaches wherein a workload is defined for high level operations included in the uncompiled code (Para [0016], "compile and analyze the neural network, layer-by-layer, to determine which layers are compute-intensive and which are memory-intensive") Examiner Comments: Paul’s layer-by-layer analysis defines workloads at the operation/layer level of the neural network, where layers constitute the high level operations of the uncompiled code. Regarding Claim 8, Paul and Mohammad teach The method of Claim 7. Paul further teaches wherein retrieving the respective workload for each of the plurality of processing devices includes calculating the respective workload for a processing device by identifying one or more operations in the respective portion of uncompiled code set for execution by the processing device and retrieving a workload associated with each of the identified one or more operations (Para [0016], "The determination of whether the phases or layers of the neural network are compute-intensive or memory-intensive may be based on whether the performance of a layer is limited, bound, constrained, or the like, by the compute bandwidth or the memory bandwidth of the VPU IP") Examiner Comments: Paul teaches identifying individual layers (operations) in the neural network code and determining the associated workload (compute vs. memory bound characteristics) for each identified operation. Regarding Claim 10, Paul and Mohammad teach The method of Claim 7. Paul and Mohammad further teach monitoring an execution of an operation on a processing device of the plurality of processing devices (Paul, Para [0014], "Compiler-based static prediction may be combined with dynamic telemetry obtained from performance counters and temperature sensors or voltage droop circuit sensors"; Mohammad, Para [0014], "This MIMO control system monitors power consumption of each compute element under its control") Examiner Comments: Both Paul and Mohammad teach monitoring the execution of operations on processing devices using performance counters and telemetry; determining the workload for the operation based on the monitoring of the execution of the operation (Mohammad, Abstract, "dynamically reacting to changes in the switching capacitance that present workload induces in each heterogeneous compute element and learning the coefficients of a power-frequency model for each compute element for the present workload") Examiner Comments: Mohammad teaches determining workload characteristics from monitored execution data for each compute element. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul’s compiler-based power management technique with Mohammad’s distributed workload-aware power limiting across multiple heterogeneous compute elements in order to extend the energy efficiency benefits of compiler-based per-layer DVFS to distributed computing environments comprising multiple heterogeneous processing devices, thereby improving overall system performance under power constraints (Mohammad, Abstract). Both Paul and Mohammad are directed to workload-aware power management for processing devices, and applying Paul’s per-layer voltage/frequency adjustments to each device in Mohammad’s distributed system would yield the predictable result of improved energy efficiency across the distributed computing system. Regarding Claim 11, Paul and Mohammad teach The method of Claim 10. Mohammad further teaches wherein the workload for the operation is updated dynamically based on the monitoring of the execution of the operation on the processing device (Mohammad, Abstract, "The scheme rapidly re-learns coefficients of the power model and rapidly adapts the frequency as the workload’s characteristics shift ensuring that compute elements run at the maximum frequency they can while not exceeding the input power limit") Examiner Comments: Mohammad explicitly teaches dynamically updating the workload model as execution characteristics change, enabling adaptive frequency control. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul’s compiler-based power management technique with Mohammad’s distributed workload-aware power limiting across multiple heterogeneous compute elements in order to extend the energy efficiency benefits of compiler-based per-layer DVFS to distributed computing environments comprising multiple heterogeneous processing devices, thereby improving overall system performance under power constraints (Mohammad, Abstract). Both Paul and Mohammad are directed to workload-aware power management for processing devices, and applying Paul’s per-layer voltage/frequency adjustments to each device in Mohammad’s distributed system would yield the predictable result of improved energy efficiency across the distributed computing system. Regarding Claim 12, Paul and Mohammad teach The method of Claim 1. Paul and Mohammad further teach wherein injecting the power state instruction into the compiled code is further based on a cost model (Paul, Para [0024], "Operation 306 may include determining the optimal compute and memory power for each layer to send to the VPU. The determination may include a voltage and/or frequency to send to the DPU 110 and/or the CMX 112 … a determination may be made of the optimal power level or frequency for the DPU 110, which may lead to the lowest energy within a performance degradation threshold for execution of the neural network"; Mohammad, Abstract, "learning the coefficients of a power-frequency model for each compute element for the present workload") Examiner Comments: Both Paul’s optimization model for determining optimal power levels and Mohammad’s learned power-frequency model constitute cost models used to determine power state settings. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul’s compiler-based power management technique with Mohammad’s distributed workload-aware power limiting across multiple heterogeneous compute elements in order to extend the energy efficiency benefits of compiler-based per-layer DVFS to distributed computing environments comprising multiple heterogeneous processing devices, thereby improving overall system performance under power constraints (Mohammad, Abstract). Both Paul and Mohammad are directed to workload-aware power management for processing devices, and applying Paul’s per-layer voltage/frequency adjustments to each device in Mohammad’s distributed system would yield the predictable result of improved energy efficiency across the distributed computing system. Regarding Claim 13, Paul and Mohammad teach The method of Claim 12. Paul further teaches predict locations in the compiled code to inject power state instructions (Paul, Para [0016], "a compiler-based proactive prediction of compute and memory bound, or intensive, phases or layers of neural network may be employed by the VPU … a compiler or performance simulator … may compile and analyze the neural network, layer-by-layer, to determine which layers are compute-intensive and which are memory-intensive") Examiner Comments: Paul’s layer-by-layer analysis predicts where in the compiled code to inject power state instructions; predict a voltage setting and a frequency setting for each of the power state instructions (Paul, Para [0017], "it may be possible to fully estimate the execution latency and energy use or requirement of the layers of the neural network … From the estimation, fractions of the total power budget available to the VPU 100 may be allocated to the compute and memory units. This fraction may be represented by a ratio of Fcompute and FCMX which may be passed to the hardware of the VPU 100 at runtime") Examiner Comments: Both references teach predicting the voltage/frequency settings for the power state instructions based on the cost model. Regarding Claim 14, Paul and Mohammad teach The method of Claim 13. Paul and Mohammad further teach wherein the cost model is trained using monitored device traces of a processing device executing high level operations (Paul, Para [0014], "dynamic telemetry obtained from performance counters and temperature sensors"; Mohammad, Abstract, "learning the coefficients of a power-frequency model for each compute element for the present workload") Examiner Comments: Mohammad’s power-frequency model is learned/trained using monitored device execution data (performance counters and telemetry), which constitutes device traces from processing devices executing operations. Regarding Claim 15, Paul and Mohammad teach The method of Claim 12. Mohammad further teaches wherein predictions made by the cost model are further based at least in part on utility costs (Mohammad, Abstract, "the scheme forecasts a maximum frequency that the compute element can run at without exceeding an input power limit for a given workload. The scheme rapidly re-learns coefficients of the power model and rapidly adapts the frequency as the workload’s characteristics shift ensuring that compute elements run at the maximum frequency they can while not exceeding the input power limit") Examiner Comments: Mohammad teaches that the cost model operates subject to external power constraints (input power limits); the input power limit is an external constraint that can incorporate utility/energy cost considerations. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul's compiler-based power management technique with Mohammad's distributed workload-aware power limiting across multiple heterogeneous compute elements in order to extend the energy efficiency benefits of compiler-based per-layer DVFS to distributed computing environments comprising multiple heterogeneous processing devices, thereby improving overall system performance under power constraints (Mohammad, Abstract). Both Paul and Mohammad are directed to workload-aware power management for processing devices, and applying Paul's per-layer voltage/frequency adjustments to each device in Mohammad's distributed system would yield the predictable result of improved energy efficiency across the distributed computing system. Regarding Claim 16, Paul and Mohammad teach The method of Claim 12. Paul further teaches wherein the respective workload for each of the plurality of processing devices is defined by at least one of: (a) floating-point operations per second (FLOP) utilization; (b) memory bandwidth utilization; or (c) any combination of (a) and (b) (Paul, Para [0016], "whether the performance of a layer is limited, bound, constrained, or the like, by the compute bandwidth or the memory bandwidth of the VPU IP"; see also FIG. 2 showing FLOPS profile 204) Examiner Comments: Paul explicitly shows FLOPS utilization profiles (FLOP utilization) and memory bandwidth utilization as workload metrics in the analysis, directly teaching that the workload is defined by FLOP utilization, memory bandwidth utilization, or a combination thereof. Regarding Claim 17, Paul and Mohammad teach The method of Claim 1. Paul further teaches wherein the computer program is for training a machine-learning model (Abstract, "A system for autonomous and proactive power management for energy efficient execution of machine learning workloads") Examiner Comments: Paul explicitly discloses execution of machine learning workloads, which encompasses training of machine-learning models. Regarding Claim 19, is a system claim corresponding to the method claim above (Claim 1) and, therefore, is rejected for the same reasons set forth in the rejection of claim 1. Regarding Claim 20, Paul and Mohammad teach The system of Claim 19. Mohammad further teaches wherein at least some of the compiled code is targeted to execute synchronously across the plurality of computing devices, and the compiler sets the voltage setting and the frequency setting in the power state instruction according to a sampled workload of the plurality of computing devices (Mohammad, Para [0023], "This MIMO control system monitors power consumption of each compute element under its control and maximizes each element’s … frequency subject to a constraint on the overall power consumed across all compute elements under control") Examiner Comments: Mohammad’s MIMO system monitors workload across all compute elements and sets frequency/voltage subject to a global constraint, which teaches setting the voltage and frequency settings in the power state instruction according to a sampled workload across synchronously-executing devices. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul’s compiler-based power management technique with Mohammad’s distributed workload-aware power limiting across multiple heterogeneous compute elements in order to extend the energy efficiency benefits of compiler-based per-layer DVFS to distributed computing environments comprising multiple heterogeneous processing devices, thereby improving overall system performance under power constraints (Mohammad, Abstract). Both Paul and Mohammad are directed to workload-aware power management for processing devices, and applying Paul’s per-layer voltage/frequency adjustments to each device in Mohammad’s distributed system would yield the predictable result of improved energy efficiency across the distributed computing system. Regarding Claim 21, Paul and Mohammad teach The system of Claim 19. Paul and Mohammad further teach wherein the compiler coordinates injection of a plurality of power state instructions into the compiled code for execution across the plurality of computing devices to improve performance and reduce energy usage (Paul, Abstract, "autonomous and proactive power management for energy efficient execution"; Mohammad, Abstract, "improve performance of power-constrained computers") Examiner Comments: Coordinating power state instructions across distributed devices to improve performance and reduce energy is the predictable result of combining Paul’s compiler-based DVFS with Mohammad’s distributed MIMO power management. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul’s compiler-based power management technique with Mohammad’s distributed workload-aware power limiting across multiple heterogeneous compute elements in order to extend the energy efficiency benefits of compiler-based per-layer DVFS to distributed computing environments comprising multiple heterogeneous processing devices, thereby improving overall system performance under power constraints (Mohammad, Abstract). Both Paul and Mohammad are directed to workload-aware power management for processing devices, and applying Paul’s per-layer voltage/frequency adjustments to each device in Mohammad’s distributed system would yield the predictable result of improved energy efficiency across the distributed computing system. Regarding Claim 22, Paul and Mohammad teach The system of Claim 19. Paul further teaches wherein the compiler is configured to inject power state instructions that increase a voltage and a frequency of one or more computing devices of the plurality of computing devices to improve performance (Para [0016], "the determination of whether the layers or phases are compute-intensive or memory-intensive may be based on the power consumption of the VPU 100, the DPU 110, or the CMX 112 as the layers of the neural network are executed … the compute bandwidth may be limited by the number of compute units, the frequency of operation of the compute units, and the utilization of the compute units") Examiner Comments: Paul teaches that for compute-bound operations, the system increases voltage and frequency to the DPU to improve execution performance. Regarding Claim 23, Paul and Mohammad teach The system of Claim 19. Paul further teaches wherein the compiler is configured to inject power state instructions that lower a voltage and a frequency of one or more computing devices of the plurality of computing devices to reduce power consumption (Para [0024], "during a loop through all of the layers, for compute intensive layers, the DPU frequency may be reduced along the DPU’s V-F curve, while calculating power and energy repeatedly until the performance drop is close to the performance degradation threshold. For memory intensive layers, the DPU frequency may be reduced repeatedly until the DPU is able to complete the compute cycles before data transfer is completed in cache or memory") Examiner Comments: Paul teaches that for memory-bound operations where the compute units are not fully utilized, the system reduces voltage and frequency to save power. Regarding Claim 24, Paul (US 2023/0273832 A1) teaches A processing device of a distributed computing system comprising: at least one processor (Abstract, "an apparatus such as system-on-chip (SoC) comprising an accelerator configurable to load and execute a neural network") Examiner Comments: Paul teaches a processing device (SoC/accelerator) having at least one processor; execute, with the at least one processor, the compiled code (Para [0042], "the accelerator executes the neural network") Examiner Comments: Paul teaches execution of compiled neural network code on the accelerator; set a voltage and frequency of the at least one processor when a power state instruction is executed, wherein the power state instruction is injected by the compiler at the centralized computing device based on a workload assigned to a portion of uncompiled code for a respective portion of the compiled code, the workload defining an analysis of processor utilization for the portion of the uncompiled code that is to be executed by the processing device and being obtained by the centralized computing device during compilation of the computer program for the distributed computing system (Para [0028], "The determination of Fcompute and FCMX may be passed to the local PMU at runtime through the use of a special power-management instruction from the compiler … based at least in part on the determination of Fcompute and FCMX derived during compilation of the neural network"; Para [0008], "using a local power management unit (PMU) on the VPU that can dynamically change the voltage and/or the frequency of a signal to the VPU, while still operating within a fixed power budget from a central processing unit (CPU) level power-management unit") Examiner Comments: Paul teaches that the compiler inserts a special power-management instruction that directs the local PMU to set voltage and frequency based on the compiler’s workload analysis of processor utilization, wherein that determination is expressly obtained/derived during compilation of the neural network, thereby teaching the workload being obtained during compilation. Paul did not specifically teach a network interface configured to electronically communicate with a centralized computing device configured to operate a compiler and manage the distributed computing system; wherein the processing device is configured to receive compiled code for a computer program from the compiler of the distributed computing system. However, Mohammad (US 2021/0124404 A1) teaches a network interface configured to electronically communicate with a centralized computing device configured to operate a compiler and manage the distributed computing system; receiving compiled code from the compiler of the distributed computing system (Para [0010], "a server rack comprising a mix of CPU and GPU blades" (where the server blades are networked together and managed by a centralized control system)) Examiner Comments: Mohammad’s server rack architecture includes networked processing devices that communicate with and receive workloads from a centralized management and compilation system. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul’s teaching with Mohammad’s in order to extend the compiler-based power management to networked processing devices in a distributed computing system, thereby achieving system-wide energy efficiency and performance optimization across multiple heterogeneous compute elements (Mohammad, Abstract). 12. Claim(s) 4, 9, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Paul (US 2023/0273832 A1) in view of Mohammad (US 2021/0124404 A1), further in view of Jasoliya (US 10,699,369 B2). Regarding Claim 4, Paul and Mohammad teach The method of Claim 1. Paul and Mohammad did not specifically teach wherein the uncompiled code defines high level operations for a computational graph defining a machine-learning model. However, Jasoliya (US 10,699,369 B2) teaches wherein the uncompiled code defines high level operations for a computational graph defining a machine-learning model (Col. 11, ln. 60-67, "In one embodiment the additional fixed function logic 516 can also include machine-learning acceleration logic, such as fixed function matrix multiplication logic, for implementations including optimizations for machine learning training or inferencing") Examiner Comments: Jasoliya teaches processing of machine-learning workloads (matrix multiplication high level operations) that define a computational graph of a machine-learning model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul and Mohammad’s teaching with Jasoliya’s in order to apply the combination to ML code, as Jasoliya teaches DVFS optimizes energy for machine-learning computations by reducing energy in ML training on distributed systems, as ML workloads are computation-intensive. Regarding Claim 9, Paul and Mohammad teach The method of Claim 8. Paul and Mohammad did not specifically teach wherein the one or more operations are high level operations for training a machine-learning model. However, Jasoliya (US 10,699,369 B2) teaches wherein the one or more operations are high level operations for training a machine-learning model (Col. 11, ln. 60-67, "In one embodiment the additional fixed function logic 516 can also include machine-learning acceleration logic, such as fixed function matrix multiplication logic, for implementations including optimizations for machine learning training or inferencing") Examiner Comments: Jasoliya teaches high level operations (matrix multiplication logic) for machine-learning training. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul and Mohammad’s teaching with Jasoliya’s in order to apply the combination to ML code, as Jasoliya teaches DVFS optimizes energy for machine-learning computations by reducing energy in ML training on distributed systems, as ML workloads are computation-intensive. Regarding Claim 18, Paul and Mohammad teach The method of Claim 17. Paul and Mohammad did not teach wherein the machine-learning model is a large language model. However, Jasoliya (US 10,699,369 B2) teaches wherein the machine-learning model is a large language model (Col. 11, ln. 60-67, "In one embodiment the additional fixed function logic 516 can also include machine-learning acceleration logic, such as fixed function matrix multiplication logic, for implementations including optimizations for machine learning training or inferencing") Examiner Comments: Jasoliya teaches machine-learning acceleration logic used for machine-learning models, which encompasses large language models. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Paul and Mohammad’s teaching with Jasoliya’s in order to apply the combination to ML code, as Jasoliya teaches DVFS optimizes energy for machine-learning computations by reducing energy in ML training on distributed systems, as ML workloads are computation-intensive. 13. Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Paul (US 2023/0273832 A1) in view of Mohammad (US 2021/0124404 A1), further in view of Zhang (US 2023/0334292 A1). Regarding Claim 5, Paul and Mohammad teach The method of Claim 1. Paul and Mohammad did not specifically teach wherein the compiling is performed by an accelerated linear algebra (XLA) compiler. However, Zhang (US 2023/0334292 A1) teaches wherein the compiling is performed by an accelerated linear algebra (XLA) compiler (Para [0089], "The XLA compiler is a linear algebra compiler for a specific field, and can speed up operating of a TensorFlow model, possibly without changing source code") Examiner Comments: Zhang expressly teaches an accelerated linear algebra (XLA) compiler used to compile a machine-learning (TensorFlow) model. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined Paul and Mohammad’s teaching with Zhang’s in order to compile the machine-learning computational graph using an XLA compiler, as Zhang teaches that an XLA compiler speeds up operation of a machine-learning model without requiring changes to the source code (Zhang, Para [0089]), yielding the predictable benefit of faster and more efficient compilation of the machine-learning workload. Response to Arguments Applicant argues (Remarks, p. 7) that a skilled artisan reading the specification would understand “compiler” to be the name of the structure that performs the recited function, and therefore 35 U.S.C. 112(f) should not apply. This argument is not persuasive. Under MPEP § 2181, the term “compiler” as recited in Claims 19 and 22-23 is a generic placeholder (a non-structural term coupled with functional language, “configured to” receive, compile, identify, retrieve, and inject) that does not, by itself, connote sufficient structure to perform the entirety of the claimed functions, which include retrieving a respective workload defining an analysis of processor utilization and injecting a power state instruction that identifies a voltage setting and a frequency setting. These claimed functions extend beyond the ordinary meaning of a “compiler” (a program that translates source code into object code) and therefore “compiler configured to …” is properly interpreted under 35 U.S.C. 112(f). It is noted that this interpretation is not adverse to Applicant: the specification discloses corresponding structure for the claimed functions (compiler 102 and cost model 108, [0021]-[0026], and the algorithms of FIGS. 2-3), and accordingly no rejection under 35 U.S.C. 112(b) arises from the 112(f) interpretation. If Applicant does not intend the limitation(s) to be interpreted under 35 U.S.C. 112(f), Applicant may amend the claim(s) to recite sufficient structure to perform the claimed function. The 112(f) interpretation is maintained. Applicant argues (Remarks, pp. 7-8) that the amended limitation, namely “retrieving, during the compiling of the uncompiled code, a respective workload for each of the plurality of processing devices defining an analysis of processor utilization for the respective portion of the uncompiled code that is to be executed by the processing device,” is not taught or suggested by the cited references. Specifically, Applicant argues that Mohammad teaches a MIMO control system that monitors workloads and dynamically makes changes during runtime (citing Mohammad [0010], [0014]), and therefore Mohammad does not teach retrieving a workload “during the compiling of the uncompiled code.” Examiner respectfully disagrees. Applicant’s argument attacks Mohammad individually for a feature that the rejection relies on Paul, not Mohammad, to teach. One cannot show nonobviousness by attacking references individually where the rejection is predicated upon a combination of prior art disclosures. As set forth in the rejection of Claim 1 above, Mohammad is relied upon only for the distributed-execution aspect (a plurality of processing devices in a distributed computing system and the identification of a respective portion of code for each device). The limitation of retrieving the respective workload during compilation is taught by Paul. Contrary to Applicant’s characterization, Paul expressly teaches that the workload analysis is performed by the compiler during compilation, not merely at runtime. Paul discloses that “a compiler or performance simulator … may compile and analyze the neural network, layer-by-layer, to determine which layers are compute-intensive and which are memory-intensive” (Paul, [0016]), and that “it may be possible to fully estimate the execution latency and energy use or requirement of the layers of the neural network … From the estimation, fractions of the total power budget available to the VPU 100 may be allocated to the compute and memory units” (Paul, [0017]). Most directly, Paul states that the voltage/frequency determinations applied at runtime are “based at least in part on the determination of Fcompute and FCMX derived during compilation of the neural network” (Paul, [0028]). The compute-versus-memory (i.e., processor utilization) analysis in Paul is thus expressly derived during compilation, which teaches “retrieving, during the compiling of the uncompiled code, a respective workload … defining an analysis of processor utilization.” Applicant’s further argument that Paul does not disclose distributed execution across a plurality of processing devices is acknowledged; that is precisely why Mohammad is combined with Paul. Mohammad teaches distributed execution across a heterogeneous plurality of processing devices (Mohammad, [0010], [0014]), and the combination applies Paul’s compile-time, per-portion workload analysis and per-portion power state instruction to each of the plurality of processing devices in Mohammad’s distributed system. The motivation to combine is set forth in the rejection above and is reproduced in substance: to extend the energy efficiency benefits of Paul’s compiler-based per-layer DVFS to distributed computing environments comprising multiple heterogeneous processing devices, thereby improving overall system performance under power constraints (Mohammad, Abstract). Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMIR SOLTANZADEH whose telephone number is (571)272-3451. The examiner can normally be reached M-F, 9am - 5pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Wei Mui can be reached at (571) 272-3708. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /AMIR SOLTANZADEH/Examiner, Art Unit 2191 /Ted T. Vo/Primary Examiner, Art Unit 2191
Read full office action

Prosecution Timeline

Feb 20, 2024
Application Filed
Mar 18, 2026
Non-Final Rejection mailed — §103, §112
Jul 16, 2026
Response Filed
Sep 08, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748683
LARGE LANGUAGE MODEL ANALYSIS AND GROUPING OF SOFTWARE REQUIREMENTS TO GENERATE TEST CASES FOR SOFTWARE TESTING
2y 6m to grant Granted Sep 29, 2026
Patent 12743270
SYSTEMS AND METHODS FOR LOADING MODIFIED CLASSES INTO A RUNNING APPLICATION
2y 7m to grant Granted Sep 22, 2026
Patent 12737167
TEMPLATE TRANSPILATION TOOL
2y 8m to grant Granted Sep 15, 2026
Patent 12737162
USING GENERATIVE AI TO MAKE A NATURAL LANGUAGE INTERFACE
2y 7m to grant Granted Sep 15, 2026
Patent 12730618
SYSTEM AND METHOD FOR IDENTIFICATION, TOKENIZATION, AND DEPENDENCY MAPPING OF SOURCE CODE IN A NETWORK ENVIRONMENT
2y 6m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
81%
Grant Probability
98%
With Interview (+16.6%)
2y 6m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 434 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month