DETAILED ACTION
Claims 1-27 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-5, 12, 16-21 and 26 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Taylor et al. (US 2014/0055347 A1).
Regarding claim 1, Taylor teaches the invention as claimed including a device () comprising:
a memory configured to store data indicating performance metrics of a plurality of hardware processing units for each of a plurality of process types ([0018]; [0024]; [0028] According to some embodiments, the method 300 may comprise determining (e.g., by the specially-programmed computer processing device) one or more characteristics of the set of image processing tasks, at 304. The data descriptive of the image processing tasks may be analyzed, for example, to infer and/or obtain attribute data regarding the type(s), quantity, priority, and/or interdependencies of the image processing tasks, and/or such attribute data may be looked-up and/or otherwise determined. According to some embodiments, the characteristics may include data descriptive of how and/or when such tasks have previously been performed (and/or performance metrics associated therewith--such as a score). [0030] A database and/or cache data store may, for example, store an indication for each available processing resource, such indication being descriptive of a variety of characteristics of each resource. Such characteristics may include, but are not limited to, (i) an indication of an availability associated with the set of heterogeneous processing resources, (ii) an indication of a performance metric associated with the set of heterogeneous processing resources, (iii) an indication of power consumption associated with the set of heterogeneous processing resources, and/or (iv) an indication of a proximity of stored data in association with the set of heterogeneous processing resources. According to some embodiments, the characteristics may include data descriptive of how and/or when such processing resources have previously been utilized and/or how they performed (and/or performance metrics associated therewith--such as a score, execution time, etc.).); and
a first hardware processing unit is configured to:
execute software code, which includes a processing job to be executed ([0031] the system and/or device may cause those resources to process the allocated imaging tasks in accordance with an allocation and/or schedule determined by the system and/or device.);
select at least one of the hardware processing units from the plurality of hardware processing units to perform the processing job based on a given process type of the processing job and the performance metrics of the hardware processing units for the given process type, wherein the selected at least one hardware processing unit is configured to process the processing job ([0018]; [0023] For example, in the case that a particular type of hardware processing resource 212-1a-d, 212-2a-f, 212-3a-d, 212-4a-d of the SoC device 212 (e.g., a GPU device 212-3a-d) has an affinity for a particular type of task (e.g., as measured by one or more performance metrics), if the particular type of hardware processing resource 212-1a-d, 212-2a-f, 212-3a-d, 212-4a-d (e.g., a GPU device 212-3a-d) is available, and/or is not currently overburdened with other tasks, then the particular type of hardware processing resource 212-1a-d, 212-2a-f, 212-3a-d, 212-4a-d (e.g., a GPU device 212-3a-d) may be the preferred (e.g., highest weighted and/or scored) resource for execution of any tasks of the particular type that require processing by the SoC device 212.; [0031] the system and/or device may cause those resources to process the allocated imaging tasks in accordance with an allocation and/or schedule determined by the system and/or device.).
Regarding claim 2, Taylor teaches wherein the hardware processing units include at least one of: a central processing unit (CPU); a graphics processing unit (GPU) ([0018] According to some embodiments, the system 200 may comprise a System-on-Chip (SoC) device 212. The SoC device 212 may, in some embodiments, comprise a plurality of heterogeneous processing resources such as a plurality of processing cores 212-1a-d, a plurality of Image Signal Processor (ISP) devices 212-1a-f, a plurality of Graphics Processing Unit (GPU) devices 212-3a-d, and/or a plurality of Fixed-Function Hard-Ware (FFHW) devices 212-4a-d.); a video decoder; a matrix multiplication unit; a neural engine block; an encryption engine; a decryption engine; or a hardware accelerator.
Regarding claim 3, Taylor teaches wherein the data includes metric descriptor fingerprints and corresponding performance metrics ([0028]; [0030] (ii) an indication of a performance metric associated with the set of heterogeneous processing resources, (iii) an indication of power consumption associated with the set of heterogeneous processing resources… According to some embodiments, the characteristics may include data descriptive of how and/or when such processing resources have previously been utilized and/or how they performed (and/or performance metrics associated therewith--such as a score, execution time, etc.). In such a manner, for example, the method 300 may take into account previous executions of the method 300 and/or otherwise take into account previous data regarding how well previous imaging tasks were executed by the available resources (e.g., by the processing array)).
Regarding claim 4, Taylor teaches wherein one of the metric descriptor fingerprints includes any one or more of the following: a metric name ([0028]; [0030] execution time, power consumption); a processing unit identifier; a process type ([0028] The data descriptive of the image processing tasks may be analyzed, for example, to infer and/or obtain attribute data regarding the type(s), quantity, priority, and/or interdependencies of the image processing tasks, and/or such attribute data may be looked-up and/or otherwise determined); a performance domain; a metric creator identifier; and a metric creation timestamp.
Regarding claim 5, Taylor teaches wherein one or the performance metrics includes one or more of the following: a processing speed metric; a latency metric; a power consumption metric ([0030] power consumption); a performance metric based on processing speed and latency; and a performance metric based on processing speed, latency and power consumption.
Regarding claim 12, Taylor teaches wherein the first hardware processing unit is configured to select the at least one hardware processing unit based on at least one power consumption metric of the performance metrics of the hardware processing units for the given process type ([0022] According to some embodiments, various attributes of the hardware processing resources 212-1a-d, 212-2a-f, 212-3a-d, 212-4a-d of the SoC device 212 (e.g., as determined via the libraries 238-1, 238-2, 238-3) may be utilized to determine how the imaging tasks should be distributed for execution. Whether a particular hardware processing resource 212-1a-d, 212-2a-f, 212-3a-d, 212-4a-d is currently (or expected to be) available, and/or a performance metric, power consumption metric, and/or location (e.g., within the SoC device 212) of a hardware processing resource 212-1a-d, 212-2a-f, 212-3a-d, 212-4a-d may be utilized, for example, to determine which processing tasks should be executed by the various available hardware processing resource 212-1a-d, 212-2a-f, 212-3a-d, 212-4a-d.).
Regarding claim 16, it is a method type claim having similar limitations as claim 1 above. Therefore, it is rejected under the same rationale above.
Regarding claim 17, it is a system type claim having similar limitations as claim 1 above. Therefore, it is rejected under the same rationale above. Further the additional limitation an integrated circuit is taught by Taylor in [0018] “a System-on-Chip (SoC) device 212”
Regarding claim 18, it is a system type claim having similar limitations as claim 2 above. Therefore, it is rejected under the same rationale above.
Regarding claim 19, it is a system type claim having similar limitations as claim 3 above. Therefore, it is rejected under the same rationale above.
Regarding claim 20, it is a system type claim having similar limitations as claim 4 above. Therefore, it is rejected under the same rationale above.
Regarding claim 21, it is a system type claim having similar limitations as claim 5 above. Therefore, it is rejected under the same rationale above.
Regarding claim 26, it is a system type claim having similar limitations as claim 12 above. Therefore, it is rejected under the same rationale above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 7, 8, 10, 11, 13, 24, 25 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over Taylor et al. (US 2014/0055347 A1) in further view of Alt et al. (US 11677681 B1).
Regarding claim 7, Taylor does not teach further comprising a first die, a second die, and a data communication bus between the first die and second die, wherein the hardware processing units include a second hardware processing unit disposed on the first die and a third hardware processing unit disposed on the second die, the first hardware processing unit being configured to select the second hardware processing unit to perform at least part of the processing job based on the given process type of the processing job, and the performance metrics of the second hardware processing unit and the third hardware processing unit for the given process type.
However Wang teaches a scheduling system with a plurality of networked devices and storing utilization characteristics in memory to aid in workload placement. Further Wang teaches further comprising a first die, a second die, and a data communication bus between the first die and second die, wherein the hardware processing units include a second hardware processing unit disposed on the first die and a third hardware processing unit disposed on the second die, the first hardware processing unit being configured to select the second hardware processing unit to perform at least part of the processing job based on the given process type of the processing job, and the performance metrics of the second hardware processing unit and the third hardware processing unit for the given process type (Fig. 2 and 3; [0023-27]; [0029] FIG. 4 is a flow diagram of a method 400 of allocating workloads to computing nodes based on energy efficiency, in accordance with some embodiments. Method 400 may be performed by processing logic that may comprise hardware (e.g., circuitry, dedicated logic, programmable logic, a processor, a processing device, a central processing unit (CPU), a system-on-chip (SoC), etc.), software (e.g., instructions running/executing on a processing device), firmware (e.g., microcode), or a combination thereof. In some embodiments, at least a portion of method 400 may be performed by workload scheduler 115 of FIGS. 1-3.; [0031] Method 400 begins at block 410, where the processing logic obtains an energy consumption profile for a set of computing nodes. For example, the processing logic may retrieve an energy consumption profile for a computing system type and/or hardware type for each of the computing nodes of a cluster of computing nodes. Thus, if a computing cluster includes multiple types of computing systems or different computing hardware, multiple energy consumption profiles may be retrieved (e.g., one for each system or hardware type). In some examples, each energy consumption profile may include energy consumption metrics, computing resource utilization metrics, performance metrics, etc. collected during execution of one or more benchmark workloads on a particular type of system. The energy consumption profiles based on the benchmark workloads may be generated external to the computing cluster. In another example, the energy consumption profile(s) may be generated by running a benchmark workload on one or more of the computing nodes of the cluster.; [0034] At block 440, the processing logic determines placement of the new workload on one or more of the computing nodes in view of the estimated energy consumption for each of the computing nodes and resource requirements of the new workload. For example, the processing logic may allocate the new workload to the computing node that is estimated to use the least amount of energy to execute the new workload. In some examples, the processing logic may place the new workload to balance performance and energy consumption. For example, the processing logic may place the new workload to meet a minimum performance threshold and also minimize energy consumption for the new workload. The processing logic may place the workload in any manner to track and reduce energy consumption for the new workload.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Wang of scheduling workloads among computing nodes based on utilization metrics with the teachings of Taylor. The modification would have been motivated by the desire of combining known scheduling methods to yield predictable results.
Regarding claim 8, Taylor as cited above teaches an SoC. In addition, Wang teaches further comprising a system-on-chip comprising a first chiplet and a second chiplet, wherein the hardware processing units include a second hardware processing unit disposed on the first chiplet and a third hardware processing unit disposed on the second chiplet, the first hardware processing unit being configured to select the second hardware processing unit to perform at least part of the processing job based on the given process type of the processing job and the performance metric of the second hardware processing unit and the third hardware processing unit for the given process type (Fig. 2 and 3; [0023-27]; [0029] FIG. 4 is a flow diagram of a method 400 of allocating workloads to computing nodes based on energy efficiency, in accordance with some embodiments. Method 400 may be performed by processing logic that may comprise hardware (e.g., circuitry, dedicated logic, programmable logic, a processor, a processing device, a central processing unit (CPU), a system-on-chip (SoC), etc.), software (e.g., instructions running/executing on a processing device), firmware (e.g., microcode), or a combination thereof. In some embodiments, at least a portion of method 400 may be performed by workload scheduler 115 of FIGS. 1-3.; [0031]).
Regarding claim 10, Alt teaches wherein the first hardware processing unit is configured to select the at least one hardware processing unit based on a maximum allowed latency of the processing job and at least one latency metric of the performance metrics of the hardware processing units for the given process type (Col. 2, lines 28-33: For example, one QoS level may be a minimum bandwidth required between GPUs allocated to the job, or a maximum power consumption level, maximum power budget, minimum cost, minimum memory bandwidth, or minimum memory quantity or configuration for the job.).
Regarding claim 11, Alt teaches wherein the selected at least one hardware processing unit includes a second hardware processing unit and a third hardware processing unit, the first hardware processing unit is configured to apportion the processing job between the second hardware processing unit and the third hardware processing unit of the hardware processing units according to a ratio of processing speed metrics of the performance metrics of the second hardware processing unit to the third hardware processing unit (Col. 1, lines 29-37: As the time required for a single system or processor to complete many of these tasks would be too great, they are typically divided into many smaller tasks that are distributed to large numbers of processors such as central processing units (CPUs) or graphics processing units (GPUs) that work in parallel to complete them more quickly. Specialized computing systems having large numbers of processors that work in parallel have been designed to aid in completing these tasks more quickly and efficiently.; Col. 3, lines 46-67: the scheduler may select of fractional portions of a computing resource (e.g., half of a GPU), and may oversubscribe resources (e.g., allocate 2 jobs to the same GPU at the same time) and permit two or more jobs to concurrently share a GPU. The scheduler may select resources based on performance feedback collected from the execution of earlier similar jobs to achieve best-fit across multiple jobs awaiting scheduling. In another example, the scheduler may select resources using a multi-dimensional best fit analysis based on one or more of the following: processor interconnect bandwidth, processor interconnect latency, processor-to-memory bandwidth and processor-to-memory latency. The scheduler may also be configured to select computing resources for a job according to a predefined placement affinity (e.g., all-to-all, tile, ring, closest, or scattered). For example, if a closest affinity is selected, the scheduler may select nodes that are closest to a particular resource (e.g., a certain non-volatile memory holding the data to be processed). In tile affinity, assigning jobs to processors in a single node (or leaf or branch in a hierarchical configuration) may be preferred when selecting resources.).
Regarding claim 13, Alt teaches wherein the first hardware processing unit is configured to select the at least one hardware processing unit based on the at least one power consumption metric responsively to the device being in a power save mode (Col. 3, lines 40-45: The scheduler may be configured to mask/unmask selected resources based on user or administrator input or other system-level information (e.g. avoiding nodes/processors that are unavailable, that are experiencing abnormally high temperatures or that are on network switches that are experiencing congestion).; Col. 8, lines 52-54: As resources are allocated and come online and go offline for various reasons (e.g., maintenance), this system configuration information may be updated.).
Regarding claim 24, it is a system type claim having similar limitations as claim 10 above. Therefore, it is rejected under the same rationale above.
Regarding claim 25, it is a system type claim having similar limitations as claim 11 above. Therefore, it is rejected under the same rationale above.
Regarding claim 27, it is a system type claim having similar limitations as claim 13 above. Therefore, it is rejected under the same rationale above.
Claims 6, 9, 14, 15, 22, and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Taylor et al. (US 2014/0055347 A1) in further view of Wang et al. (US 20230418688 A1).
Regarding claim 6, Taylor as cited in [0028-30] teaches previous executions of tasks of the same type but does not teach further comprising a second hardware processing unit, wherein:
the second hardware processing unit is configured to cause a test process to be processed on at least some of the hardware processing units;
the at least some hardware processing units are configured to process the test process; the second hardware processing unit is configured to perform measurements related to performance of the at least some hardware processing units processing the test process; and
the second hardware processing unit is configured to update the data of the performance metrics responsively to the performed measurements.
However, Alt teaches further comprising a second hardware processing unit, wherein:
the second hardware processing unit is configured to cause a test process to be processed on at least some of the hardware processing units (Col. 2, lines 57-63: The computing system may comprise a number of nodes, in one or more clusters, both local and remote (e.g., cloud resources). The configuration information from the computing system may gathered from one or more system configuration files, or it may be empirically generated by running test jobs that are instrumented to measure values such as maximum/average bandwidth and latency.);
the at least some hardware processing units are configured to process the test process (Col. 3, lines 23-26: For example, the mapper may be configured to run one or more test jobs to measure the available bandwidths between the computing resources and include those in the mesh model.);
the second hardware processing unit is configured to perform measurements related to performance of the at least some hardware processing units processing the test process (Col. 3, lines 23-26); and
the second hardware processing unit is configured to update the data of the performance metrics responsively to the performed measurements (Col. 8, lines 37-54: Turning now to FIG. 8, a flowchart of an example embodiment of a method for allocating computing devices in a computing system is shown. Configuration information about the distributed computing system is gathered (step 800). This may include reading system configuration files to determine the quantity and location of available computing resources in the distributed computing system (e.g., type and number of processes, interconnect types, memory quantity and location, and storage locations). This may also include running test jobs (e.g., micro-benchmarks) that are timed to measure the interconnectivity of the computing resources. While the earlier examples above illustrated GPU and CPU interconnectivity, interconnectivity to other types of resources (e.g., memory and storage bandwidth and latency) can also be used in selecting which computing resources are allocated to jobs. As resources are allocated and come online and go offline for various reasons (e.g., maintenance), this system configuration information may be updated.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Alt to perform test jobs to gather information on the distributed computing system with the previously executed workloads of Taylor. The modification would have been motivated by the desire of establishing benchmarks and scheduling decisions that will inform subsequent task placements.
Regarding claim 9, Alt teaches wherein the selected at least one hardware processing unit includes a second hardware processing unit and a third hardware processing unit, the first hardware processing unit being configured to apportion the processing job between the second hardware processing unit and the third hardware processing unit according to a ratio of the performance metrics of the second hardware processing unit to the third hardware processing unit (Col. 1, lines 29-37: As the time required for a single system or processor to complete many of these tasks would be too great, they are typically divided into many smaller tasks that are distributed to large numbers of processors such as central processing units (CPUs) or graphics processing units (GPUs) that work in parallel to complete them more quickly. Specialized computing systems having large numbers of processors that work in parallel have been designed to aid in completing these tasks more quickly and efficiently.; Col. 3, lines 46-67: the scheduler may select of fractional portions of a computing resource (e.g., half of a GPU), and may oversubscribe resources (e.g., allocate 2 jobs to the same GPU at the same time) and permit two or more jobs to concurrently share a GPU. The scheduler may select resources based on performance feedback collected from the execution of earlier similar jobs to achieve best-fit across multiple jobs awaiting scheduling. In another example, the scheduler may select resources using a multi-dimensional best fit analysis based on one or more of the following: processor interconnect bandwidth, processor interconnect latency, processor-to-memory bandwidth and processor-to-memory latency. The scheduler may also be configured to select computing resources for a job according to a predefined placement affinity (e.g., all-to-all, tile, ring, closest, or scattered). For example, if a closest affinity is selected, the scheduler may select nodes that are closest to a particular resource (e.g., a certain non-volatile memory holding the data to be processed). In tile affinity, assigning jobs to processors in a single node (or leaf or branch in a hierarchical configuration) may be preferred when selecting resources.).
Regarding claim 14, the combination teaches wherein the first hardware processing unit is configured to run an operating system on which to execute the software code, which includes the processing job to be executed, the operating system being configured to select the at least one hardware processing units to perform the processing job based on the given process type of the processing job and the performance metrics of the hardware processing units for the given process type (Taylor’s [0012]; Wang’s Fig. 1 shows Host OS running Workload Scheduler 115; [0026-27, 0031, 34]).
Regarding claim 15, Wang teaches wherein the operating system is configured to cause the at least one hardware processing unit to process the processing job ([0017] Host OS 120 manages the hardware resources of the computer system and provides functions such as inter-process communication, scheduling, memory management, and so forth.; [0020] In some examples, host system 110A may include a workload scheduler 115 to schedule and allocate computing workloads to computing nodes of the computing cluster (e.g., among host system 110A-B and any additional host systems of the cluster). The workload scheduler may receive a workload and/or an instruction to execute a workload from client device 105 (e.g., a device of a user or customer of the computing platform). The workload scheduler 115 may determine resource requirements of the workload and allocate the workload to a computing node of the computer system 100 for optimal energy efficiency.).
Regarding claim 22, it is a system type claim having similar limitations as claim 6 above. Therefore, it is rejected under the same rationale above.
Regarding claim 23, it is a system type claim having similar limitations as claim 9 above. Therefore, it is rejected under the same rationale above.
Response to Arguments
Applicant’s arguments with respect to claims 1-27 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Applicant's amendment necessitated the new grounds of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JORGE A CHU JOY-DAVILA whose telephone number is (571)270-0692. The examiner can normally be reached Monday-Friday, 6:00am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee J Li can be reached at (571)272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JORGE A CHU JOY-DAVILA/Primary Examiner, Art Unit 2195