Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-3 and 6-22 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-3, 10-13, 17, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang (US 2021/0295174) and further in view of Chandra (US 2019/0266015) and further in view of Xu (US 2020/0301739).
Regarding claim 1, Zhang teaches: A task method for task execution, wherein the method comprises:
determining a plurality of deep learning tasks to be concurrently executed and an artificial intelligence model (¶ 33, “allowing for flexibility in determining an appropriate resource-accuracy tradeoff for each of several concurrently-running deep learning models involves taking into account several factors and preferences for each deep learning model”) for implementing a respective deep learning task of the plurality of deep learning tasks (¶ 33, “a device can employ a runtime scheduler that selects the most suitable sub-model of a multi-capacity model for each application”);
obtaining an execution policy of the respective deep learning task (¶ 33, “allowing for flexibility in determining an appropriate resource-accuracy tradeoff for each of several concurrently-running deep learning models”), wherein the execution policy indicates a scheduling mode and a model variant of the respective deep learning task, and the model variant of the respective deep learning task is obtained according to the artificial intelligence model for implementing the respective deep learning task (¶ 33, “a device can employ a runtime scheduler that selects the most suitable sub-model of a multi-capacity model for each application and determines the optimal amount of runtime resources to allocate to each selected sub-model to jointly maximize the accuracy and minimize the inference latency of concurrent vision applications”); and
executing the respective deep learning task according to the execution policy of the respective deep learning task (¶ 42, “the scheduler 58 examines the model profiles 52 of all running tasks that use multi-capacity models, selects the appropriate descendant model for each application 62”).
Zhang does not teach as clearly as Chandra teaches: obtaining an execution policy of the respective deep learning task (¶ 50, “The allocator 106 handles the task of resource allocation for the DNN workloads. The allocator 106 takes as input the system resource utilization from the profiler 110, and then determines an efficient resource allocation scheme”).
It would have been obvious to a person having ordinary skill in the art, at the effective filing date of the invention, to have applied the known technique of obtaining an execution policy of the respective deep learning task, as taught by Chandra, in the same way to the method, as taught by Zhang. Both inventions are in the field of executing deep neural network workloads, and combining them would have predictably resulted in a method that “determines an efficient resource allocation scheme,” as indicated by Chandra (¶ 50).
Zhang and Chandra do not teach, however, Xu teaches: the scheduling mode indicates to execute the respective deep learning task concurrently with another deep learning task (¶ 50, “the resource usage optimizer 303 can assign at least a part of the computational resources of the accelerator to concurrently execute two or more neural networks”) the another deep learning task is selected by: for each remaining deep learning task that is in the plurality of the deep learning tasks and that is different from the respective deep learning task, determining a sum resource utilization rate by adding a resource utilization rate of the respective deep learning task with a resource utilization rate of the each remaining deep learning task (¶ 37, “Workload analyzer 301 can determine an amount of resources for executing each neural network of the received two or more neural networks;” ¶ 68, “At step S620, the total resources needed to process the received two or more neural networks are determined”); and selecting a remaining deep learning task that corresponds to a greatest sum resource utilization rate as the another deep learning task (¶ 74, “if concurrent processing of the multiple neural network can lead to maximizing resource usage, the process proceeds to step S650. At step S650, the received two or more neural networks can be scheduled for concurrent execution on an accelerator based on the optimization result at step S630”).
It would have been obvious to a person having ordinary skill in the art, at the effective filing date of the invention, to have applied the known technique of the scheduling mode indicates to execute the respective deep learning task concurrently with another deep learning task, wherein the another deep learning task is selected by: for each remaining deep learning task that is in the plurality of the deep learning tasks and that is different from the respective deep learning task, determining a sum resource utilization rate by adding a resource utilization rate of the respective deep learning task with a resource utilization rate of the each remaining deep learning task; and selecting a remaining deep learning task that corresponds to a greatest sum resource utilization rate as the another deep learning task, as taught by Xu, in the same way to the task execution method, as taught by Zhang and Chandra. Both inventions are in the field of concurrent multi-neural-network scheduling, and combining them would have predictably resulted in “increasing resource utilization rate on accelerators and thus improving overall throughput of the accelerators,” as indicated by Xu (¶ 35).
Regarding claim 2, Zhang teaches: The method according to claim 1, wherein the executing the respective deep learning task according to the execution policy of the respective deep learning task comprises: executing, by using a model variant indicated by an execution policy of any deep learning task (¶ 42, “the scheduler 58 examines the model profiles 52 of all running tasks that use multi-capacity models, selects the appropriate descendant model for each application 62”), the respective deep learning task in a scheduling mode indicated by the execution policy of the respective deep learning task (¶ 42, “and determines the amount of runtime resources to allocate to each selected descendant model 64 to jointly maximize the accuracy and minimize the inference latency of those applications”).
Regarding claim 3, Zhang teaches: The method according to claim 1, wherein the scheduling mode indicates the execution priority of the respective deep learning task (¶ 42, “The scheduler 58 can also monitor various environmental factors from sensor input, and/or the output of various deep learning models, to assess whether circumstances are present that would elevate or change the criticality, priority, or importance of various applications”).
Regarding claim 10, Zhang teaches: The method according to claim 1, wherein the model variant indicated by the execution policy of the respective deep learning task is obtained by compressing the artificial intelligence model for implementing the respective deep learning task (¶ 54, “By pruning these redundant filters, the VGG-16 model was effectively compressed without any meaningful accuracy degradation”).
Regarding claim 11, Zhang teaches: The method according to claim 10, wherein the model variant indicated by the execution policy of the respective deep learning task is obtained by compressing the artificial intelligence model for implementing the respective deep learning task and by adjusting a weight parameter of the compressed artificial intelligence model (¶ 58, “due to the retraining step within each pruning iteration, these pruned or compressed models may have different model parameters, or differently grouped weights, or other idiosyncrasies resulting from the model pruning/compression phase”).
Claims 12, 13, 17, and 18 recite commensurate subject matter as claims 1 and 2. Therefore, they are rejected for the same reasons.
Claim(s) 6, 14, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, Chandra, and Xu, as applied above, and further in view of Nakazawa (US 10,203,985).
Regarding claim 6, Zhang, Chandra, and Xu do not teach; however, Nakazawa teaches: dividing the respective deep learning task into a plurality of subtasks (col. 8:13-17, “In step S502, the task 103, based on the task detail information created in step S501, divides the task into a plurality of subtasks and creates the subtask definition data 109 that represents respective subtasks”);
determining a priority of each subtask in the respective deep learning task among subtasks of a same type comprised in the plurality of deep learning tasks (col. 9:50-52, “In step S604, the subtask controller 102 determines whether or not the obtained subtask queue is the subtask queue of the high degree of priority”); and
executing the respective deep learning task based on the execution policy of the respective deep learning task and the priority of the subtask (col. 2:20-22, “obtaining a subtask from one of the plurality of queues and causing the obtained subtask to execute by newly creating a thread”).
It would have been obvious to a person having ordinary skill in the art, at the effective filing date of the invention, to have applied the known technique of dividing the respective deep learning task into a plurality of subtasks; determining a priority of each subtask in the respective deep learning task among subtasks of a same type comprised in the plurality of deep learning tasks; and executing the respective deep learning task based on the execution policy of the respective deep learning task and the priority of the subtask, as taught by Nakazawa, in the same way to the respective deep learning task, as taught by Zhang, Chandra, and Xu. Both inventions are in the field of executing tasks, and combining them would have predictably resulted in “perform processing efficiently when processing a task using a queue,” as indicated by Nakazawa (col. 2:63-64).
Claims 14 and 19 recite commensurate subject matter as claim 6. Therefore, they are rejected for the same reasons.
Claim(s) 7-9, 15, 16, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, Chandra, and Xu, as applied above, and further in view of Dias (US 11,315,014).
Regarding claim 7, Zhang, Chandra, and Xu do not teach; however, Dias teaches: obtaining a plurality of candidate execution policies of the respective deep learning task, wherein at least scheduling modes or model variants indicated by any two candidate execution policies are different (col. 9:41-47, “the analytical engine generates several possible new states exploring available resources and the parameter space. In this embodiment, the analytical engine may consider new states with more or less resources, different scheduling, alternative algorithms, different data placement and alternative data partitioning strategies”);
obtaining performance data for executing the respective deep learning task according to each candidate execution policy (claim 1, “executing the set of subworkflows, collecting provenance data from the execution, and collecting monitoring data that represents a state of the set of resources”); and
selecting the execution policy of the respective deep learning task from the plurality of candidate execution policies based on the performance data of the plurality of candidate execution policies (col. 9:48-52, “In this embodiment, candidates are evaluated using the DNN that predicts the outcome if that state is considered. In this embodiment, the state that minimizes the cost function associated with the user-defined quality requirements is selected as the best state”).
It would have been obvious to a person having ordinary skill in the art, at the effective filing date of the invention, to have applied the known technique of obtaining a plurality of candidate execution policies of the respective deep learning task, wherein at least scheduling modes or model variants indicated by any two candidate execution policies are different; obtaining performance data for executing the respective deep learning task according to each candidate execution policy; and selecting the execution policy of the respective deep learning task from the plurality of candidate execution policies based on the performance data of the plurality of candidate execution policies, as taught by Dias, in the same way to the obtaining an execution policy of the respective deep learning task, as taught by Zhang, Chandra, and Xu. Both inventions are in the field of executing deep learning workflows, and combining them would have predictably resulted in “optimizing an allocation of resources of the set of resources to each task of the sets of tasks to ensure compliance with a user-defined quality metric based on the deep neural network output,” as indicated by Dias (abstract).
Regarding claim 8, Chandra teaches: The method according to claim 7, wherein the performance data comprises real-time data (¶ 42, “Runtime parameters 200 that used by the profiler include a sampling rate 202, a batch size 204, a precision 206, and a CPU core utilization 208”), and the real-time data is obtained through prediction according to a pretrained artificial intelligence model (¶ 42, “The profiler 110 uses these inputs 200 and predicts the run time 220 and memory usage 222 of a DNN workload based on the DNN model”).
Regarding claim 9, Zhang teaches: The method according to claim 7, wherein the performance data comprises accuracy data (¶ 41, “The processing latency 46, memory footprint 48, and accuracy 50 characteristics of each descendant model 40 of the multi-capacity model 44 are identified and stored in an overall model profile 52 of the multi-capacity model 44”), and the accuracy data is obtained based on precision of a model variant indicated by a respective candidate execution policy (¶ 53, “As an example, the topmost curve (marked as blue dotted line with blue triangle markers and three hashmarks) shows the test accuracies are 89.75%, 89.72% and 87.40% when 0% (i.e., the full VGG-16 model with no filters pruned), 50% and 90% of the filters within the 13th convolutional layer conv.sub.13 are pruned, respectively”).
Claims 15, 16, and 20 recite commensurate subject matter as claims 7 and 8. Therefore, they are rejected for the same reasons.
Claim(s) 21 and 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang, Chandra, and Xu, as applied above, and further in view of Nagpal (US 11,561,826).
Regarding claim 21, Zhang, Chandra, and Xu do not teach; however, Nagpal teaches: the respective deep learning task comprises a plurality of execution units (col. 1:66-67 and col. 2:1-4, “The method includes generating a graph in a memory by the computer processor. The graph has nodes and edges, each node represents a task and specifies an assignment of the task to one or more of the kernel objects, and each edge represents a data dependency between nodes”), and each of the plurality of execution units implements a different function of the respective deep learning task (col. 1:41-45, “A neural network application involves, in addition to the inference stage, compute-intensive stages such as pre-processing and post-processing of data. Pre-processing can include reading data from retentive storage, decoding, resizing, color space conversion, scaling, cropping, etc.”), wherein the dividing is based on execution bodies and task properties of the plurality of execution units (col. 6:6-12, “A system can be configured to include different compute circuits to execute different tasks of the ML application(s). For example, a CPU can be programmed to perform a pre-processing task of a graph, and an FPGA can be configured to perform tasks of tensor operations as part of inference”).
It would have been obvious to a person having ordinary skill in the art, at the effective filing date of the invention, to have applied the known technique of the respective deep learning task comprises a plurality of execution units, and each of the plurality of execution units implements a different function of the respective deep learning task, wherein the dividing is based on execution bodies and task properties of the plurality of execution units, as taught by Nagpal, in the same way to the obtaining an execution policy of the respective deep learning task, as taught by Zhang, Chandra, and Xu. Both inventions are related to neural network workload scheduling across computing resources, and combining them would have predictably resulted in a method adapted “in order to improve throughput,” as indicated by Nagpal (col. 1:54).
Regarding claim 22, Nagpal teaches: The method according to claim 21, wherein the execution bodies comprise a central processing unit (CPU) and a graphic processing unit (GPU) (col. 3:66-67, “Reading of data is performed by CPUs, and subsequent processing is performed by accelerators such as GPUs, FPGA, etc”), and the task properties comprise a neural network inference (col. 6:9-12, “an FPGA can be configured to perform tasks of tensor operations as part of inference”) and a non-neural network inference (col. 5:15-18, “Examples of tasks include inputting data, formatting input data, computing operations associated with layers of a neural network, and those in the description of the kernels in Table 1”).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB D DASCOMB whose telephone number is (571)272-9993. The examiner can normally be reached M-F 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Vital can be reached at (571) 272-4215. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACOB D DASCOMB/Primary Examiner, Art Unit 2198