DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
Claims 1-7 and 15-21, 23-25, and 27 are pending and are examined herein.
Claims 1-7 and 15-21, 23-25, and 27 are rejected under 35 USC 103.
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 02/25/2026 has been entered.
Response to Arguments
The rejections under 35 USC 112 are withdrawn in view of Applicant’s amendment.
The rejection under 35 USC 101 is withdrawn in view of Applicant’s response. In particular, the claims as amended reflect the improvement to computer technology described at as-filed [0040] by reducing incidence of stalling.
Applicant’s arguments filed 02/25/2026 regarding the rejection under 35 USC 103 have been fully considered, but are not persuasive. Applicant argues that in Venkatesan certain steps are performed by a user, not by a device as required by the claims, and consequently Venkatesan fails to teach these steps. Examiner respectfully disagrees. Venkatesan, [0255] reads in part: ‘the term “user", as used herein, is intended to be broadly interpreted to include, for example, a computer or data processing system or a human user of a computer or data processing system, unless otherwise stated.’ Consequently, the functions being performed by a “user” does not preclude the functions being performed by a device as required by the claim. Note also that Applicant’s amendment necessitated the newly cited reference Kim.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 4-7, 15, and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over “Venkatesan” (US 2021/0174214 A1) in view of “Kim” (US 2021/0081122 A1).
Regarding claim 1, Venkatesan teaches
A method, comprising: (Venkatesan, method of Figure 4A-E)
receiving an indication to execute a portion of a machine learning (ML) model on a processing core of a device; determining, by the device, a resource allocation of the device for executing the ML model on the processing core; (Venkatesan, Figure 4A-E, especially Figure 4B, step 408, described at [0073-0075]. See also [0043, 0028]. The system processing the indication of the device to use is a determination to use that device. Note that [0255] indicates that the term “user” encompasses a computer or data processing system. The various systems including the “user” considered together are being interpreted as the claimed device. )
determining, by the device, that a layer of the ML model will use a first amount of a resource of the device that causes the resource allocation to be exceeded; (Venkatesan, Figure 4A-E, block 438 and [0176] describes determining that the memory utilization exceeds the available memory resources. [0104] indicates that the decisions about quantization may be performed with instrumentation points being between layers. Note that [0255] indicates that the term “user” encompasses a computer or data processing system. The various systems including the “user” considered together are being interpreted as the claimed device)
determining, by the device, that an adaptation may be applied to executing the layer of the ML model on the device; (Venkatesan, Figure 4A-E, step 410 receives quantization options. See description at least at [0077, 0084] and rules of applying quantization at [0086-0103]. Note that [0255] indicates that the term “user” encompasses a computer or data processing system. The various systems including the “user” considered together are being interpreted as the claimed device)
determining, by the device, whether executing the layer of the ML model using the adaptation on the device will exceed the resource allocation; when executing the layer of the ML model using the adaptation on the device is determined to exceed the resource allocation, (Venkatesan, Figure 4A-E, block 438 and [0176] describes determining whether or not the memory utilization exceeds the available memory resources. [0104] indicates that the decisions about quantization may be performed with instrumentation points being between layers. This is an iterative process, so a subsequent time through the process would be a time in which it is determined whether the layer with adaptation would exceed the resource allocation. Note that [0255] indicates that the term “user” encompasses a computer or data processing system. The various systems including the “user” considered together are being interpreted as the claimed device)
...when executing the layer of the ML model using the adaptation on the device is determined to not exceed the resource allocation: (Venkatesan, Figure 4A-E, block 438 and [0176] describes determining whether or not the memory utilization exceeds the available memory resources. [0104] indicates that the decisions about quantization may be performed with instrumentation points being between layers. This is an iterative process, so a subsequent time through the process would be a time in which it is determined whether the layer with adaptation would exceed the resource allocation. When the performance (including an indication that it can run on available resources) is acceptable, the process flows to step 446 and the model is deployed at step 450.)
executing the layer of the ML model using the adaptation on the device, wherein executing the layer using the adaptation reduces the first amount of the resource of the device used by the layer as compared to executing the layer without using the adaptation; and outputting a result of the ML model based on the executed layer. (Venkatesan, Figures 4A-E, step 450 deploys the model responsive to its performance being satisfactory. The deployment is described at [0193-0195]. The model generates outputs as described at least at [0198]. )
Venkatesan does not appear to explicitly teach
stalling the executing of the layer of the ML model using the adaptation on the device by not executing the layer of the ML model using the adaptation on the device until a resource becomes available for executing the layer of the ML model; and
However, Kim in view of Mattar teaches
stalling the executing of the layer of the ML model using the adaptation on the device by not executing the layer of the ML model using the adaptation on the device until a resource becomes available for executing the layer of the ML model; and (Kim, [0063] describes waiting (i.e., stalling) an execution of an enclave until memory becomes available. As described at [0062], the enclaves correspond to ML program execution requests (i.e., stalling an execution of an enclave means stalling the execution of the corresponding ML program). In the combination with Venkatesan, the technique taught by Kim for prioritizing jobs would be applied to the particular jobs such as the ML model using the adaptation taught by Venkatesan.)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Venkatesan by Mattar because the technqiues taught by Mattar allow for the prioritization of computer resources to high reward tasks as described by Mattar at column 18, lines 43-61
Regarding claim 4, the rejection of claim 1 is incorporated herein. Furthermore, Venkatesan teaches
wherein the resource includes at least one of an amount of memory, an amount of memory bandwidth, an amount of memory throughput, or an amount of current. (Venkatesan, Figure 4A-E, block 438 and [0176] describes determining that the memory utilization exceeds the available memory resources. [0104] indicates that the decisions about quantization may be performed with instrumentation points being between layers.)
Regarding claim 5, the rejection of claim 1 is incorporated herein. Furthermore, Venkatesan teaches
altering a number of bits used to represent features of the layer; (Venkatesan, [0227, 0240] indicate that the inputs (i.e., features) may be quantized. See also Figure 11B, “data” rows which show proposed quantizations for the data/features, which are features for at least the first layer of the network.)
altering a number of bits used to represent weights of the layer; (Venkatesan, [0039])
executing the layer on another processing core;
executing the layer using data directly from external memory;
executing the layer at a reduced speed on the processing core; or
a combination thereof.
Regarding claim 6, the rejection of claim 1 is incorporated herein. Furthermore, Venkatesan teaches
wherein adaptations applicable to the layer are predetermined. (Figure 4A-E, step 410 shows the quantization options being received. They are consequently determined at least prior to being received. The claim does not specify what they need to be “predetermined” with respect to.)
Regarding claim 7, the rejection of claim 6 is incorporated herein. Furthermore, Venkatesan teaches
wherein the adaptations applicable to the layer are provided in context information associated with the ML model and wherein the determining that the adaptation may be applied is based on the context information. (Figure 4A-E, step 410 shows the quantization options being received. The decision about what quantization to perform is based on the quantization options (see, e.g., 422).)
Regarding claim 15, Venkatesan teaches
An electronic device, comprising: a memory; and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to: (Venkatesan, [0064])
The remainder of claim 15 is substantially similar to claim 1; claim 15 is rejected with the same rationale.
Regarding claims 18-20, the rejection of claim 15 is incorporated herein. Claims 18-20 recite substantially similar subject matter to claims 4-6, respectively, and are rejected with the same rationale.
Claims 2 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over “Venkatesan” (US 2021/0174214 A1) in view of “Kim” (US 2021/0081122 A1), further in view of “Boss” (US 2016/0099884 A1).
Regarding claim 2, the rejection of claim 1 is incorporated herein. Venkatesan does not appear to explicitly teach
receiving a request, from the processing core executing the ML model, for a dynamic allocation of a second amount of the resource of the device; and determining that there is an insufficient amount of the resource to allocate the second amount to the processing core.
However, Venkatesan in view of Boss teaches
receiving a request, from the processing core executing the ML model, for a dynamic allocation of a second amount of the resource of the device; and determining that there is an insufficient amount of the resource to allocate the second amount to the processing core. (Boss, [0033] describes receiving a request from a virtual machine for additional resources, determining that the resource is unavailable and reallocating resources as a result. In the combination with Venkatesan, Venkatesan already teaches that the device is executing a machine learning model as described above with respect to claim 1.)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Venkatesan by Boss because this allows for the system to reallocate resources so that the requesting device has access to the needed resources as described by Boss at [0033].
Regarding claim 16, the rejection of claim 15 is incorporated herein. Claim 16 recites substantially similar subject matter to claim 2 and is rejected with the same rationale.
Claims 3 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over “Venkatesan” (US 2021/0174214 A1) in view of “Kim” (US 2021/0081122 A1), further in view of “Boss” (US 2016/0099884 A1), and further in view of “Chandra” (US 2019/0266015 A1).
Regarding claim 3, the rejection of claim 2 is incorporated herein. Furthermore, Venkatesan in view of Boss teaches
wherein the resource is dynamically allocated to another task (Boss, [0033] describes receiving a request from a virtual machine for additional resources, determining that the resource is unavailable and reallocating resources as a result (i.e., the resource is allocated to a different task run by the virtual machine to which it is reallocated).)
Venkatesan in view of Boss does not appear to explicitly teach
another executing ML model.
However, Venkatesan in view of Boss and Chandra teaches
another executing ML model. (Chandra, Abstract, [0108] indicates that the various neural network workloads are allocated to processing cores for processing. In the combination with Venkatesan and Boss, Boss already teaches reallocating the resource. In the combination, the task would be a neural network (i.e., ML model) task as taught by Chandra.)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Venkatesan in view of Boss by Chandra because the techniques of Chandra allow for distributed systems to take full advantage of all of the devices in the network when performing ML tasks resulting in more efficient schedule workloads as described by Chandra at [0018].
Regarding claim 17, the rejection of claim 16 is incorporated herein. Claim 17 recites substantially similar subject matter to claim 3 and is rejected with the same rationale.
Claims 21 and 25 are rejected under 35 U.S.C. 103 as being unpatentable over “Venkatesan” (US 2021/0174214 A1) in view of “Kim” (US 2021/0081122 A1), further in view of “Kotler” (US 2021/0342733 A1).
Regarding claim 21, the rejection of claim 1 is incorporated herein. Venkatesan does not appear to explicitly teach
wherein the resource allocation includes a first portion associated with the layer of the ML model and a second portion common to a plurality of layers of the ML model.
However, Venkatesan in view of Kotler teaches
wherein the resource allocation includes a first portion associated with the layer of the ML model and a second portion common to a plurality of layers of the ML model. (Kotler, Abstract, [0063] describes performing an allocation/assignment of different layers of a neural network to processing tiles. In particular, [0063] indicates that some tiles may be shared between layers.)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Venkatesan by Kotler because the allocation techniques taught by Kotler allow for a minimization of processing time, power, data transfers, and/or memory reads as described by Kotler at [0064].
Regarding claim 25, the rejection of claim 15 is incorporated herein. Claim 25 recites substantially similar subject matter to claim 21 and is rejected with the same rationale.
Claims 23-24 and 27 are rejected under 35 U.S.C. 103 as being unpatentable over “Venkatesan” (US 2021/0174214 A1) in view of “Kim” (US 2021/0081122 A1), further in view of “Yan” (US 2022/0188609 A1).
Regarding claim 23, the rejection of claim 1 is incorporated herein. Venkatesan does not appear to explicitly teach
determining whether to retrieve a first set of data or a second set of data from a memory based on whether execution of the layer of the ML model without the adaptation will cause the resource allocation to be exceeded.
However, Venkatesan in view of Yan teaches
determining whether to retrieve a first set of data or a second set of data from a memory based on whether execution of the layer of the ML model without the adaptation will cause the resource allocation to be exceeded. (Yan, Abstract describes storing models with varying levels of precision and choosing the precision to be loaded based on available RAM. See also [0030-0031, 0033, 0041] where a model suitable for the amount of memory available is selected. The determination of a largest model consistent with memory availability means that larger model sizes would exceed the memory availability.)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Venkatesan by Yan because doing so allows for running the model without “starving other functions of RAM...or reducing the frequency of the neural network operations” (Yan, [0041]).
Regarding claim 24, the rejection of claim 23 is incorporated herein. Furthermore, Venkatesan in view of Yan teaches
wherein the first set of data has a greater precision than the second set of data. (Yan, Abstract describes storing models with varying levels of precision and choosing the precision to be loaded based on available RAM. See also [0030-0031, 0033] where a model suitable for the amount of memory available is selected.)
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to have combined these references in this way for the same reasons given above with respect to claim 23.
Regarding claim 27, the rejection of claim 15 is incorporated herein. Claim 27 recites substantially similar subject matter to claim 23 and is rejected with the same rationale.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Markus A Vasquez whose telephone number is (303)297-4432. The examiner can normally be reached Monday to Friday 10AM to 2PM PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li Zhen can be reached at (571) 272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARKUS A. VASQUEZ/Primary Examiner, Art Unit 2121