DETAILED ACTION
Status of Claims
This action is in reply to the application filed on 06/14/2024.
Claims 6, 12, 14 and 15 have been amended in a preliminary amendment dated 06/14/2024.
Claims 17-22 have been added preliminary amendment dated 06/14/2024.
Claims 13 and 16 have been canceled preliminary amendment dated 06/14/2024.
Claims 1-12, 14, 15, and 17-22 are currently pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Allowable Subject Matter
Claim 12 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5, 14, 15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Shiraki (US 2011/0310108 A1, cited in IDS dated 09/12/2025) in view of Munshi et al. (US 2009/0307704 A1).
Claims 1, 14, and 15:
Shiraki discloses the limitations as shown in the following rejections:
A task execution method applied to a target device configured with a graphics processing unit (GPU), the task execution method comprising: obtaining [access to a database] in response to receiving an execution instruction for a specified task, wherein the specified task is a task executed based on the GPU, and the [database] comprises a mapping relationship between a device category (kind of GPU) and a preferred GPU sub-thread number (optimum thread number) (¶0009, 0075, 0078-0079, 0094),
wherein the preferred GPU sub-thread number is number of a sub-thread of the GPU used by a representative device in the device category taking shortest time to execute the specified task based on the GPU of the representative device; (¶0074, 0081-0088, 0099) “number of threads that the GPU 121 may process at the highest speed under an image-processing condition (¶0074)…determines the thread parameter making the processing time the shortest (enabling processing at the highest speed) from all the search-target thread parameters as the optimum thread parameter” (¶0088).
determining a target device category to which the target device belongs based on the [database]; and executing the specified task by using the preferred GPU sub-thread number corresponding to the target device category (¶0079, 0092-0094).
[Claims 14, 15] a processor; and a memory configured to store executable instructions for the processor; wherein the processor is configured to read the executable instructions from the memory and execute the executable instructions (¶0040; FIG. 2, and 7)
As shown above, Shiraki discloses accessing a database over a network to obtain the optimum GPU thread number and does not clearly anticipate a configuration file storing the information.
Munshi, however, discloses (¶0033-0034, 0051, 0066-0068, 0073-0075) an analogous method for optimizing GPU execution including obtaining a configuration file (compute application library) in response to receiving an execution instruction for a specified task, wherein the specified task is a task executed based on the GPU, and the configuration file comprises a mapping relationship between a device category (GPU vendor and/or version) and an executable optimized for the device including and optimal thread group size (preferred GPU sub-thread number).
It would have been obvious to one of ordinary skill in the art prior to the filing date of the invention to modify Shiraki to utilize a local application library file as taught by Munshi to increase operational flexibility by allowing the application to be run on a variety of CPU/GPU platforms (Munshi ¶0005-0007) and to avoid network overhead each time a GPU task is executed (Munshi FIG. 1).
Claims 2, 17, and 20:
The combination of Shiraki/Munshi discloses the limitations as shown in the rejections above. Munshi further discloses wherein the configuration file is recorded with device identifiers of a plurality of devices corresponding to each device category; and the determining of the target device category to which the target device belongs based on the configuration file comprises: searching for the device category (e.g. capabilities, type, vendor, and/or version) corresponding to a device identifier of the target device in the configuration file, and using the device category searched as the target device category to which the target device belongs (¶0028, 0033-0034, 0059, 0073-0075). See also Shiraki (¶0074, 0078, 0091) disclosing searching the database using an ID “which is a combination of the kind of the GPU 121, the size of image data, and processing contents (kind of effect, effect parameter) of an effect”.
Claims 3, 18, and 21:
The combination of Shiraki/Munshi discloses the limitations as shown in the rejections above. Munshi further discloses wherein the configuration file is recorded with a categorization method of the device category, the categorization method comprising categorization by a GPU manufacturer name (GPU vendor) or categorization by a GPU model (GPU version); and the determining of the target device category to which the target device belongs based on the configuration file comprises: determining the target device category to which the target device belongs based on the categorization method of the device category recorded in the configuration file and GPU information of the target device (see at least ¶0033-0034, 0059, 0073-0075). Exemplary quotation(s): “determination may be based on, for example, the version of the physical computing device. In one embodiment, process 600 may determine that an existing compute kernel executable is optimized for a physical computing device if the version of target physical computing device in the description data matches the version of the physical computing device (¶0059)…designate multiple target physical computing devices for the API function at block 903. A target physical computing device may be designated according to types, such as CPU or GPU, versions or vendors (¶0073)…each executable may be stored with description data including, for example, the type, version and vendor of the target physical computing device” (¶0074)
Claims 4, 18, and 21:
The combination of Shiraki/Munshi discloses the limitations as shown in the rejections above. Munshi further discloses wherein the GPU information comprises the GPU manufacturer name (GPU vendor) or the GPU model (GPU version) (see at least ¶0033-0034, 0059, 0073-0075).
Claim 5:
The combination of Shiraki/Munshi discloses the limitations as shown in the rejections above. Munshi further discloses wherein all devices corresponding to a device category have a specific GPU commonality, and different device categories correspond to different GPU commonalities (see at least ¶0033-0034, 0059, 0073-0075). See also Shiraki (¶0074, 0078, 0091).
Claims 6-11 are rejected under 35 U.S.C. 103 as being unpatentable over Shiraki (US 2011/0310108 A1, cited in IDS dated 09/12/2025) in view of Munshi et al. (US 2009/0307704 A1) Jiang et al. (“Profiling and Optimizing Deep Learning Inference on Mobile GPUs”, 2020).
Claim 6:
The combination of Shiraki/Munshi discloses the limitations as shown in the rejections above. Munshi further discloses (¶0034, 0073-0074, 0064-0066) generating the application library by obtaining GPU information of a plurality of devices (designated devices), wherein each of the plurality of devices is configured with a GPU; categorizing the plurality of devices into a plurality of device categories based on the GPU information, each device category of the plurality of device categories (e.g. GPU vendors, versions) corresponding to a plurality of devices…obtaining the preferred GPU sub-thread number corresponding to the [designated devices] of the each device category, and using the preferred GPU sub-thread number as the preferred GPU sub-thread number corresponding to the each device category; and generating the configuration file based on the mapping relationship between the each device category and the preferred GPU sub-thread number. See also Shiraki ¶0074, 0078-0091) disclosing obtaining optimum thread parameters for various “kinds of GPUs” and storing it in the database mapped to a generated ID.
Shiraki/Munshi do not specifically disclose for the each device category, selecting a representative device from all devices corresponding to the each device category and/or obtaining the preferred GPU sub-thread number corresponding to the representative device to include in the configuration file mapping information.
Jiang, however, discloses an analogous method for identifying optimal work group size settings (preferred GPU sub-thread number) for particular GPU hardware configurations including categorizing the plurality of devices into a plurality of (model/manufacture Mali, Adreno) device categories (pg. 78-79). Jiang further discloses (pg. 81, sect. 5) known methods “to tune the work group size towards improving GPU performance” include “Look-up table. This method obtains optimal settings for popular hardware offline and saves into the look-up table. In run time the selected work group size is read from the table.” Thus, it was known in the art to identify a representative (popular) device from all devices corresponding to the each device category and obtain the preferred GPU sub-thread number corresponding to the representative (popular) device for subsequent use via a runtime lookup table (configuration file) that comprises a mapping relationship between a device category and a preferred GPU sub-thread number.
It would have been obvious to one of ordinary skill in the art prior to the filing date of the invention to modify the designated devices of Shiraki/Munshi’s library preparation to target popular (representative) device candidates as taught by Jiang to increase the effectiveness of the optimized settings upon deployment (pg. 81, sect. 5 and 6); and thus represents the application of a known technique to a known optimization method to yield predictable results.
Claim 7:
The combination of Shiraki/Munshi/Jiang discloses the limitations as shown in the rejections above. Jiang further discloses the obtaining of the preferred GPU sub-thread number corresponding to the representative device of the each device category comprises: obtaining a plurality of candidate GPU sub-thread numbers (Jiang “enumerate all possible work group sizes”); for each of the plurality of candidate GPU sub-thread numbers, using the each of the plurality of candidate GPU sub-thread numbers as a parameter for a preset open computing language OpenCL program, and obtaining time consumption for the representative device of the each device category to adopt the preset OpenCL program to execute the specified task; and using a candidate GPU sub-thread number corresponding to shortest time consumption as the preferred GPU sub-thread number corresponding to the representative device of the each device category in at least pg. 78; pg. 79, FIG. 4; pg. 81, sect. 5)
“we break down the neural network into operators and execute the single operator on both the Adreno and Mali GPU. We pick the typical operator convolution (Conv), since Conv is considered as one of the most time-consuming operators. We select 3×3 as the kernel size, and set 56×56×256 as the input shape, which is a common Conv shape shared across various networks [10, 17]. We enumerate all possible work group sizes and then execute Conv with them. The inference latency and the APU utilization are collected” (pg. 78, col. 1).
Examiner notes “work group size” is OpenCL specific nomenclature/variable name for the configuration setting. See also Munshi (¶0031, 0041, 0064-0067) disclosing an analogous process in the context of OpenCL. See additionally Shiraki (FIG. 8 and 9; ¶0073-0075, 0081-0088) disclosing an equivalent process for CUDA implemented executables.
Claims 8 and 9:
The combination of Shiraki/Munshi/Jiang discloses the limitations as shown in the rejections above. Jiang further discloses wherein the GPU information comprises a GPU manufacturer/model name (Mali/Ardeno); and the categorizing of the plurality of devices into the plurality of device categories based on the GPU information comprises: categorizing the plurality of devices into a plurality of manufacturer/model categories based on the GPU manufacturer/model name of each of the plurality of devices, wherein all devices corresponding to each of the plurality of manufacturer categories have a same manufacturer name in at least Jiang pg. 77-78, sect. 3.2; pg. 78, Fig. 4; and pg. 79, Fig. 9. Examiner notes that a GPU’s model brand also denotes the manufacturer, this interpretation is consistent with Applicant’s Specification (see AppSpec ¶0049 where “Mali” identifies either/both GPU manufacturer and model).
Claims 10 and 11:
The combination of Shiraki/Munshi/Jiang discloses the limitations as shown in the rejections above. Jiang further discloses wherein the selecting of the representative device from all devices corresponding to the each device category comprises: obtaining index data of each of the plurality of devices corresponding to the each device category based on a preset measurement index, wherein the preset measurement index comprises a market coverage rate (i.e. popularity) and/or device performance; and selecting the representative device from all the devices corresponding to the each device category based on the index data of the each of the plurality of devices by selecting a most typical device (popular) from all the devices corresponding to the each device category as the representative device by comparing the index data of the plurality of devices (see at least Jiang 78 and pg. 81, sect. 5). Selecting popular devices for inclusion in the lookup table inherently requires recognition of a market coverage rate of the devices.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure:
The following references are directed to identifying optimized GPU work group/thread block sizes: “Auto-Generation and Auto-Tuning of 3D Stencil Codes on GPU Clusters”; “On-Device Neural Net Inference with Mobile GPUs”; “Parameter based tuning model for optimizing performance on GPU”; US 20150187040 A1.
US 20190102212 A1 is directed to various examples for platform independent GPU profiles for improved utilization.
Any inquiry of a general nature or relating to the status of this application or concerning this communication or earlier communications from the Examiner should be directed to Paul Mills whose telephone number is 571-270-5482. The Examiner can normally be reached on Monday-Friday 11:00am-8:00pm. If attempts to reach the examiner by telephone are unsuccessful, the Examiner’s supervisor, April Blair can be reached at 571-270-1014.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/P. M./
Paul Mills
08/08/2026
/APRIL Y BLAIR/Supervisory Patent Examiner, Art Unit 2196