Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claims 1, 3-8, 10-15, and 17-23 have been considered but are moot because the arguments do not apply to any of the references being used in the current rejection
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made
Claims 1, 3-8, 10-15, and 17-23 are rejected under 35 U.S.C. 103 as being unpatentable over Hirota et al. (US 2023/0289211, hereinafter Hirota) in view of Fontaine et al. (US 2023/0222619), and further in view of Storer et al. (US 2020/0382572).
Regarding claim 1, Hirota discloses
One or more processors comprising (fig. 1-28):
circuitry to, in response to an application programming interface (API) call (paragraphs [0027]-[0029]: Cooperative Groups API for defining and synchronizing groups of threads; paragraph [0113]: dynamic parallelism API … can be nested to further control the placement of SM resources at each level), allocate one or more data structures to indicate which of one or more processing units of one or more processors are to be used to perform one or more software threads (paragraph [0182]: the CWD 420 generates sm_masks from the “launch” registers accumulated in the speculative/shadow state launch process (this sm_masks data structure stores reservation information (FIG. 16A, block 2010; see FIG. 16B) for each CTA to be run on each SM in the relevant hardware domain for the CGA launch), and moves on to a next CGA. The hardware allocates a CGA sequential number and attaches it to each sm_mask, which specifies which SM gets which CTA at launch; paragraph [0226]: Mapping table 5004 identifies to the SM which other SMs the other CTAs in the CGA are executing on … CTA_ID -> CGA_ID, SM_ID, and TPC_ID; Note: The hardware circuitry (CWD 420) responds to the API call by allocating data structures (sm_masks) that explicitly indicate which processing units (SMs) will execute which software threads (CTAs)).
Hirota does not disclose API call including arguments comprising a memory location to store a sub-context corresponding to the one or more data structures, the sub-context corresponding to a subset of resources of an existing context and configured to avoid context switching. Fontaine discloses API call including arguments comprising a memory location to store a sub-context corresponding to the one or more data structures, the sub-context corresponding to a subset of resources of an existing context (paragraph [0053]: a parameter in the call to the load API is a pointer to a location in memory 128; paragraph [0064]: parameter const void *image is a pointer to a memory location for contextual information to be loaded; paragraph [0054]: processor 104 loads the set of contextual information specified by the call into existing GPU contexts; Note: A module or set of contextual information loaded into an existing context implicitly acts as a sub-context that utilizes a subset of the resources of that existing context). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to configure the API of Hirota to include API arguments comprising a memory location pointer to load contextual information (i.e., a sub-context) into an existing context as taught by Fontaine. The motivation would have been to provide performance advantages and enable a use of contexts on logical GPUs to provide greater scaling ability while also maintaining an efficient allocation of memory, without needing to track and control loading into each particular context (Fontaine paragraph [0055]).
Hirota in view of Fontaine does not disclose configured to avoid context switching. Storer discloses configured to avoid context switching (paragraph [0033]: passing the IR code to the driver for compilation and execution has the benefit of avoiding the need for a context switch into the user application; paragraph [0038]: removing the sandbox so context switches into the code may be faster or avoided, reducing latency and introducing less overhead). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to configure the API of Hirota in view of Fontaine to execute the distributed contextual information directly at the driver level to avoid context switching as taught by Storer. The motivation would have been to avoid the need for a context switch into the user application, thereby reducing latency and introducing less overhead, which improves the power efficiency of the processor running the code (Storer paragraphs [0033] and [0038]).
Regarding claim 8 referring to claim 1, Hirota discloses A computer-implemented method comprising: … (See the rejection for claim 1).
Regarding claim 15 referring to claim 1, Hirota discloses A computer system comprising: one or more processors and memory storing executable instructions that, if executed d by the one or more processors (paragraph [0020]).
Regarding claims 3 and 17, Hirota discloses
wherein the API call includes additional arguments comprising a resource descriptor indicative of the one or more processing units (paragraph [0097]: a CGA can define or specify the hardware domain on which all CTAs in the CGA shall run; paragraph [0103]: hardware domain specifier associated with that certain CGA; paragraph [0113]: These example levels (Grid, GPU_CGA, μGPU_CGA, GPC_CGA, and CTA—see FIG. 10B) can be nested to further control the placement of SM resources at each level; paragraph [0182]: the CWD 420 generates sm_masks… which specifies which SM gets which CTA at launch; Note: The API call used to launch the thread groups includes additional arguments/specifiers acting as a resource descriptor (e.g., the hardware domain specifier) that indicates the specific placement of SMs (processing units) to be used).
Regarding claims 4, 11, and 18, Hirota discloses
wherein the API call includes additional arguments comprising an indication of a device comprising the one or more processing units upon which the circuitry is to perform at least a portion of operations in response to the API call (paragraph [0097]: a CGA can define or specify the hardware domain on which all CTAs in the CGA shall run. By way of analogy… a CGA could require the CTAs it references to all run on the same portion (GPC and/or μGPU) of a GPU, on the same GPU, on the same cluster of GPUs, etc.; paragraph [0103]: hardware domain specifier associated with that certain CGA, for example: all the CTAs for a GPU_CGA are launched onto SMs that are part of the same GPU; paragraph [0118]: Grid of GPU_CGAs of CTAs—This is a grid where the CTAs for each GPU_CGA are launched together and always placed on the same GPU. Thus, the hardware domain specified by this type of grid is “GPU.”; Note: The API call used to launch the thread groups includes arguments/specifiers, such as the hardware domain specifier, that explicitly indicate the specific device (e.g., a specific GPU) containing the SMs (processing units) where the operations are to be performed).
Regarding claims 5, 12, and 19, Hirota discloses
wherein the API call includes additional arguments comprising a set of flags at least indicating which processing units of a device comprising the one or more processing units are to be used to perform the one or more software threads (paragraph [0103]: hardware domain specifier associated with that certain CGA, for example: all the CTAs for a GPU_CGA are launched onto SMs that are part of the same GPU; paragraph [0113]: These example levels … can be nested to further control the placement of SM resources at each level; paragraph [0182]: the CWD 420 generates sm_masks … The hardware allocates a CGA sequential number and attaches it to each sm_mask, which specifies which SM gets which CTA at launch; Note: The API call includes arguments/specifiers that act as or generate a set of flags (e.g., the “sm_masks”, which are bitmasks/flags) to explicitly indicate which specific processing units (SMs) of the device are to be used to perform the software threads (CTAs)).
Regarding claims 6 and 13, Hirota does not disclose wherein the circuitry is further to indicate that the sub-context has been created. Fontaine discloses wherein the circuitry is further to indicate that the sub-context has been created paragraph [0066]: a response 202 to load contextual information API call 200 includes an operation status. In at least one embodiment, response 202 to load contextual information API call 200 indicates if load contextual information API call 200 was successful; paragraph [0061]: Identifier parameter is an outgoing parameter from load contextual information API call 200 that provides a handle to loaded contextual information; Note: Returning a successful operation status or a handle for the newly loaded contextual information explicitly indicates to the calling program that the contextual information, acting as the sub-context, has been successfully loaded/created). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to configure the API of the combined system of Hirota to indicate that the sub-context has been created, as taught by Fontaine. The motivation would have been to inform the calling application whether the API execution succeeded or failed, allowing the application to interpret the status and determine the appropriate subsequent actions to take, such as proceeding with execution using the valid handle or handling errors appropriately (Fontaine paragraph [0066]).
Regarding claims 7, 14, and 20, Hirota discloses
wherein the one or more data structures indicate a partitioning of a plurality of processing units of the one or more processors into at least the one or more processing units and an additional one or more processing units not including the one or more processing units (paragraph [0024]: GPU implementations … enable plural partitions that operate as micro GPUs such as μGPU0 and μGPU1, where each micro GPU includes a portion of the processing resources of the overall GPU. When the GPU is partitioned into two or more separate smaller μGPUs for access by different clients, resources… are also typically partitioned; paragraph [0097]: a CGA could require the CTAs it references to all run on the same portion (GPC and/or μGPU) of a GPU; paragraph [0182]: the CWD 420 generates sm_masks … which specifies which SM gets which CTA at launch; Note: The data structures, such as the sm_masks and hardware domain specifiers, indicate this partitioning by assigning the software threads to a specific partitioned portion of the processing units (e.g., μGPU0), which inherently leaves the other partitioned portions (e.g., μGPU1) as the additional processing units not included in the first set).
Regarding claim 10, Hirota discloses
wherein the API call includes additional arguments comprising an indication of the one or more processing units (paragraph [0097]: a CGA can define or specify the hardware domain on which all CTAs in the CGA shall run. By way of analogy … a CGA could require the CTAs it references to all run on the same portion (GPC and/or μGPU) of a GPU, on the same GPU, on the same cluster of GPUs, etc.; paragraph [0103]: hardware domain specifier associated with that certain CGA, for example: all the CTAs for a GPU_CGA are launched onto SMs that are part of the same GPU; paragraph [0113]: These example levels (Grid, GPU_CGA, μGPU_CGA, GPC_CGA, and CTA—see FIG. 10B) can be nested to further control the placement of SM resources at each level; Note: The API call receives arguments, such as the hardware domain specifier, that explicitly indicate the specific processing units (e.g., the specific SMs, GPC, or μGPU) where the threads are to be executed).
Regarding claim 21, Hirota does not disclose wherein the circuitry is further to return a success indicator to a calling process in response to storing the sub-context in the memory location. Fontaine discloses wherein the circuitry is further to return a success indicator to a calling process in response to storing the sub-context in the memory location (paragraph [0066]: a response 202 to load contextual information API call 200 includes an operation status. In at least one embodiment, response 202 to load contextual information API call 200 indicates if load contextual information API call 200 was successful). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to configure the API of Hirota to include API arguments comprising a memory location pointer to load contextual information (i.e., a sub-context) into an existing context as taught by Fontaine. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to configure the API of the combined system to return a success indicator to a calling process as taught by Fontaine. The motivation would have been to provide an operation status that indicates if load contextual information API call 200 was successful, has failed, or if other errors have occurred so that the status can be interpreted by a calling application and/or library (e.g., using a mapping of operation status numeric identifiers to their meaning and/or actions to take) (Fontaine paragraph [0066]).
Regarding claim 22, Hirota discloses
further comprising reserving, further in response to the API call, the subset of resources for the sub-context (paragraph [0182]: the CWD 420 generates sm_masks from the “launch” registers accumulated in the speculative/shadow state launch process (this sm_masks data structure stores reservation information (FIG. 16A, block 2010; see FIG. 16B) for each CTA to be run on each SM in the relevant hardware domain for the CGA launch); Note: As established in claim 1, the API call triggers the generation of these data structures, which implicitly store reservation information to reserve the specific processing resources (SMs) for the threads/sub-context).
Regarding claim 23, Hirota does not disclose wherein the computer system is further to determine whether the API call is valid by verifying validity of the arguments prior to allocation of the one or more data structures. Fontaine discloses wherein the computer system is further to determine whether the API call is valid by verifying validity of the arguments prior to allocation of the one or more data structures (paragraph [0066]: a response 202 to load contextual information API call 200 includes an operation status. In at least one embodiment, response 202 to load contextual information API call 200 indicates if load contextual information API call 200 was successful, has failed, or if other errors have occurred). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to configure the API of Hirota to include API arguments comprising a memory location pointer to load contextual information (i.e., a sub-context) into an existing context as taught by Fontaine. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to configure the computer system of the combined references to determine whether the API call is valid by verifying the validity of the arguments prior to allocation as taught by Fontaine’s error checking. The motivation would have been to determine if the API call was successful, has failed, or if other errors have occurred so that the resulting status can be interpreted by a calling application and/or library (e.g., using a mapping of operation status numeric identifiers to their meaning and/or actions to take) (Fontaine paragraph [0066]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
Larson et al. (US 2022/0043731) discloses systems and methods utilizing subcontexts to group and isolate specific properties or attributes for analysis (paragraphs [0060] and [0078]). Larson further discloses executing computational tasks via API calls and utilizing runtime APIs for execution control and memory management (paragraphs [0403] and [0406]), as well as enabling rapid preemption and context switching of threads executing on a processing array (paragraph [0375]).
Duluk, Jr. et al. (US 2021/0157651) discloses configuring a processor to generate multiple different processing subcontexts within a parent processing context to assign tasks to different processes (paragraphs [0005], [0007], and [0157]). Duluk further discloses hardware units configured to manage context switches independently across different engines (paragraphs [0098], [0104], and [0117]), and intercepting API calls to route processing workloads to appropriate hardware resources (paragraphs [0091] and [0332]-[0334]).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in [0037] CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to [0037] CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SISLEY N. KIM whose telephone number is (571)270-7832. The examiner can normally be reached M-F 11:30AM -7:30PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, April Y. Blair can be reached on (571)270-1014. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SISLEY N KIM/Primary Examiner, Art Unit 2196 5/10/2026