DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
The Office Action is in response to claims filed 08/03/2026.
Claims 1-23 are pending.
Claim Objections
Claims 24-26 were not elected by Applicant and should be withdrawn. Applicant elected claims 1-23. For the purposes of examination, claims 24-26 are withdrawn and claims 1-23 will be examined.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 22 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “appropriate” in claim 22 is a relative term which renders the claim indefinite. The term “appropriate” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The “number of thread blocks” has been rendered indefinite.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-23 are rejected under 35 U.S.C. 101 because the claimed invention recites a judicial exception, an abstract idea, and it has not been integrated into practical application and the claims further do not recite significantly more than the judicial exception. Examiner has evaluated the claims under the framework provided in the 2019 Patent Eligibility Guidance published in the Federal Register 01/07/2019 and has provided such analysis below.
Step 1: Claims 1-12 are directed to a computing method and fall within the statutory category of process. Claims 13-23 are directed to a graphics processing unit and fall within the statutory category of machine. Therefore, “Are the claims to a process, machine, manufacture or composition of matter?” Yes.
Step 2A Prong 1:
Claims 1 and 13: The limitation “and in response to the work assignment request, dynamically assigning a work item for the at least one kernel to perform without requiring relaunching of the at least one kernel” recites a mental process. Assigning a work item is a mental process because it observes the work assignment request and the available tasks and then performs a determination step about which work item to assign. It is understood that these limitations are to be performed within a computer environment, however, the limitations can also be performed entirely in the mind.
Therefore, Yes, claims 1 and 13 recite a judicial exception. Step 2A Prong 2 will evaluate whether the claims integrate the judicial exception into a practical application.
Step 2A Prong 2:
Claims 1 and 13: The judicial exception is not integrated into a practical application. Claims 1 recites the following additional elements – “receiving a work assignment request from the at least one kernel.” Claim 13 recites similar elements in “receive a work assignment request from the launched thread block.” These limitations are considered insignificant extra-solution activities of data transmission and data gathering (MPEP § 2106.05(g)). Claim 1 additionally recites “launching at least one kernel on a processing core” and claim 13 similarly recites “the work distributor being configured to launch a thread block to execute on at least one of the plurality of processing cores.” These limitations are means to apply an exception (MPEP § 2106.05(f)). Claim 13 also recites “a graphics processing unit comprising: a work distributor, and a plurality of processing cores.” This limitation is generic computing components used as a means to apply an exception (MPEP § 2106.05(f)). These additional elements do not integrate the judicial exception into a practical application.
Therefore, “Do the claims recite additional elements that integrate the judicial exception in a practical application?” No, these additional elements do not integrate the abstract idea into a practical application and they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
After having evaluated the inquiries set forth in Steps 2A Prong 1 and 2, it has been concluded that claims 1 and 13 not only recite a judicial exception but that the claims are directed to the judicial exception as the judicial exception has not been integrated into practical application.
Step 2B:
Claims 1 and 13: The claims do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above, the additional elements only amount to insignificant extra-solution activity and means to apply an exception. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” is a well understood, routine, and conventional computer function. Launching a kernel on a processing core is data transmission and receiving a work assignment request from the kernel is data gathering. Both involve moving information in a network.
Therefore, “Do the claims recite additional elements that amount to significantly more than the judicial exception? No, these additional elements, alone or in combination, do not amount to significantly more than the judicial exception.
Having concluded analysis within the provided framework, claims 1 and 13 do not recite eligible subject matter under 35 U.S.C. § 101.
With regard to claim(s) 2 and 14, it recites “further including the at least one kernel executing a programmatic instruction to generate the work assignment request.” This limitation is means to apply an exception (MPEP § 2106.05(f)) because it amounts to using computer components to “apply it.” It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 2 and 14 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 3 and 15, it recites “wherein dynamically assigning comprises sending the at least one kernel a work identifier indexing into a three dimensional grid array.” This limitation is insignificant extra-solution activity of data transmission (MPEP § 2106.05(g)). It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” is a well understood, routine, and conventional computer function. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 3 and 15 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 4 and 16, it recites “wherein dynamically assigning includes broadcasting or multicasting a response to a plurality of kernels.” This limitation is insignificant extra-solution activity of data transmission (MPEP § 2106.05(g)). It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” is a well understood, routine, and conventional computer function. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 4 and 16 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 5, it recites “wherein dynamically assigning work items for the kernels to perform in response to plural work assignment requests from the plurality of kernels load balances between the plurality of kernels.” This limitation is a mental process because assigning and load balancing are mental processes. They are steps of observing work assignment requests and a determination/planning step of assigning work items. Therefore, the claim(s) recite a judicial exception and fail(s) Step 2A Prong 1. The claim(s) do not include any additional elements that integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 5 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 6 and 18, it recites “wherein launching the at least one kernel includes giving the at least one kernel an initial work assignment to execute.” This limitation further limits the insignificant extra-solution activity of data transmission (MPEP § 2106.05(g)) in claim(s) 1 and 13 and is also considered as such. It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. When reevaluating the insignificant extra-solution activities for an inventive concept that is significantly more, the claims do not add an inventive concept that is other than what is well understood, routine, and conventional in the field. MPEP § 2106.05(d)(II) lists that “Receiving or transmitting data over a network” is a well understood, routine, and conventional computer function. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 6 and 18 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 7 and 23, it recites “wherein launching the at least one kernel includes dynamically launching additional kernels to utilize any new processing cores that become available.” This limitation further limits the means to apply an exception (MPEP § 2106.05(f)) in claim(s) 1 and 13 and is also considered as such. It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination with other limitations, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 7 and 23 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 8, it recites “wherein the at least one kernel comprises at least one thread block.” This limitation is field of use/technological environment (MPEP § 2106.05(h)) because it limits what the kernel comprises of in the environment. It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 8 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 9 and 19, it recites “wherein the at least one kernel comprises a CTA within a CGA.” This limitation is field of use/technological environment (MPEP § 2106.05(h)) because it limits what the kernel comprises of in the environment. It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 9 and 19 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 10 and 20, it recites “wherein dynamically assigning includes assigning more than one work item for the at least one kernel to execute.” This limitation further limits the mental process in claim(s) 1 and 13 and is also considered as such. Therefore, the claim(s) recite a judicial exception and fail(s) Step 2A Prong 1. The claim(s) do not include any additional elements that integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 10 and 20 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 11 and 21, it recites “further including persistently executing the at least one kernel on the processing core.” This limitation is means to apply an exception (MPEP § 2106.05(f)). It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 10 and 20 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 12 and 22, it recites “automatically choosing a number of kernels to be launched on processing cores to support persistent execution of a total amount of work.” This limitation is a mental process because choosing a number of kernels is observing the kernels and determining a number based on the observation. Therefore, the claim(s) recite a judicial exception and fail(s) Step 2A Prong 1. The claim(s) further recite “further including specifying a total amount of work for persistent execution without explicitly specifying a number of kernels to be launched on processing cores.” This limitation is field of use/technological environment (MPEP § 2106.05(h)) because it limits the environment of specifying a total amount of work. It does not integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 10 and 20 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
With regard to claim(s) 17, it recites “wherein the work distributor is configured to selectively decline work assignment requests in order to load balance based at least in part on responses from the executing thread blocks.” This limitation is a mental process because it observes responses and forms a determination/judgement based on the responses. Therefore, the claim(s) recite a judicial exception and fail(s) Step 2A Prong 1. The claim(s) do not include any additional elements that integrate the judicial exception into a practical application, so the claim(s) fail Step 2A Prong 2. When reevaluating the claim limitations, alone or in combination, no inventive concept that amounts to significantly more was found. Therefore, the claim(s) fail Step 2B. Therefore, claim(s) 10 and 20 do/does not recite patent eligible subject matter under 35 U.S.C. 101.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 6, 8-11, 13-15, and 18-21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Perelygin et al. Pat. No. US 11080111 B1 (hereafter Perelygin) in view of Egger et al. Pat. No. US 20180181443 A1 (hereafter Egger).
With regard to claim 1, Perelygin teaches a computing method comprising (Col. 69 Line 67 – Col. 70 Line 3 states “Terms “system” and “method” are used herein interchangeably insofar as system may embody one or more methods and methods may be considered a system.”):
launching at least one kernel on a processing core (Col. 5 Lines 28-35 states “a persistent kernel 214 is a sequence of instructions implementing a portion of operations from a CPU thread 212 that are loaded when a CPU thread 212 is launched and executed on a GPU when needed by an associated CPU thread 212. In at least one embodiment, a persistent kernel 214 remains on a GPU in a wait state 228 during CPU thread 212 execution and said persistent kernel 214 performs work when requested by said CPU thread 212.”)
and in response to the work assignment request, dynamically assigning a work item for the at least one kernel to perform without requiring relaunching of the at least one kernel (Col. 5 Lines 42-45 states “when a CPU thread 212, during persistent kernel operation 204, requires work to be performed by its persistent kernel 214, said CPU thread 212 requests work to be performed.” Col. 5 Lines 50-55 states “In at least one embodiment, a persistent kernel 214, during operation 204, will then wait 228 for further requests from an associated CPU thread 212. In at least one embodiment, additional requests by a CPU thread 1912 during persistent kernel operation 204 do not require a persistent kernel 214 to be re-initialized.”).
Perelygin does not explicitly teach receiving a work assignment from the kernel.
However, in an analogous art, Egger teaches launching at least one kernel on a processing core (¶ [0057] states “a leaf control core receives information about work groups allocated by high-level control cores and may directly allocate the work groups to the processing cores.” ¶ [0004] states “OpenCL is an open general-purpose parallel computing framework for programs executing across multiple platforms such as general-purpose multi-core central processing units (CPUs), field-programmable gate arrays (FPGAs), and graphics processing units (GPUs).” ¶ [0005] states “At least one embodiment of the inventive concept provides a method for processing an Open Computing Language (OpenCL) kernel and a computing device for the method, wherein a hierarchical control core group allocates work groups for executing the OpenCL kernel to a processing core group.” Examiner’s Note: the work group that executes on a core is a kernel),
receiving a work assignment request from the at least one kernel (¶ [0062] states “the leaf control core 550 may receive a request for an additional work group from the processing core 560.”),
and in response to the work assignment request, dynamically assigning a work item for the at least one kernel to perform without requiring relaunching of the at least one kernel (¶ [0062] states “if a work group allocated to a processing core 560 has completely processed, a work group that has been allocated to another processing core may be re-allocated to the processing core 560 by a leaf control core 550”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the core requesting additional work groups from the leaf control core and the assignment of work groups of Egger with the persistent kernel launching of Perelygin. As a result, the persistent kernel can request additional tasks. A person having ordinary skill in the art would have been motivated to make this combination because reallocating work groups in response to a core becoming available avoids having tasks waiting to be executed by the processor they were originally allocated to (¶ [0064] states “For example, if a first processing core has finished processing a first workgroup and a second processing core is still processing a second workgroup, instead of waiting for the second processing core to complete processing of the second work group, the leaf control core associated with the second processing core can send a third workgroup to the first processing core even though the third workgroup was scheduled to be next processed on the second processing core”). In this way, work groups are executed more efficiently. Further, ¶ [0121] states “since control cores for allocating work groups to processing cores are hierarchically grouped, work groups may be efficiently distributed.”
With regard to claim 2, Perelygin and Egger teach the computing method of claim 1. Egger additionally teaches further including the at least one kernel executing a programmatic instruction to generate the work assignment request (¶ [0062] states “In this case, the leaf control core 550 may receive a request for an additional work group from the processing core.” Examiner’s Note: the request from the processing core is generated by an instruction).
With regard to claim 3, Perelygin and Egger teach the computing method of claim 1. Perelygin additionally teaches wherein dynamically assigning comprises sending the at least one kernel a work identifier indexing into a three dimensional grid array (Col. 67 Lines 19-24 states “FIG. 38 illustrates how threads of an exemplary CUDA grid 3820 are mapped to different compute units 3740 of FIG. 37, in accordance with at least one embodiment. In at least one embodiment and for explanatory purposes only, grid 3820 has a GridSize of BX by BY by 1 and a BlockSize of TX by TY by 1.” See FIG. 38).
With regard to claim 6, Perelygin and Egger teach the computing method of claim 1. Perelygin additionally teaches wherein launching the at least one kernel includes giving the at least one kernel an initial work assignment to execute (Col. 5 Lines 39-44 states “persistent kernel operation 204 comprises a CPU thread 212 that begins by performing work including initializing or loading a persistent kernel 214. In at least one embodiment, when a CPU thread 212, during persistent kernel operation 204, requires work to be performed by its persistent kernel 214, said CPU thread 212 requests work to be performed.”).
With regard to claim 8, Perelygin and Egger teach the computing method of claim 1. Egger additionally teaches wherein the at least one kernel comprises at least one thread block (¶ [0005] states “a hierarchical control core group allocates work groups for executing the OpenCL kernel to a processing core group.”).
With regard to claim 9, Perelygin and Egger teach the computing method of claim 1. Perelygin additionally teaches wherein the at least one kernel comprises a CTA within a CGA (Col. 7 Lines 7-10 states “In at least one embodiment, a grid or CG 302 is a group of thread blocks or CTAs 304 as well as threads 304 that can be individually synchronized with other threads 306 or thread blocks 304 in a kernel 308.” See FIG. 3 and 5).
With regard to claim 10, Perelygin and Egger teach the computing method of claim 1. Perelygin additionally teaches wherein dynamically assigning includes assigning more than one work item for the at least one kernel to execute (Col. 5 Lines 42-45 states “In at least one embodiment, when a CPU thread 212, during persistent kernel operation 204, requires work to be performed by its persistent kernel 214, said CPU thread 212 requests work to be performed.” Col. 5 Lines 50-52 states “In at least one embodiment, a persistent kernel 214, during operation 204, will then wait 228 for further requests from an associated CPU thread 212.” See FIG. 2. Examiner’s Note: in FIG. 2 Persistent Kernel 204, the persistent kernel is assigned work twice in the two “Perform Work” boxes).
With regard to claim 11, Perelygin and Egger teach the computing method of claim 1. Perelygin additionally teaches further including persistently executing the at least one kernel on the processing core (Col. 5 Lines 28-38 states “a persistent kernel 214 is a sequence of instructions implementing a portion of operations from a CPU thread 212 that are loaded when a CPU thread 212 is launched and executed on a GPU when needed by an associated CPU thread 212. In at least one embodiment, a persistent kernel 214 remains on a GPU in a wait state 228 during CPU thread 212 execution and said persistent kernel 214 performs work when requested by said CPU thread 212. In at least one embodiment, a kernel 214 is persistent if it remains on a GPU during an associated CPU thread's 212 execution lifecycle.”).
With regard to claim 13, Perelygin teaches a graphics processing unit comprising (Col. 3 Lines 13-16 states “FIG. 1 is a block diagram illustrating a multi-process service for executing multiple kernels 104 from multiple central processing unit (CPU) processes on a graphics processing unit (GPU)”):
a work distributor (Col. 5 Lines 42-45 states “when a CPU thread 212, during persistent kernel operation 204, requires work to be performed by its persistent kernel 214, said CPU thread 212 requests work to be performed.” Examiner’s Note: the CPU thread is the work distributor),
and a plurality of processing cores (Col. 46 Lines 51-54 states “each DPC 2506 included in GPC 2500 comprise … one or more SMs 2514”),
the work distributor being configured to launch a thread block to execute on at least one of the plurality of processing cores (Col. 5 Lines 28-35 states “a persistent kernel 214 is a sequence of instructions implementing a portion of operations from a CPU thread 212 that are loaded when a CPU thread 212 is launched and executed on a GPU when needed by an associated CPU thread 212. In at least one embodiment, a persistent kernel 214 remains on a GPU in a wait state 228 during CPU thread 212 execution and said persistent kernel 214 performs work when requested by said CPU thread 212.” Col. 47 Lines 48-50 states “scheduler unit 2604 receives tasks from a work distribution unit and manages instruction scheduling for one or more thread blocks assigned to SM 2600.” Col. 7 Lines 7-10 states “In at least one embodiment, a grid or CG 302 is a group of thread blocks or CTAs 304 as well as threads 304 that can be individually synchronized with other threads 306 or thread blocks 304 in a kernel 308.” Examiner’s Note: thread blocks is considered synonymous with kernel),
and in response to the work assignment request, dynamically assign a work item for the thread block to execute without requiring relaunch of the thread block (Col. 5 Lines 42-45 states “when a CPU thread 212, during persistent kernel operation 204, requires work to be performed by its persistent kernel 214, said CPU thread 212 requests work to be performed.” Col. 5 Lines 50-55 states “In at least one embodiment, a persistent kernel 214, during operation 204, will then wait 228 for further requests from an associated CPU thread 212. In at least one embodiment, additional requests by a CPU thread 1912 during persistent kernel operation 204 do not require a persistent kernel 214 to be re-initialized.”).
Perelygin does not explicitly teach receiving a work assignment from the kernel.
However, in an analogous art, Egger teaches the work distributor being configured to launch a thread block to execute on at least one of the plurality of processing cores (¶ [0057] states “a leaf control core receives information about work groups allocated by high-level control cores and may directly allocate the work groups to the processing cores.” ¶ [0004] states “OpenCL is an open general-purpose parallel computing framework for programs executing across multiple platforms such as general-purpose multi-core central processing units (CPUs), field-programmable gate arrays (FPGAs), and graphics processing units (GPUs).” ¶ [0005] states “At least one embodiment of the inventive concept provides a method for processing an Open Computing Language (OpenCL) kernel and a computing device for the method, wherein a hierarchical control core group allocates work groups for executing the OpenCL kernel to a processing core group.” ¶ [0044] states “The control core group 210 is a group of control cores allocating work groups for executing an OpenCL kernel to low-level control cores and processing cores.” See FIG. 3. Examiner’s Note: the work group that executes on a core is a thread block. The cores in FIG. 3 Control Core Group 320 are the work distributor),
receive a work assignment request from the launched thread block (¶ [0062] states “the leaf control core 550 may receive a request for an additional work group from the processing core 560.”),
and in response to the work assignment request, dynamically assign a work item for the thread block to execute without requiring relaunch of the thread block (¶ [0062] states “if a work group allocated to a processing core 560 has completely processed, a work group that has been allocated to another processing core may be re-allocated to the processing core 560 by a leaf control core 550”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the core requesting additional work groups from the leaf control core and the assignment of work groups of Egger with the persistent kernel launching of Perelygin. As a result, the persistent kernel/thread block can request additional tasks. A person having ordinary skill in the art would have been motivated to make this combination because reallocating work groups in response to a core becoming available avoids having tasks waiting to be executed by the processor they were originally allocated to (¶ [0064] states “For example, if a first processing core has finished processing a first workgroup and a second processing core is still processing a second workgroup, instead of waiting for the second processing core to complete processing of the second work group, the leaf control core associated with the second processing core can send a third workgroup to the first processing core even though the third workgroup was scheduled to be next processed on the second processing core”). In this way, work groups are executed more efficiently. Further, ¶ [0121] states “since control cores for allocating work groups to processing cores are hierarchically grouped, work groups may be efficiently distributed.”
With regard to claim 14, Perelygin and Egger teach the graphics processing unit of claim 13. Egger additionally teaches further including the at least one of the plurality of processing cores executing a programmatic instruction within the thread block to generate the work assignment request (¶ [0062] states “In this case, the leaf control core 550 may receive a request for an additional work group from the processing core.” Examiner’s Note: the request from the processing core is generated by an instruction).
With regard to claim 15, Perelygin and Egger teach the graphics processing unit of claim 13. Perelygin additionally teaches wherein the work distributor is further configured to send the launched thread block a work identifier indexing into a three dimensional grid array (Col. 67 Lines 19-24 states “FIG. 38 illustrates how threads of an exemplary CUDA grid 3820 are mapped to different compute units 3740 of FIG. 37, in accordance with at least one embodiment. In at least one embodiment and for explanatory purposes only, grid 3820 has a GridSize of BX by BY by 1 and a BlockSize of TX by TY by 1.” See FIG. 38).
With regard to claim 18, Perelygin and Egger teach the graphics processing unit of claim 13. Perelygin additionally teaches wherein the work distributor gives the thread block an initial work assignment to execute at launch (Col. 5 Lines 39-44 states “persistent kernel operation 204 comprises a CPU thread 212 that begins by performing work including initializing or loading a persistent kernel 214. In at least one embodiment, when a CPU thread 212, during persistent kernel operation 204, requires work to be performed by its persistent kernel 214, said CPU thread 212 requests work to be performed.”).
With regard to claim 19, Perelygin and Egger teach the graphics processing unit of claim 13. Perelygin additionally teaches wherein the thread block comprises a CTA within a CGA (Col. 7 Lines 7-10 states “In at least one embodiment, a grid or CG 302 is a group of thread blocks or CTAs 304 as well as threads 304 that can be individually synchronized with other threads 306 or thread blocks 304 in a kernel 308.” See FIG. 3 and 5).
With regard to claim 20, Perelygin and Egger teach the graphics processing unit of claim 13. Egger additionally teaches Perelygin additionally teaches wherein the work distributor is further configured to assign more than one work item to the thread block to execute (Col. 5 Lines 42-45 states “In at least one embodiment, when a CPU thread 212, during persistent kernel operation 204, requires work to be performed by its persistent kernel 214, said CPU thread 212 requests work to be performed.” Col. 5 Lines 50-52 states “In at least one embodiment, a persistent kernel 214, during operation 204, will then wait 228 for further requests from an associated CPU thread 212.” See FIG. 2. Examiner’s Note: in FIG. 2 Persistent Kernel 204, the persistent kernel is assigned work twice in the two “Perform Work” boxes).
With regard to claim 21, Perelygin and Egger teach the graphics processing unit of claim 13. Perelygin additionally teaches wherein one of the processing cores persistently executes the thread block (Col. 5 Lines 28-38 states “a persistent kernel 214 is a sequence of instructions implementing a portion of operations from a CPU thread 212 that are loaded when a CPU thread 212 is launched and executed on a GPU when needed by an associated CPU thread 212. In at least one embodiment, a persistent kernel 214 remains on a GPU in a wait state 228 during CPU thread 212 execution and said persistent kernel 214 performs work when requested by said CPU thread 212. In at least one embodiment, a kernel 214 is persistent if it remains on a GPU during an associated CPU thread's 212 execution lifecycle.”).
Claim(s) 4-5 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Perelygin in view of Egger and further in view of Mei et al. Pat. No. US 20240220254 A1 (hereafter Mei).
With regard to claim 4, Perelygin and Egger teach the computing method of claim 1. Perelygin and Egger do not explicitly teach broadcasting or multicasting.
However, in an analogous art, Mei teaches wherein dynamically assigning includes broadcasting or multicasting a response to a plurality of kernels (¶ [0360] states “In some embodiments, an apparatus, system, or process provides for data broadcasting (also referred to herein as data multicasting) directly between compute cores in a compute core cluster (also referred to as a partition) in a processing architecture, such as a GPU architecture. In some embodiments, a cluster of compute cores utilizes high-bandwidth core interconnections to broadcast fetched data from a producer compute core to one or more additional compute cores”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the data broadcasting of Mei with the dynamic assignment of work groups of Egger and persistent kernels of Perelygin. As a result, the work groups are broadcast to kernels. A person having ordinary skill in the art would have been motivated to make this combination for the purpose of “reducing or eliminating repeated transmissions of data elements to provide the fetched data to each consumer compute core” (¶ [0360]).
With regard to claim 5, Perelygin, Egger, and Mei teach the computing method of claim 4. Egger additionally teaches wherein dynamically assigning work items for the kernels to perform in response to plural work assignment requests from the plurality of kernels load balances between the plurality of kernels (¶ [0064] states “For example, if a first processing core has finished processing a first workgroup and a second processing core is still processing a second workgroup, instead of waiting for the second processing core to complete processing of the second work group, the leaf control core associated with the second processing core can send a third workgroup to the first processing core even though the third workgroup was scheduled to be next processed on the second processing core.”).
With regard to claim 16, Perelygin and Egger teach the graphics processing unit of claim 13. Perelygin and Egger do not explicitly teach broadcasting or multicasting.
However, in an analogous art, Mei teaches wherein the work distributor is further configured to cause broadcast or multicast of a response to a plurality executing thread blocks on a respective plurality of processing cores (¶ [0360] states “In some embodiments, an apparatus, system, or process provides for data broadcasting (also referred to herein as data multicasting) directly between compute cores in a compute core cluster (also referred to as a partition) in a processing architecture, such as a GPU architecture. In some embodiments, a cluster of compute cores utilizes high-bandwidth core interconnections to broadcast fetched data from a producer compute core to one or more additional compute cores”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the data broadcasting of Mei with the dynamic assignment of work groups of Egger and persistent kernels of Perelygin. As a result, the work groups are broadcast to thread blocks. A person having ordinary skill in the art would have been motivated to make this combination for the purpose of “reducing or eliminating repeated transmissions of data elements to provide the fetched data to each consumer compute core” (¶ [0360]).
Claim(s) 7 and 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Perelygin in view of Egger and further in view of Oro Garcia et al. Pat. No. US 20130243329 A1 (hereafter Oro Garcia).
With regard to claim 7, Perelygin and Egger teach the computing method of claim 1. Perelygin and Egger do not explicitly teach launching additional kernels as new processing cores become available.
However, in an analogous art, Oro Garcia teaches wherein launching the at least one kernel includes dynamically launching additional kernels to utilize any new processing cores that become available (¶ [0039] states “If the number of thread blocks 902 is greater than the number of cores available in the PPU 102, the remaining thread blocks 902 are dynamically enqueued and dequeued as the PPU 102 resources of each core become available.”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the launching of additional thread blocks as cores become available of Oro Garcia with the launching of kernels of Perelygin and Egger. A person having ordinary skill in the art would have been motivated to make this combination because it “achieves the full occupancy of both the cores and the functional units of the abovementioned computer system, while reducing the latency of memory operations through an efficient usage of the underlying cache hierarchy. Thus, the overall speed and throughput of the object detection process is dramatically improved relative to prior art techniques” (¶ [0011]).
With regard to claim 23, Perelygin and Egger teach the graphics processing unit of claim 13. Perelygin and Egger do not explicitly teach launching additional kernels as new processing cores become available.
However, in an analogous art, Oro Garcia teaches wherein the work distributor dynamically launches additional thread blocks to utilize any new processing cores that become available (¶ [0039] states “If the number of thread blocks 902 is greater than the number of cores available in the PPU 102, the remaining thread blocks 902 are dynamically enqueued and dequeued as the PPU 102 resources of each core become available.”).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the launching of additional thread blocks as cores become available of Oro Garcia with the launching of kernels of Perelygin and Egger. A person having ordinary skill in the art would have been motivated to make this combination because it “achieves the full occupancy of both the cores and the functional units of the abovementioned computer system, while reducing the latency of memory operations through an efficient usage of the underlying cache hierarchy. Thus, the overall speed and throughput of the object detection process is dramatically improved relative to prior art techniques” (¶ [0011]).
Claim(s) 12 and 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Perelygin in view of Egger and further in view of Levin et al. Pat. No. US 20160239348 A1 (hereafter Oro Garcia).
With regard to claim 12, Perelygin and Egger teach the computing method of claim 1. Perelygin additionally teaches further including specifying a total amount of work for persistent execution without explicitly specifying a number of kernels to be launched on processing cores, and automatically choosing a number of kernels to be launched on processing cores to support persistent execution of a total amount of work (Col. 5 Lines 5-7 states “In at least one embodiment, a CPU thread 206, 212 offloads a portion of its operations to be performed on a GPU by a kernel 208, 210, 214.” Col. 5 Lines 21-25 states “In at least one embodiment, multiple CPU threads 226, 212 submit one or more kernels 208, 210, 214 to a multi-process service, as described above, resulting in one or more kernels 208, 210, 214 being executed on one or more GPUs.”).
Perelygin and Egger do not explicitly teach that specifying a total amount of work for persistent execution is done without explicitly specifying a number of kernels to be launched on processing cores.
However, in an analogous art, Levin teaches further including specifying a total amount of work for persistent execution without explicitly specifying a number of kernels to be launched on processing cores, and automatically choosing a number of kernels to be launched on processing cores to support persistent execution of a total amount of work (¶ [0017] states “evaluating a first number of available cores of a first type of the multi-processor system and a second number of available cores of a second type of the multi-processor system; determining a first number of loops of the computational block for binding with the cores of the first type and a second number of loops of the computational block for binding with the cores of the second type;” ¶ [0020] states “Determining the first number and the second number of loops is according to a load balancing relation with respect to the available cores of the first type and the available cores of the second type reduce programmer efforts on developing parallel applications for heterogeneous hardware.” Examiner’s Note: the programmer does not have to specify the number of loops for the workload).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine evaluating available cores and then determining a number of loops to bind to the available cores of Levin with the kernel launching of Perelygin and Egger. As a result, the portion of work that needs to be offloaded to the GPU does not need to explicitly specify the number of kernels that the workload will be distributed to. Instead, the system can determine the number of kernels to launch.
A person having ordinary skill in the art would have been motivated to make this combination for the purpose of making it easier to developing parallel applications, and reducing labor costs of programmers developing or porting code to different architectures (¶ [0020] states “This kind of effect results in making the process of developing parallel application for multi-processor systems such as MMCHCS hardware easier. Before that, the programmers needed to spend a lot of time considering how to split the processors among cores, now this work can be done automatically. The specific determining leads to decreasing of labor costs of either software developing or effective porting of existing code to specific architecture.”).
With regard to claim 22, Perelygin and Egger teach the graphics processing unit of claim 13. Perelygin additionally teaches wherein an application specifies a total amount of work for persistent execution without explicitly specifying a number of thread blocks to be launched on processing cores, and the work distributor automatically chooses an appropriate number of thread blocks to launch on the processing cores to support persistent execution of the total amount of work (Col. 5 Lines 5-7 states “In at least one embodiment, a CPU thread 206, 212 offloads a portion of its operations to be performed on a GPU by a kernel 208, 210, 214.” Col. 5 Lines 21-25 states “In at least one embodiment, multiple CPU threads 226, 212 submit one or more kernels 208, 210, 214 to a multi-process service, as described above, resulting in one or more kernels 208, 210, 214 being executed on one or more GPUs.”).
Perelygin and Egger do not explicitly teach that specifying a total amount of work for persistent execution is done without explicitly specifying a number of kernels to be launched on processing cores.
However, in an analogous art, Levin teaches wherein an application specifies a total amount of work for persistent execution without explicitly specifying a number of thread blocks to be launched on processing cores, and the work distributor automatically chooses an appropriate number of thread blocks to launch on the processing cores to support persistent execution of the total amount of work (¶ [0017] states “evaluating a first number of available cores of a first type of the multi-processor system and a second number of available cores of a second type of the multi-processor system; determining a first number of loops of the computational block for binding with the cores of the first type and a second number of loops of the computational block for binding with the cores of the second type;” ¶ [0020] states “Determining the first number and the second number of loops is according to a load balancing relation with respect to the available cores of the first type and the available cores of the second type reduce programmer efforts on developing parallel applications for heterogeneous hardware.” Examiner’s Note: the programmer does not have to specify the number of loops for the workload).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine evaluating available cores and then determining a number of loops to bind to the available cores of Levin with the kernel launching of Perelygin and Egger. As a result, the portion of work that needs to be offloaded to the GPU does not need to explicitly specify the number of kernels that the workload will be distributed to. Instead, the system can determine the number of kernels to launch.
A person having ordinary skill in the art would have been motivated to make this combination for the purpose of making it easier to developing parallel applications, and reducing labor costs of programmers developing or porting code to different architectures (¶ [0020] states “This kind of effect results in making the process of developing parallel application for multi-processor systems such as MMCHCS hardware easier. Before that, the programmers needed to spend a lot of time considering how to split the processors among cores, now this work can be done automatically. The specific determining leads to decreasing of labor costs of either software developing or effective porting of existing code to specific architecture.”).
Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Perelygin in view of Egger and Mei and further in view of Xu et al. Pat. No. US 20140181839 A1 (hereafter Xu).
With regard to claim 17, Perelygin, Egger, and Mei teach the graphics processing unit of claim 16. Perelygin, Egger, and Mei do not explicitly teach the work distributor being configured to selectively decline work assignment requests.
However, in an analogous art, Xu teaches wherein the work distributor is configured to selectively decline work assignment requests in order to load balance based at least in part on responses from the executing thread blocks (¶ [0026] states “In step S301, the task executing node sends a request for acquiring a task to the scheduling node, wherein the request carries a current load value and an available memory space of the task executing node therein.” ¶ [0030] states “In step S302, the scheduling node decides whether the current load value is less than a threshold. If the decision result is "YES", that is, if the current load value is less than a threshold, step S304 is executed, and if the decision result is "NO", that is, if the current load value is greater than or equal to the threshold, step S303 is executed.” ¶ [0033] states “In step S303, the task executing node is rejected to be assigned a task.” ¶ [0022] states “The multi-task scheduling system 1 comprises a task scheduling apparatus 11 and at least one task executing apparatus 12.” See FIG. 2 and FIG. 3. Examiner’s Note: the work assignment request is also considered a response because it is from the task executing node. There are a plurality of task executing nodes, so it is understood that there are multiple responses. By considering the load value, the scheduler can selectively reject work assignment requests in order to load balance).
It would have been obvious to a person having ordinary skill in the art prior to the effective filing date to combine the scheduler rejecting a request for acquiring of task of Xu with the persistent kernel of Perelygin, the processing core requesting additional work of Egger, and the broadcasting of Mei. As a result, when the cores of Egger request additional work, the cores in the control core group can selectively decline requests for additional work in order to load balance between cores. A person having ordinary skill in the art would have been motivated to make this combination for the purpose of reducing problems of overload, increasing the utilization ratio of the resources, and improve the efficiency of task scheduling and executing (¶ [0012] states “so as to reduce problems of overload and insufficient of the load and memory of the task executing node effectively, and increase the utilization ratio of the resource of the task executing node and efficiency of task scheduling and executing.”).
Response to Amendment
Claims 24-26 have an improper status identifier. The claims should be withdrawn because the Remarks/Arguments filed 08/03/2026 elected Group I (claims 1-23). As explained in the claim objections, claims 24-26 are withdrawn for the purposes of examination. Only claims 1-23 will be examined.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20220206851 A1
teaches
REGENERATIVE WORK-GROUPS
US 20140237477 A1
teaches
SIMULTANEOUS SCHEDULING OF PROCESSES AND OFFLOADING COMPUTATION ON MANY-CORE COPROCESSORS
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PETER L YUAN whose telephone number is (571)272-5737. The examiner can normally be reached Mon-Fri 7:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at 571-272-3338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PETER LI YUAN/Examiner, Art Unit 2197
/BRADLEY A TEETS/Supervisory Patent Examiner, Art Unit 2197