Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-31 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 2 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 2 recites the limitation “a subset of the SMs” in line 6. There is insufficient antecedent basis for this limitation in the claim. For the purpose of examination, the “subset of the SMs” is interpreted “subset of the plurality of the SMs.” Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-31 is/are rejected under 35 U.S.C. 103 as being unpatentable over Johnson (US 2020/0081748) and further in view of Ray (US 2018/0285110) and further in view of Munshi (US 2008/0276220).
Regarding claim 1, Johnson teaches: A central processing unit (CPU) comprising: circuitry to:
cause a graphics processing unit (GPU) to reserve a subset of a plurality of streaming multiprocessors (SMs) within the GPU (¶ 188, “Volta MPS provides control for MPS clients to specify what fraction of the GPU (GPU fraction 2108, GPU fraction 2110, and GPU fraction 2112) is necessary for execution”) to perform one or more software threads (¶ 191, “the parallel processing unit 2200 is a multi-threaded processor that is implemented on one or more integrated circuit devices”).
Johnson does not teach, however, Ray teaches: one or more indications of a number of SMs to be usable by one or more contexts associated with the one or more software threads (¶ 132, “At operation 725 the thread scheduler 610 may determine whether one or more contexts have “extra” threads that don't fit within the number of streaming microprocessors (SMs) 620 allocated to the context”); and
cause the one more software threads associated with a respective context to be executed by the subset of SMs determined for the respective context (¶ 131, “At operation 720 the thread scheduler 610 dispatches threads from each context only to the streaming microprocessors (SMs) 620 assigned to the contexts for the respective threads”).
It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the invention, to have applied the known technique of one or more indications of a number of SMs to be usable by one or more contexts associated with the one or more software; and cause the one more software threads associated with a respective context to be executed by the subset of SMs determined for the respective context, as taught by Ray, in the same way to the one or more software threads, as taught by Johnson. Both inventions are in the field of GPU scheduling, and combining them would have predictably resulted in “efficient multi-context thread distribution in a processing system,” as indicated by Ray (¶ 1).
Johnson and Ray do not teach, however, Munshi teaches: the one or more indications are provided by the CPU to the GPU (¶ 37, “an application 303 and a platform layer 305 may be running in a host CPU 301” and “Each of the physical compute devices Physical_Compute_Device-1 305 . . . Physical_Compute_Device-N 311 may be one of the CPUs 117 or GPUs 115 of FIG. 1”) via one or more application programming interface (API) calls that create the one or more contexts (¶ 43, “process 400 may be based on a plurality of APIs including cuCreateContext, cuRetainContext and cuReleaseContext. The API cuCreateContext creates a compute context. A compute context may correspond to a compute context object”).
It would have been obvious to a person having ordinary skill in the art, before the effective filing date of the invention, to have applied the known technique of the one or more indications are provided by the CPU to the GPU via one or more application programming interface (API) calls that create the one or more contexts, as taught by Munshi, in the same way to the one or more indications, as taught by Johnson and Ray. Both inventions are in the field of GPU scheduling, and combining them would have predictably resulted in “application interface for data parallel computing across both CPUs (Central Processing Units) and GPUs (Graphical Processing Units),” as indicated by Munshi (¶ 2).
Regarding claim 2, Johnson teaches: The CPU of claim 1, wherein the subset of SMs is to be reserved, at least in part, by a multi-process service (MPS) serving one or more clients for performing the one or more software threads and a server to reserve the subset of the plurality of SMs (¶ 187, “a multi-process service environment 2100 using Volta Multi-Process Service (MPS 2118) is a feature of the Volta GV100 architecture enabling improved performance and isolation for multiple compute applications sharing the GPU”), the one or more clients providing the one or more indications (¶ 188, “olta MPS provides control for MPS clients to specify what fraction of the GPU (GPU fraction 2108, GPU fraction 2110, and GPU fraction 2112) is necessary for execution”), wherein the circuitry is further to cause the GPU to reserve based, at least in part, on the one or more indications, a subset of the SMs within the GPU for the one or more contexts to perform the one or more software threads prior to scheduling performance of the one or more software threads by the GPU (¶ 188, “This control to restrict each client to only a fraction of the GPU execution resources reduces or eliminates head-of-line blocking where work from one MPS client may overwhelm GPU execution resources”).
Regarding claim 3, Johnson teaches: The CPU of claim 1, wherein the one or more indications comprise environment variable data set by one or more software programs for performing the one or more software threads (¶ 188, “Volta MPS provides control for MPS clients to specify what fraction of the GPU (GPU fraction 2108, GPU fraction 2110, and GPU fraction 2112) is necessary for execution”) and provided via the one or more API calls that create the one or more contexts (¶ 202, “An application may generate instructions (e.g., API calls) that cause the driver kernel to generate one or more tasks for execution by the parallel processing unit 2200”).
Regarding claim 4 Johnson teaches: The CPU of claim 1, wherein circuitry is further to combine the one or more contexts with one or more other contexts associated with one or more other software threads (¶ 187, “NVIDIA introduced a software-based multi-process service (MPS) and MPS server that allowed multiple different CPU processes (application contexts) to be combined into a single application context and run on the GPU, attaining higher GPU resource utilization”).
Regarding claim 5, Johnson teaches: The CPU of claim 4, wherein the one or more contexts are combined based, at least in part, on many-to-one context mapping (¶ 187, “NVIDIA introduced a software-based multi-process service (MPS) and MPS server that allowed multiple different CPU processes (application contexts) to be combined into a single application context and run on the GPU, attaining higher GPU resource utilization”).
Regarding claim 6, Johnson teaches: The CPU of claim 1, wherein one or more multi-process service (MPS) clients provide the one or more indications via the one or more API calls that create the one or more contexts for performing the one or more software threads to a MPS server (¶ 187, “Starting with Kepler GK110 GPUs, NVIDIA introduced a software-based multi-process service (MPS) and MPS server that allowed multiple different CPU processes (application contexts) to be combined into a single application context and run on the GPU, attaining higher GPU resource utilization”), the MPS server to manage access to the subset of the plurality of SMs by the one or more software threads (¶ 187, “The MPS 2118 may be used to implement improved thread convergence in accordance with the methods disclosed herein”).
Regarding claim 7, Johnson teaches: The CPU of claim 1, wherein the circuitry is further to execute a daemon process to communicate with one or more clients for executing the one or more software threads (¶ 187, “This process acts as the intermediary to submit work 2120 to the work queues 2122 inside the GPU 2124 for concurrent kernel execution”).
Claims 14 recites commensurate subject matter as claim 1. Therefore, it is rejected for the same reasons.
Regarding claim 19, Johnson teaches: the group of SMs comprises two or more cores and memory (¶ 215, “each of the SM 2500 modules may implement an L1 cache. The L1 cache is private memory that is dedicated to a particular SM 2500”).
Claims 8-13, 15-18, and 20-31 recite commensurate subject matter as claims 1-7, 14, and 19. Therefore, they are rejected for the same reasons.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB D DASCOMB whose telephone number is (571)272-9993. The examiner can normally be reached M-F 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Vital can be reached at (571) 272-4215. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACOB D DASCOMB/Primary Examiner, Art Unit 2198