Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 04/20/2026, 06/01/2026 and 08/31/2026. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-36 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claim 1-36 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 10, 19 and 28 of copending Application No. 17/955,094 in view of Munshi “The OpenCL Specification” hereafter Munshi in view of Nvidia “CUDA C++ PROGRAMMING GUIDE” hereafter Nvidia.
This is a provisional nonstatutory double patenting rejection.
17/955,110
17/955,094
Claim 1. One or more processors comprising: circuitry to: respond to an application programming interface (API) call, wherein to respond to the call, the circuitry at least: returns a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads of a software kernel capable of being performed concurrently on an accelerator, wherein at least one group of each block cluster comprises multiple blocks of threads.
Claim 1. One or more processors: circuitry to, in response to a call to an application programming interface (API), return a value indicative of a maximum number of blocks of threads that can be included in a block cluster to perform a software kernel on a graphics processing unit (GPU), wherein the maximum number of blocks is determined based at least on a configuration of the software kernel to be performed. In view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 2. The one or more processors of claim 1, wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator.
Claim 1 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 3. The one or more processors claim 1, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads.
Claim 1 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 4. The one or more processors claim 1, wherein the maximum number is a limit on a number of groups of blocks of threads.
Claim 1 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 5. The one or more processors of claim 1, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads that can be capable of being performed concurrently on the accelerator.
Claim 1 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 6. The one or more processors claim 1, wherein the maximum number is based, at least in part, on an architecture of the accelerator.
Claim 1 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 7 . The one or more processors claim 1, wherein the API is to return the maximum number.
Claim 1 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 8 . The one or more processors of claim 1, wherein the accelerator performs the maximum number of block clusters.
Claim 1 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 9 . The one or more processors of claim 1, wherein the circuitry, to respond to the API call returns a number of partitions of blocks of a grid of blocks of threads.
Claim 1 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 10. A computer-implemented method comprising: receiving an application programming interface (API) call comprising one or more parameters indicating a software kernel to be performed on an accelerator; and in response to receiving the API call, returning a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads capable of being performed in parallel within the indicated kernel, wherein at least one group of each block comprises multiple blocks of threads.
A computer-implemented method comprising: receiving a call to an application programming interface (API), and in response to the call to the API: returning a value indicative of a maximum number of blocks of threads that can be included in a block cluster to perform a software kernel on a graphics processing unit (GPU), wherein the maximum number of blocks is determined based at least on a configuration of the software kernel to be performed. In view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 11 . The computer-implemented method of wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator.
Claim 10 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 12. The computer-implemented method of claim 10, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads.
Claim 10 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 13. The computer-implemented method of claim 10, wherein the maximum number the maximum number is a limit on a number of groups of blocks of threads.
Claim 10 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 14. The computer-implemented method of claim 10, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads capable of being performed concurrently on the accelerator.
Claim 10 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 15. The computer-implemented method of claim 10, wherein the maximum number is based, at least in part, on an architecture of the accelerator.
Claim 10 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 16. The computer-implemented method of claim 10, wherein the API returns the maximum number.
Claim 10 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 17. The computer-implemented method of claim 10,
PNG
media_image1.png
9
4
media_image1.png
Greyscale
further comprising in response to receiving the API call, causing the maximum number to be applied.
Claim 10 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 18 . The computer-implemented method of claim 10, further comprising: in response to receiving the API call, returning a number of partitions of blocks of a grid of blocks.
Claim 10 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 19. A computer system comprising: one or more processors and memory storing executable instructions that, if performed by the one or more processors, are to cause the one or more processors to, respond to an application programming interface (API) call, wherein to respond to the call, the one or more processors at least: return a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads of a software kernel capable of being performed concurrently on an accelerator, wherein at least one group of each block cluster comprises multiple blocks of threads.
Claim 19. A computer system comprising: one or more processors and memory storing executable instructions that, [[if]]when performed by the one or more processors, are to in response to a call to an application programming interface (API), return a value indicative of a maximum number of blocks of threads that can be included in a block cluster to perform a software kernel on a graphics processing unit (GPU), wherein the maximum number of blocks is determined based at least on a configuration of the software kernel to be performed. In view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 20. The computer system of claim 19, wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator.
Claim 19 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 21. The computer system of claim 19, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads.
Claim 19 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 22. The computer system of claim 19, wherein the maximum number is a limit on number of groups of blocks of threads .
Claim 19 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 23 . The computer system of claim 19, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads capable of being performed concurrently on the accelerator.
Claim 19 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 24 . The computer system of claim 19, wherein the maximum number is based, at least in part, on an architecture of the accelerator.
Claim 19 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 25 . The computer system of claim 19, wherein the API returns the maximum number.
Claim 19 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 26 . The computer system of claim 19, wherein the accelerator is to perform the maximum number of block clusters.
Claim 19 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 27 . The computer system of claim 19, wherein to respond to the API call, the one or more processors return a number of partitions of blocks of a grid of blocks of blocks of threads.
Claim 19 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 28 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, are to cause the one or more processors to, respond to an application programming interface (API) call, wherein to respond to the call, the one or more processors at least: return a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads of a software kernel capable of being performed concurrently on an accelerator, wherein at least one group of each block cluster comprises multiple blocks of threads.
Claim 28 . A computer system comprising: one or more processors and memory storing executable instructions that, when performed by the one or more processors, are to in response to a call to an application programming interface (API), return a value indicative of a maximum number of blocks of threads that can be included in a block cluster to perform a software kernel on a graphics processing unit (GPU), wherein the maximum number of blocks is determined based at least on a configuration of the software kernel to be performed. In view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 29 . The non-transitory machine-readable medium of claim 28, wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator.
Claim 28 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 30 . The non-transitory machine-readable medium of claim 28, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads.
Claim 28 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 31 . The non-transitory machine-readable medium of claim 28, wherein the maximum number is a limit on a number of groups of blocks of threads.
Claim 28 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 32 . The non-transitory machine-readable medium of claim 28, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads capable of being performed concurrently on the accelerator.
Claim 28 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 33 . The non-transitory machine-readable medium of claim 28, wherein the maximum number is based, at least in part, on an architecture of an accelerator-a graphics processing unit (GPU).
Claim 28 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 34 The non-transitory machine-readable medium of claim 28, wherein the API call returns the maximum number.
Claim 28 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 35 . The non-transitory machine-readable medium of claim 28, wherein the accelerator is to perform the maximum number of block clusters.
Claim 28 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 36 . non-transitory machine-readable medium of claim 28, wherein the API is to return a number of partitions of blocks of a grid of blocks of threads.
Claim 28 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 1-36 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 9, 18, 27 and 36 of copending Application No. 17/955,123 in view of Munshi “The OpenCL Specification” in view of Nvidia “CUDA C++ PROGRAMMING GUIDE” .
17/955,110
17/955,123
Claim 1. One or more processors comprising: circuitry to: respond to an application programming interface (API) call, wherein to respond to the call, the circuitry at least: returns a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads of a software kernel capable of being performed concurrently on an accelerator, wherein at least one group of each block cluster comprises multiple blocks of threads.
Claim 9. The one or more processors of claim 1, wherein to respond to the call to the API, the circuitry at least: indicates a maximum number of blocks per group of blocks of threads.. In view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 2. The one or more processors of claim 1, wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator.
Claim 9 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 3. The one or more processors claim 1, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads.
Claim 9 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 4. The one or more processors claim 1, wherein the maximum number is a limit on a number of groups of blocks of threads.
Claim 9 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 5. The one or more processors of claim 1, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads that can be capable of being performed concurrently on the accelerator.
Claim 9 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 6. The one or more processors claim 1, wherein the maximum number is based, at least in part, on an architecture of the accelerator.
Claim 9 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 7 . The one or more processors claim 1, wherein the API is to return the maximum number.
Claim 9 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 8 . The one or more processors of claim 1, wherein the accelerator performs the maximum number of block clusters.
Claim 9 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 9 . The one or more processors of claim 1, wherein the circuitry, to respond to the API call returns a number of partitions of blocks of a grid of blocks of threads.
Claim 9 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below
Claim 10. A computer-implemented method comprising: receiving an application programming interface (API) call comprising one or more parameters indicating a software kernel to be performed on an accelerator; and in response to receiving the API call, returning a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads capable of being performed in parallel within the indicated kernel, wherein at least one group of each block comprises multiple blocks of threads.
Claim 18. The computer-implemented method of further comprising: in response responding to the call to the API by at least indicating a maximum number of blocks per group of blocks of threads. In view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 11 . The computer-implemented method of wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator.
Claim 18 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 12. The computer-implemented method of claim 10, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads.
Claim 18 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 13. The computer-implemented method of claim 10, wherein the maximum number the maximum number is a limit on a number of groups of blocks of threads.
Claim 18 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 14. The computer-implemented method of claim 10, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads capable of being performed concurrently on the accelerator.
Claim 18 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 15. The computer-implemented method of claim 10, wherein the maximum number is based, at least in part, on an architecture of the accelerator.
Claim 18 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 16. The computer-implemented method of claim 10, wherein the API returns the maximum number.
Claim 18 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 17. The computer-implemented method of claim 10,
PNG
media_image1.png
9
4
media_image1.png
Greyscale
further comprising in response to receiving the API call, causing the maximum number to be applied.
Claim 18 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 18 . The computer-implemented method of claim 10, further comprising: in response to receiving the API call, returning a number of partitions of blocks of a grid of blocks.
Claim 18 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 19. A computer system comprising: one or more processors and memory storing executable instructions that, if performed by the one or more processors, are to cause the one or more processors to, respond to an application programming interface (API) call, wherein to respond to the call, the one or more processors at least: return a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads of a software kernel capable of being performed concurrently on an accelerator, wherein at least one group of each block cluster comprises multiple blocks of threads.
Claim 27. The computer system of claim 19, wherein to respond to the call to the API, the one or more processors are to at least indicate a maximum number of blocks per group of blocks of threads.. In view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 20. The computer system of claim 19, wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator.
Claim 27 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 21. The computer system of claim 19, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads.
Claim 27 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 22. The computer system of claim 19, wherein the maximum number is a limit on number of groups of blocks of threads .
Claim 27 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 23 . The computer system of claim 19, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads capable of being performed concurrently on the accelerator.
Claim 27 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 24 . The computer system of claim 19, wherein the maximum number is based, at least in part, on an architecture of the accelerator.
Claim 27 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 25 . The computer system of claim 19, wherein the API returns the maximum number.
Claim 27 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 26 . The computer system of claim 19, wherein the accelerator is to perform the maximum number of block clusters.
Claim 27 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 27 . The computer system of claim 19, wherein to respond to the API call, the one or more processors return a number of partitions of blocks of a grid of blocks of blocks of threads.
Claim 27 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 28 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, are to cause the one or more processors to, respond to an application programming interface (API) call, wherein to respond to the call, the one or more processors at least: return a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads of a software kernel capable of being performed concurrently on an accelerator, wherein at least one group of each block cluster comprises multiple blocks of threads.
Claim 36 . The non-transitory machine-readable medium of claim 28, wherein to respond to the call to the API, the one or more processors at least: indicate a maximum number of blocks per group of blocks of threads. In view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 29 . The non-transitory machine-readable medium of claim 28, wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator.
Claim 36 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 30 . The non-transitory machine-readable medium of claim 28, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads.
Claim 36 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 31 . The non-transitory machine-readable medium of claim 28, wherein the maximum number is a limit on a number of groups of blocks of threads.
Claim 36 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 32 . The non-transitory machine-readable medium of claim 28, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads capable of being performed concurrently on the accelerator.
Claim 36 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 33 . The non-transitory machine-readable medium of claim 28, wherein the maximum number is based, at least in part, on an architecture of an accelerator-a graphics processing unit (GPU).
Claim 36 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 34 The non-transitory machine-readable medium of claim 28, wherein the API call returns the maximum number.
Claim 36 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 35 . The non-transitory machine-readable medium of claim 28, wherein the accelerator is to perform the maximum number of block clusters.
Claim 36 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim 36 . non-transitory machine-readable medium of claim 28, wherein the API is to return a number of partitions of blocks of a grid of blocks of threads.
Claim 36 in view of Munshi and Nvidia. it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Munshi and Nvidia to coordinate thread allocation, see art rejection below.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-36 are rejected under 35 U.S.C. 103 as being unpatentable over Munshi “The OpenCL Specification” hereafter Munshi in view of Nvidia “CUDA C++ PROGRAMMING GUIDE” hereafter Nvidia.
Regarding claim 1, Munshi teaches:
One or more processors comprising: circuitry to: respond to an application programming interface (API) call, wherein to respond to the call, the circuitry at least: (Page 165. …returns information about the kernel object that may be specific to a device. kernel specifies the kernel object being queried. E.N.: Munshi expressly teaches an API query concerning a particular kernel/device)
returns a value indicative of a maximum number of block clusters each comprising two or more groups of blocks of threads of a software kernel capable of being performed concurrently on an accelerator, wherein at least one group of each block cluster comprises multiple blocks of threads. (Page 166, CL_KERNEL_WORK_GROUP_S IZE, his provides a mechanism for the application to query the maximum work group size that can be used to execute a kernel on a specific device given by device. The OpenCL implementation uses the resource requirements of the kernel (register usage etc.) to determine what this work-group size should be.)
Munshi does not appear to explicitly teach block clusters each comprising two or more groups of blocks of threads … each block cluster comprises multiple blocks of threads..
However, Nvidia teaches: Page 9, figure 4 Grid of Thread Blocks The number of threads per block and the number of blocks per grid specified in the <<>> syntax can be of type int or dim3. Two-dimensional blocks or grids can be specified as in the example above. In Page 121, Nvidia teaches configuring a kernel according to user input. In page 381 L.2.4.1. Device Properties Unified Memory is supported only on devices with compute capability 3.0 or higher. A program may query whether a GPU device supports managed memory by using cudaGetDeviceProperties() and checking the new managedMemory property.
Accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention, having the teachings of Munshi and Nvidia before them, to include implement Nvidia’s execution structure where a software kernel is executed using a grid comprising of multiple thread blocks, further partitioned into groups of threads and device occupancy based grid configuration technique in Munshi’s API kernel query system such that the system returns a value of the maximum of grids capable of execution on the accelerator.
Regarding claim 2, Munshi teaches:
The one or more processors of claim 1, wherein each block cluster is to perform the software kernel, concurrently, on a plurality of compute units of the accelerator. (Page 172. These work-group instances are executed in parallel across multiple compute units or concurrently on the same compute unit.)
Nvidia also teaches: Page 112. When a CUDA program on the host CPU invokes a kernel grid, the blocks of the grid are enumerated and distributed to multiprocessors with available execution capacity. The threads of a thread block execute concurrently on one multiprocessor, and multiple thread blocks can execute concurrently on one multiprocessor. Refer to claim 1 for the motivation to combine.
Regarding claim 3, Nvidia teaches:
The one or more processors of claim 1, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads. (Page 8, There is a limit to the number of threads per block, since all threads of a block are expected to reside on the same processor core and must share the limited memory resources of that core. On current GPUs, a thread block may contain up to 1024 threads.). Refer to claim 1 for the motivation to combine.
Regarding claim 4, Nvidia teaches:
The one or more processors of claim 1, wherein the maximum number is a limit on a number of groups of blocks of threads. Page 8, Blocks are organized into a one-dimensional, two-dimensional, or three-dimensional grid of thread blocks as illustrated by Figure 4. The number of thread blocks in a grid is usually dictated by the size of the data being processed, which typically exceeds the number of processors in the system. Refer to claim 1 for the motivation to combine.
Regarding claim 5, Nvidia teaches:
The one or more processors of claim 1, wherein the maximum number is a limit on a number of partitions of blocks of threads of a grid of blocks of threads capable of being performed concurrently on the accelerator. (Page 9 . The number of threads per block and the number of blocks per grid specified in the <<>> syntax can be of type int or dim3. Two-dimensional blocks or grids can be specified as in the example above). Refer to claim 1 for the motivation to combine.
Regarding claim 6, Nvidia teaches:
The one or more processors of claim 1, wherein the maximum number is based, at least in part, on an architecture of the accelerator. (page 114. In particular, each multiprocessor has a set of 32-bit registers that are partitioned among the warps, and a parallel data cache or shared memory that is partitioned among the thread blocks. The number of blocks and warps that can reside and be processed together on the multiprocessor for a given kernel depends on the amount of registers and shared memory used by the kernel and the amount of registers and shared memory available on the multiprocessor. There are also a maximum number of resident blocks and a maximum number of resident warps per multiprocessor.). Refer to claim 1 for the motivation to combine.
Regarding claim 7, Munshi teaches:
The one or more processors of claim 1, wherein the value returned by the API is the maximum number. (Page 166, CL_KERNEL_WORK_GROUP_S IZE, his provides a mechanism for the application to query the maximum work group size that can be used to execute a kernel on a specific device given by device. The OpenCL implementation uses the resource requirements of the kernel (register usage etc.) to determine what this work-group size should be.)
Regarding claim 8, Nvidia teaches:
The one or more processors of claim 1, wherein the accelerator performs the maximum number of block clusters. (Page 36, The maximum number of kernel launches that a device can execute concurrently depends on its compute capability and is listed in Table 15.). Refer to claim 1 for the motivation to combine.
Regarding claim 9, Nvidia teaches:
The one or more processors of claim 1, wherein the circuitry, to respond to the API call returns a number of partitions of blocks of a grid of blocks of threads. (Page 142. B.4. Built-in Variables Built-in variables specify the grid and block dimensions and the block and thread indices. They are only valid within functions that are executed on the device. B.4.1. gridDim This variable is of type dim3 (see dim3) and contains the dimensions of the grid.). Refer to claim 1 for the motivation to combine.
Regarding claims 10-18, 19-27 and 28-36 recite commensurate subject matter as claims 1-9. Therefore, they are rejected for the same reasons.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARLOS A ESPANA whose telephone number is (703)756-1069. The examiner can normally be reached Monday - Friday 8 a.m - 5 p.m EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, LEWIS BULLOCK JR can be reached at (571)272-3759. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/C.A.E./Examiner, Art Unit 2199
/LEWIS A BULLOCK JR/Supervisory Patent Examiner, Art Unit 2199