Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
1. Claims 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over Ray et al (US 2021/0191868, herein Ray A) in view of Ray et al (US 2021/0303481, herein Ray B).
Regarding claim 11, Ray teaches a method comprising:
receiving a memory access message at a message sequencer included within memory access circuitry of a graphics processor ([0078], GPU memory access control, [0184]);
determining if the memory access message is directed to a set of memory banks that are local to the memory access circuitry ([0053], [0056], graphics core local memory, [0200], local memory banks);
in response to determining that the memory access message is not directed to the set of memory banks that are local to the memory access circuitry, routing the memory access message to a memory access port, the memory access port associated with device memory of the graphics processor ([0038], [0078], external/device memory accessible by GPU, [0101], [0109], port for connecting to system or other memory).
Ray A fails to teach the method comprising in response to determination that the memory access message is directed to the set of memory banks that are local to the memory access circuitry, partitioning the memory access message into multiple message partitions and independently routing the multiple message partitions to the set of memory banks that are local to the memory access circuitry via a crossbar interconnect coupled with the set of memory banks.
Ray B teaches a method for accessing memory of a graphics processor comprising determining that a memory access message is directed to a set of memory banks that are local to memory access circuitry ([0070-0071], local memory partitions, [0413-0416], memory banks), partitioning the memory access message into multiple message partitions ([0069-0073], [0264], memory partitions, [0432], breaking memory access request into multiple commands), and independently routing the multiple message partitions to the set of memory banks that are local to the memory access circuitry via a crossbar interconnect coupled with the set of memory banks ([0432], [0062], [0071], [0369], memory crossbars of graphics processor).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the teachings of Ray A&B to utilize submessages for accessing data within a graphics processing system. While Ray A does not explicitly disclose that the partitioning of local memory into various banks and shared memory partitions may also entail the use of partitioned memory access messages, one of ordinary skill in the art would understand that breaking memory access operations into multiple micro-ops or subinstructions is a routine and conventional aspect of the microprocessor art. As both Ray A&B disclose techniques for handling a shared local memory in a graphics processor, the combination would merely entail a simple substitution of known prior art elements to achieve predictable results, and thus would have been obvious to one of ordinary skill in the art.
Regarding claim 12, the combination of Ray A & Ray B teaches the method of claim 11, wherein the set of memory banks that are local to the memory access circuitry are associated with a cache memory and a shared local memory (SLM) of a graphics core of the graphics processor (Ray A [0053], [0200-0203], shared local memory).
Regarding claim 13, the combination of Ray A & Ray B teaches the method of claim 12, wherein partitioning the memory access message into multiple message partitions includes: partitioning the memory access message into multiple submessages; fragmenting each of the multiple submessages into one or more submessage fragments (Ray B [0264], memory partitions, [0432], breaking memory access request into multiple commands); dividing each submessage fragment into one or more single instruction multiple data (SIMD) slot messages that are sized according to a data element size of a SIMD vector supported by execution resources of the graphics processor (Ray A [0103-0106], SIMD execution with different sizing & Ray B [0088], [0290-0293], SIMD graphics processing); and independently routing the multiple SIMD slot message partitions to the set of memory banks via the crossbar interconnect coupled with the set of memory bank (Ray B [0432], [0062], [0071], [0369], memory crossbars of graphics processor).
Regarding claim 14, the combination of Ray A & Ray B teaches the method of claim 13, wherein independently routing the multiple SIMD slot message partitions to the set of memory banks includes: storing the SIMD slot messages to a first-in-first-out request buffer configured to enable out-of-order access to buffer entries; and configuring the crossbar interconnect to service bank requests out-of-order of receipt (Ray A [0093], [0147], buffers for storing memory request messages & Ray B [0386], [0407], instruction processing out of order).
2. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Edwards (US 2021/0294638) in view of Ray B.
Regarding claim 17, Edwards teaches a memory access circuit within a graphics core of a graphics processor, the memory access circuit comprising:
a first circuitry configured:
store a plurality of request candidates to a first-in-first-out (FIFO) buffer of a plurality of FIFO buffers, the plurality of FIFO buffers configured to enable out-of-order access to buffer entries ([0061], FIFO buffers, [0195], out-of-order execution, [0171], memory requests to memory partitions); and
a second circuitry to:
determine a first crossbar grant configuration based on a first set of pending requests selected from the plurality of FIFO buffers and destination memory banks associated with requests of the first set of pending requests ([0062-0063], memory destinations of copy operations, [0225], pending task pool, [0094], [0165-0166], task scheduling, [0173], [0242], crossbar configuration);
in response to a determination that the first crossbar grant configuration matches each memory bank of the plurality of memory banks with a request in the first set of pending requests, configure a crossbar for memory bank access based on the first crossbar grant configuration ([0171-0172], memory access via crossbar, [0173], [0242], crossbar configuration);
in response to a determination that the first crossbar grant configuration does not match each memory bank of the plurality of memory banks with a pending request in the first set of pending requests:
determine a second crossbar grant configuration via out-of-order access to the buffer entries of the plurality of FIFO buffer ([0170], scheduling tasks on graphics processing array, [0171-0172], partitioned memory access via crossbar, [0173], [0242], crossbar configuration); and
configure the crossbar for memory bank access based on the second crossbar grant configuration ([0173], [0242], crossbar configuration)
Edwards fails to teach wherein the circuitry is to partition a memory access message for a memory bank of a plurality of memory banks into a plurality of request candidates.
Ray B teaches a memory access circuit within a graphics core of a graphics processor configured to partition a memory access message for a memory bank of a plurality of memory banks into a plurality of request candidates ([0432], breaking messages into suboperations, [0062], [0071], [0369], memory crossbars of graphics processor).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the teachings of Edwards and Ray B to utilize submessages for accessing data within a graphics processing system. While Edwards does not explicitly disclose that the partitioning of local memory into various banks and shared memory partitions may also entail the use of partitioned memory access messages, one of ordinary skill in the art would understand that breaking memory access operations into multiple micro-ops or subinstructions is a routine and conventional aspect of the microprocessor art. As both Edwards and Ray B disclose techniques for handling a shared local memory in a graphics processor, the combination would merely entail a simple substitution of known prior art elements to achieve predictable results, and thus would have been obvious to one of ordinary skill in the art.
Allowable Subject Matter
3. Claims 15 and 17-20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Diamant (US 12,645,605) discloses a processor for implementing a clos network to access GPU memory partitions.
Edwards (US 2022/0365882) discloses a processor for implementing out-of-order operations on a GPGPU with a shared local memory.
Andrei (US 2020/0293367) discloses a GPGPU with a partitioned shared local memory.
Pappachan (US 2020/0134208) discloses a GPGPU that handles memory access requests to a partitioned local memory.
Lacy (US 2019/0243645) discloses a GPGPU that implemented SIMD execution with a crossbar to interconnect partitioned memory banks.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL J METZGER whose telephone number is (571)272-3105. The examiner can normally be reached Monday-Friday 8:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached at 571-270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL J METZGER/ Primary Examiner, Art Unit 2183