DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 21-40 are pending in this office action and presented for examination. Claims 1-20 are newly cancelled, and claims 21-40 are newly added, by the preliminary amendment received July 14, 2025.
Claim 32 recites the limitation “load a data element to the memory” in lines 1-2. Applicant may have intended to recite “load a data element from the memory” in view of, for example, claims 23-24 and 31.
Drawings
The drawings are objected to because:
In FIG. 4D, it is unclear as to how an effective address can comprise a processor, system memory, an accelerator integration slice, and a graphics acceleration (as is indicated via the underlining of reference character 493, which indicates a surface area).
In FIG. 4D, reference characters 492 and 490 appear to be directed to a same surface area.
In FIG. 4E, it is unclear as to how an effective address can comprise a processor, system memory, an accelerator integration slice, and a graphics acceleration (as is indicated via the underlining of reference character 493, which indicates a surface area).
In FIG. 4E, reference characters 492 and 490 appear to be directed to a same surface area.
In FIG. 11, reference characters 1102, 1112, and 1108 appear to be directed to a same surface area.
In FIG. 11, reference characters 1104 and 1106 appear to be directed to a same surface area.
In FIG. 20, reference characters 2000, 2010, 2030, 2040, 2042, 2044, 2046, 2048, and 2050 appear to be directed to a same surface area.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-2, 4-10, 12-20 of U.S. Patent No. 12333310. Although the claims at issue are not identical, they are not patentably distinct from each other because all the limitations of each of the aforementioned instant claims are taught by a corresponding claim of the ‘310 patent. As an exemplary case, see the table below, wherein standard-format limitations in the left column correlate to italicized limitations in the right column; note that the correlation is further explained below the table.
Claim 21 of Instant Application: 19203556
Claim 1 of U.S. Patent No. 12333310
21. (New) A graphics processor comprising:
1. A graphics processor comprising:
a graphics core including execution circuitry and
a graphics core including a plurality of processing resources, each having a plurality of processor lanes to perform a parallel processing operation on a plurality of data elements stored in a memory;
memory access circuitry configured to receive offload of memory address calculations from the execution circuitry,
and memory access circuitry configured to receive offload of memory address calculations for the plurality of data elements from the plurality of processor lanes of the plurality of processing resources,
wherein the memory access circuitry is configured to:
wherein the memory access circuitry is configured to:
determine byte addresses for a plurality of data elements stored in memory,
determine byte addresses for the plurality of data elements stored in the memory,
the byte addresses determined based on a base address and an offset between addresses of data elements of the plurality of data elements,
the byte addresses determined based on a base address, an offset between addresses of data elements of the plurality of data elements, and a scale factor to apply to the offset,
wherein the byte addresses are byte granularity addresses of data elements to be processed by the execution circuitry; and
wherein the byte addresses are byte granularity addresses of data elements to be processed by the plurality of processor lanes; and
submit a memory access request to the memory on behalf of the execution circuitry to access the plurality of data elements at the byte addresses determined for the plurality of data elements.
submit a memory access request to the memory on behalf of the plurality of processor lanes to access the plurality of data elements at the byte addresses determined for the plurality of data elements.
All the limitations of claim 22 are taught by corresponding claim 2 of the ‘310 patent.
All the limitations of claim 23 are taught by corresponding claim 4 of the ‘310 patent.
All the limitations of claim 24 are taught by corresponding claim 5 of the ‘310 patent.
All the limitations of claim 25 are taught by corresponding claim 6 of the ‘310 patent.
All the limitations of claim 26 are taught by corresponding claim 6 of the ‘310 patent.
All the limitations of claim 27 are taught by corresponding claim 7 of the ‘310 patent.
All the limitations of claim 28 are taught by corresponding claim 8 of the ‘310 patent.
All the limitations of claim 29 are taught by corresponding claim 9 of the ‘310 patent.
All the limitations of claim 30 are taught by corresponding claim 10 of the ‘310 patent.
All the limitations of claim 31 are taught by corresponding claim 12 of the ‘310 patent.
All the limitations of claim 32 are taught by corresponding claim 13 of the ‘310 patent.
All the limitations of claim 33 are taught by corresponding claim 14 of the ‘310 patent.
All the limitations of claim 34 are taught by corresponding claim 9 of the ‘310 patent.
All the limitations of claim 35 are taught by corresponding claim 15 of the ‘310 patent.
All the limitations of claim 36 are taught by corresponding claim 16 of the ‘310 patent.
All the limitations of claim 37 are taught by corresponding claim 17 of the ‘310 patent.
All the limitations of claim 38 are taught by corresponding claim 18 of the ‘310 patent.
All the limitations of claim 39 are taught by corresponding claim 19 of the ‘310 patent.
All the limitations of claim 40 are taught by corresponding claim 20 of the ‘310 patent.
Claims 21-40 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 4-9, 12-17, and 19-20 of U.S. Patent No. 12014183. Although the claims at issue are not identical, they are not patentably distinct from each other because all the limitations of each of the aforementioned instant claims are taught by a corresponding claim of the ‘183 patent. As an exemplary case, see the table below, wherein standard-format limitations in the left column correlate to italicized limitations in the right column; note that the correlation is further explained below the table.
Claim 21 of Instant Application: 19203556
Claim 1 of U.S. Patent No. 12014183
21. (New) A graphics processor comprising:
1. A graphics processor comprising:
a graphics core including execution circuitry
a graphics core configured to perform parallel processing operations on a plurality of data elements stored in a memory, the graphics core including processing resources, each of the processing resources having a plurality of processor lanes to perform the parallel processing operations;
and memory access circuitry configured to receive offload of memory address calculations from the execution circuitry,
and memory access circuitry configured to facilitate access to the memory by functional units of the graphics core, wherein the memory access circuitry includes hardware logic to receive offload of memory address calculations for the plurality of data elements from the processing resources
wherein the memory access circuitry is configured to: determine byte addresses for a plurality of data elements stored in memory, the byte addresses determined based on a base address and an offset between addresses of data elements of the plurality of data elements, wherein the byte addresses are byte granularity addresses of data elements to be processed by the execution circuitry; and submit a memory access request to the memory on behalf of the execution circuitry to access the plurality of data elements at the byte addresses determined for the plurality of data elements.
and the memory access circuitry is configured to: receive a message to access data elements of the plurality of data elements stored in the memory, the message to indicate a base address for the plurality of data elements, an offset between addresses of data elements of the plurality of data elements, and a scale factor to apply to the offset; determine a per-lane address for each of the plurality of processor lanes based on the base address, the offset between the addresses of the data elements, and the scale factor to apply to the offset, wherein the per-lane address is an address of a data element to be processed by an associated processor lane of the plurality of processor lanes; and submit a memory access request to the memory to access the data elements at the per-lane addresses determined for the plurality of processor lanes, wherein the offset between the addresses of the data elements is a first offset, the memory access circuitry is additionally configured to add a second offset to the base address, the second offset is a global offset, and the per-lane address for each of the plurality of processor lanes is a byte address.
All the limitations of claim 22 are taught by corresponding claim 1 of the ‘183 patent.
All the limitations of claim 23 are taught by corresponding claim 4 of the ‘183 patent.
All the limitations of claim 24 are taught by corresponding claim 5 of the ‘183 patent.
All the limitations of claim 25 are taught by corresponding claim 6 of the ‘183 patent.
All the limitations of claim 26 are taught by corresponding claim 1 of the ‘183 patent.
All the limitations of claim 27 are taught by corresponding claim 7 of the ‘183 patent.
All the limitations of claim 28 are taught by corresponding claim 8 of the ‘183 patent.
All the limitations of claim 29 are taught by corresponding claim 9 of the ‘183 patent.
All the limitations of claim 30 are taught by corresponding claim 9 of the ‘183 patent.
All the limitations of claim 31 are taught by corresponding claim 12 of the ‘183 patent.
All the limitations of claim 32 are taught by corresponding claim 13 of the ‘183 patent.
All the limitations of claim 33 are taught by corresponding claim 14 of the ‘183 patent.
All the limitations of claim 34 are taught by corresponding claim 9 of the ‘183 patent.
All the limitations of claim 35 are taught by corresponding claim 15 of the ‘183 patent.
All the limitations of claim 36 are taught by corresponding claim 16 of the ‘183 patent.
All the limitations of claim 37 are taught by corresponding claim 17 of the ‘183 patent.
All the limitations of claim 38 are taught by corresponding claim 17 of the ‘183 patent.
All the limitations of claim 39 are taught by corresponding claim 19 of the ‘183 patent.
All the limitations of claim 40 are taught by corresponding claim 20 of the ‘183 patent.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 29-36 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 29 recites the limitation “a plurality of data elements” in line 6. However, it is indefinite as to whether this plurality of data elements is the same as, or different from, “a plurality of data elements” as recited in claim 29, lines 2-3.
Claim 29 recites the limitation “the plurality of data elements” in lines 8-9. However, it is indefinite as to whether the antecedent basis for this limitation “a plurality of data elements” in claim 29, lines 2-3, or “a plurality of data elements” in line 6. Note that this limitation is also recited in claim 29, line 12, and claim 29, line 13.
Claim 29 recites the limitation “the execution circuitry” in lines 11-12. However, it is indefinite as to whether the antecedent basis for this limitation is “execution circuitry” in claim 29, lines 6-7, or “execution circuitry” in claim 29, line 10. Note that this limitation is also recited in claim 34, line 1.
Claims 30-36 are rejected for failing to alleviate the rejections of claim 29 above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 21-40 is/are rejected under 35 U.S.C. 103 as being unpatentable over Minkin et al. (Minkin) (US 20230289292 A1) in view of Cepulis (US 6223271 B1).
Consider claim 21, Minkin discloses a graphics processor ([0049], line 2, GPU) comprising: a graphics core including execution circuitry ([0049], lines 4-9, the plurality of processors comprises multicore processors for example, streaming multiprocessors (SM), 102a . . . 102n (collectively 102). Each SM 102 includes a plurality of processing cores such as functional units 104a . . . 104m (collectively 104)) and memory access circuitry ([0041], line 7, Tensor Memory Access Unit (TMAU)) configured to receive offload of memory address calculations from the execution circuitry ([0042], lines 5-8, the TMAU enables the SM(s) to be more computationally efficient by offloading a significant portion of the related data access operations from the kernels running on the SM(s) to the TMAU; [0044], lines 8-13, the TMAU, in contrast to the LDGSTS instruction executed by the SM, enables the SM to asynchronously transfer a much larger block of data with a single instruction and to also offload the associated address calculations and the like from the threads on the SM to the TMAU; [0059], lines 1-6, each TMAU 112 enables the circuitry of processing cores in the corresponding SM to continue math and other processing of application program kernels while the address calculations and memory access operations are outsourced to closely coupled circuitry dedicated to address calculations and memory accesses. As described below, a TMAU 112, coupled to an SM 102 and having its own hardware circuitry to calculate memory addresses and to read and write shared memory and global memory, enables the coupled SM 102 to improve overall application program kernel performance by outsourcing to the TMAU accesses to any type of data; [0064], lines 12-18, by delegating to hardware in the TMAU the numerous address calculations necessary for obtaining a large amount of data of a data structure or block and the associated coordination of the loads and stores of the respective subblocks of the large amount of data, the SM's power consumption may also be reduced), wherein the memory access circuitry is configured to: determine addresses for a plurality of data elements stored in memory ([0042], lines 5-8, the TMAU enables the SM(s) to be more computationally efficient by offloading a significant portion of the related data access operations from the kernels running on the SM(s) to the TMAU; [0044], lines 8-13, the TMAU, in contrast to the LDGSTS instruction executed by the SM, enables the SM to asynchronously transfer a much larger block of data with a single instruction and to also offload the associated address calculations and the like from the threads on the SM to the TMAU; [0059], lines 1-6, each TMAU 112 enables the circuitry of processing cores in the corresponding SM to continue math and other processing of application program kernels while the address calculations and memory access operations are outsourced to closely coupled circuitry dedicated to address calculations and memory accesses. As described below, a TMAU 112, coupled to an SM 102 and having its own hardware circuitry to calculate memory addresses and to read and write shared memory and global memory, enables the coupled SM 102 to improve overall application program kernel performance by outsourcing to the TMAU accesses to any type of data; [0064], lines 12-18, by delegating to hardware in the TMAU the numerous address calculations necessary for obtaining a large amount of data of a data structure or block and the associated coordination of the loads and stores of the respective subblocks of the large amount of data, the SM's power consumption may also be reduced), the addresses determined based on a base address ([0091], line 7, base global address, for example; also see, for example, [0089]-[0092]; [0132], line 12, base address) and an offset between addresses of data elements of the plurality of data elements ([0090], line 5, stride; [0091], line 4, stride; [0132], line 11, offset), wherein the addresses are addresses of data elements to be processed by the execution circuitry ([0042], lines 16-17, the fetched data can be consumed by multiple threads executing on that SM or on multiple SMs); and submit a memory access request to the memory on behalf of the execution circuitry to access the plurality of data elements at the addresses determined for the plurality of data elements ([0047], lines 6-9, in response to a single request for a data block from the requesting SM, the TMAU is capable of generating multiple requests each for a respective (different) portion of the requested block).
To any extent to which Minkin does not explicitly, implicitly, or inherently disclose that the addresses are byte addresses which are byte granularity addresses, Cepulis explicitly discloses an address that is a byte address that is a byte granularity address (col. 1, lines 14-19, modern microprocessors and/or computers address according to byte granularity. This means memory is generally organized and accessed as a sequence of bytes, and a byte address is used to address memory. Byte addressable memory encompasses what is generally referred to as memory address space). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Cepulis with the invention of Minkin in order to implement the capability of byte addressing in particular. Alternatively, this modification merely entails combining prior art elements (the prior art elements of Minkin as cited above, and the prior art element of byte addresses as explicitly disclosed by Cepulis) according to known methods (Examiner submits that byte addresses are known, as reflected by the disclosure of Cepulis) to yield predictable results (the invention of Minkin, wherein the addresses are byte addresses in particular), which is an example of a rationale that may support a conclusion of obviousness as per MPEP 2143.
Consider claim 22, the overall combination entails the graphics processor of claim 21 (see above), wherein the offset between the addresses of the data elements is a first offset (Minkin, for example, [0090], line 5, stride), the memory access circuitry is additionally configured to add a second offset (for example, [0091], line 4, stride) to the base address (for example, [0090], lines 2-3, base global address), and the second offset is a global offset (for example, [0090], lines 2-4, base global address … is incremented by the tensor stride).
Consider claim 23, the overall combination entails the graphics processor of claim 21 (see above), wherein the memory access request is a request to store a data element to the memory (Minkin, [0057], line 9, TMAU 112 has read/write access; [0081], line 5, stores).
Consider claim 24, the overall combination entails the graphics processor of claim 21 (see above), wherein the memory access request is a request to load a data element from the memory (Minkin, [0057], line 9, TMAU 112 has read/write access; [0081], line 5, loads).
Consider claim 25, the overall combination entails the graphics processor of claim 21 (see above), wherein the memory access request is a request to prefetch a data element to a cache memory (Minkin, [0081], line 5, prefetch; [0144], lines 1-2, the TMAU supports data prefetch requests).
Consider claim 26, the overall combination entails the graphics processor of claim 21 (see above), wherein the execution circuitry includes a plurality of processor lanes and the byte addresses are associated respectively with processor lanes of the plurality of processor lanes (Minkin, [0051], lines 8-9, single-instruction multiple-data (SIMD); [0180], lines 5-11, in an embodiment, the SM 1140 implements a SIMD (Single-Instruction, Multiple-Data) architecture where each thread in a group of threads (e.g., a warp) is configured to process a different set of data based on the same set of instructions. All threads in the group of threads execute the same instructions; [0051], line 12, single-instruction multiple-thread (SIMT); [0180], lines 11-15, in another embodiment, the SM 1140 implements a SIMT (Single-Instruction, Multiple Thread) architecture where each thread in a group of threads is configured to process a different set of data based on the same set of instructions, but where individual threads in the group of threads are allowed to diverge during execution; Cepulis, col. 1, lines 14-19, modern microprocessors and/or computers address according to byte granularity. This means memory is generally organized and accessed as a sequence of bytes, and a byte address is used to address memory. Byte addressable memory encompasses what is generally referred to as memory address space).
Consider claim 27, the overall combination entails the graphics processor of claim 26 (see above), wherein the plurality of processor lanes are associated with a plurality of single instruction multiple data (SIMD) channels (Minkin, [0051], lines 8-9, single-instruction multiple-data (SIMD); [0180], lines 5-11, in an embodiment, the SM 1140 implements a SIMD (Single-Instruction, Multiple-Data) architecture where each thread in a group of threads (e.g., a warp) is configured to process a different set of data based on the same set of instructions. All threads in the group of threads execute the same instructions).
Consider claim 28, the overall combination entails the graphics processor of claim 26 (see above), wherein the plurality of processor lanes are associated with a plurality of single instruction multiple thread (SIMT) threads (Minkin, [0051], line 12, single-instruction multiple-thread (SIMT); [0180], lines 11-15, in another embodiment, the SM 1140 implements a SIMT (Single-Instruction, Multiple Thread) architecture where each thread in a group of threads is configured to process a different set of data based on the same set of instructions, but where individual threads in the group of threads are allowed to diverge during execution).
Consider claim 29, Minkin discloses a method comprising: on a graphics core ([0049], lines 4-9, the plurality of processors comprises multicore processors for example, streaming multiprocessors (SM), 102a . . . 102n (collectively 102). Each SM 102 includes a plurality of processing cores such as functional units 104a . . . 104m (collectively 104)) configured to perform parallel processing operations on a plurality of data elements stored in a memory ([0051], lines 8-9, single-instruction multiple-data (SIMD); [0180], lines 5-11, in an embodiment, the SM 1140 implements a SIMD (Single-Instruction, Multiple-Data) architecture where each thread in a group of threads (e.g., a warp) is configured to process a different set of data based on the same set of instructions. All threads in the group of threads execute the same instructions; [0051], line 12, single-instruction multiple-thread (SIMT); [0180], lines 11-15, in another embodiment, the SM 1140 implements a SIMT (Single-Instruction, Multiple Thread) architecture where each thread in a group of threads is configured to process a different set of data based on the same set of instructions, but where individual threads in the group of threads are allowed to diverge during execution): determining addresses for the plurality of data elements stored in the memory via memory access circuitry configured to receive offload of memory address calculations for a plurality of data elements from execution circuitry of the graphics core ([0042], lines 5-8, the TMAU enables the SM(s) to be more computationally efficient by offloading a significant portion of the related data access operations from the kernels running on the SM(s) to the TMAU; [0044], lines 8-13, the TMAU, in contrast to the LDGSTS instruction executed by the SM, enables the SM to asynchronously transfer a much larger block of data with a single instruction and to also offload the associated address calculations and the like from the threads on the SM to the TMAU; [0059], lines 1-6, each TMAU 112 enables the circuitry of processing cores in the corresponding SM to continue math and other processing of application program kernels while the address calculations and memory access operations are outsourced to closely coupled circuitry dedicated to address calculations and memory accesses. As described below, a TMAU 112, coupled to an SM 102 and having its own hardware circuitry to calculate memory addresses and to read and write shared memory and global memory, enables the coupled SM 102 to improve overall application program kernel performance by outsourcing to the TMAU accesses to any type of data; [0064], lines 12-18, by delegating to hardware in the TMAU the numerous address calculations necessary for obtaining a large amount of data of a data structure or block and the associated coordination of the loads and stores of the respective subblocks of the large amount of data, the SM's power consumption may also be reduced), the addresses determined based on a base address ([0091], line 7, base global address, for example; also see, for example, [0089]-[0092]; [0132], line 12, base address) and an offset between addresses of data elements of the plurality of data elements ([0090], line 5, stride; [0091], line 4, stride; [0132], line 11, offset), wherein the addresses are addresses of data elements to be processed by execution circuitry ([0042], lines 16-17, the fetched data can be consumed by multiple threads executing on that SM or on multiple SMs); and submitting a memory access request to the memory on behalf of the execution circuitry to access the plurality of data elements at the addresses determined for the plurality of data elements ([0047], lines 6-9, in response to a single request for a data block from the requesting SM, the TMAU is capable of generating multiple requests each for a respective (different) portion of the requested block).
To any extent to which Minkin does not explicitly, implicitly, or inherently disclose that the addresses are byte addresses which are byte granularity addresses, Cepulis explicitly discloses an address that is a byte address that is a byte granularity address (col. 1, lines 14-19, modern microprocessors and/or computers address according to byte granularity. This means memory is generally organized and accessed as a sequence of bytes, and a byte address is used to address memory. Byte addressable memory encompasses what is generally referred to as memory address space). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Cepulis with the invention of Minkin in order to implement the capability of byte addressing in particular. Alternatively, this modification merely entails combining prior art elements (the prior art elements of Minkin as cited above, and the prior art element of byte addresses as explicitly disclosed by Cepulis) according to known methods (Examiner submits that byte addresses are known, as reflected by the disclosure of Cepulis) to yield predictable results (the invention of Minkin, wherein the addresses are byte addresses in particular), which is an example of a rationale that may support a conclusion of obviousness as per MPEP 2143.
Consider claim 30, the overall combination entails the method of claim 29 (see above), wherein the offset between the addresses of the data elements is a first offset (Minkin, for example, [0090], line 5, stride), the memory access circuitry is additionally configured to add a second offset (for example, [0091], line 4, stride) to the base address (for example, [0090], lines 2-3, base global address), and the second offset is a global offset (for example, [0090], lines 2-4, base global address … is incremented by the tensor stride).
Consider claim 31, the overall combination entails the method of claim 29 (see above), wherein the memory access request is a request to store a data element to the memory (Minkin, [0057], line 9, TMAU 112 has read/write access; [0081], line 5, stores).
Consider claim 32, the overall combination entails the method of claim 29 (see above), wherein the memory access request is a request to load a data element from the memory (Minkin, [0057], line 9, TMAU 112 has read/write access; [0081], line 5, loads).
Consider claim 33, the overall combination entails the method of claim 29 (see above), wherein the memory access request is a request to prefetch a data element to a cache memory (Minkin, [0081], line 5, prefetch; [0144], lines 1-2, the TMAU supports data prefetch requests).
Consider claim 34, the overall combination entails the method of claim 29 (see above), wherein the execution circuitry includes a plurality of processor lanes and the byte addresses are associated respectively with processor lanes of the plurality of processor lanes (Minkin, [0051], lines 8-9, single-instruction multiple-data (SIMD); [0180], lines 5-11, in an embodiment, the SM 1140 implements a SIMD (Single-Instruction, Multiple-Data) architecture where each thread in a group of threads (e.g., a warp) is configured to process a different set of data based on the same set of instructions. All threads in the group of threads execute the same instructions; [0051], line 12, single-instruction multiple-thread (SIMT); [0180], lines 11-15, in another embodiment, the SM 1140 implements a SIMT (Single-Instruction, Multiple Thread) architecture where each thread in a group of threads is configured to process a different set of data based on the same set of instructions, but where individual threads in the group of threads are allowed to diverge during execution; Cepulis, col. 1, lines 14-19, modern microprocessors and/or computers address according to byte granularity. This means memory is generally organized and accessed as a sequence of bytes, and a byte address is used to address memory. Byte addressable memory encompasses what is generally referred to as memory address space).
Consider claim 35, the overall combination entails the method of claim 34 (see above), wherein the plurality of processor lanes are associated with a plurality of single instruction multiple data (SIMD) channels (Minkin, [0051], lines 8-9, single-instruction multiple-data (SIMD); [0180], lines 5-11, in an embodiment, the SM 1140 implements a SIMD (Single-Instruction, Multiple-Data) architecture where each thread in a group of threads (e.g., a warp) is configured to process a different set of data based on the same set of instructions. All threads in the group of threads execute the same instructions).
Consider claim 36, the overall combination entails the method of claim 34 (see above), wherein the plurality of processor lanes are associated with a plurality of single instruction multiple thread (SIMT) threads (Minkin, [0051], line 12, single-instruction multiple-thread (SIMT); [0180], lines 11-15, in another embodiment, the SM 1140 implements a SIMT (Single-Instruction, Multiple Thread) architecture where each thread in a group of threads is configured to process a different set of data based on the same set of instructions, but where individual threads in the group of threads are allowed to diverge during execution).
Consider claim 37, Minkin discloses a data processing system comprising: a memory device configured to store a plurality of data elements ([0038], lines 7-11, a tensor memory access unit (TMAU) hardware circuitry for moving large data blocks between the shared memory of the parallel processor core and external memory such as, for example, global memory of the parallel processing system); a graphics processor coupled with the memory device ([0049], line 2, GPU), the graphics processor including execution circuitry ([0049], lines 4-9, the plurality of processors comprises multicore processors for example, streaming multiprocessors (SM), 102a . . . 102n (collectively 102). Each SM 102 includes a plurality of processing cores such as functional units 104a . . . 104m (collectively 104)) and memory access circuitry ([0041], line 7, Tensor Memory Access Unit (TMAU)) configured to receive offload of memory address calculations for the plurality of data elements from the execution circuitry ([0042], lines 5-8, the TMAU enables the SM(s) to be more computationally efficient by offloading a significant portion of the related data access operations from the kernels running on the SM(s) to the TMAU; [0044], lines 8-13, the TMAU, in contrast to the LDGSTS instruction executed by the SM, enables the SM to asynchronously transfer a much larger block of data with a single instruction and to also offload the associated address calculations and the like from the threads on the SM to the TMAU; [0059], lines 1-6, each TMAU 112 enables the circuitry of processing cores in the corresponding SM to continue math and other processing of application program kernels while the address calculations and memory access operations are outsourced to closely coupled circuitry dedicated to address calculations and memory accesses. As described below, a TMAU 112, coupled to an SM 102 and having its own hardware circuitry to calculate memory addresses and to read and write shared memory and global memory, enables the coupled SM 102 to improve overall application program kernel performance by outsourcing to the TMAU accesses to any type of data; [0064], lines 12-18, by delegating to hardware in the TMAU the numerous address calculations necessary for obtaining a large amount of data of a data structure or block and the associated coordination of the loads and stores of the respective subblocks of the large amount of data, the SM's power consumption may also be reduced), the memory access circuitry configured to: determine addresses for the plurality of data elements stored in the memory device ([0042], lines 5-8, the TMAU enables the SM(s) to be more computationally efficient by offloading a significant portion of the related data access operations from the kernels running on the SM(s) to the TMAU; [0044], lines 8-13, the TMAU, in contrast to the LDGSTS instruction executed by the SM, enables the SM to asynchronously transfer a much larger block of data with a single instruction and to also offload the associated address calculations and the like from the threads on the SM to the TMAU; [0059], lines 1-6, each TMAU 112 enables the circuitry of processing cores in the corresponding SM to continue math and other processing of application program kernels while the address calculations and memory access operations are outsourced to closely coupled circuitry dedicated to address calculations and memory accesses. As described below, a TMAU 112, coupled to an SM 102 and having its own hardware circuitry to calculate memory addresses and to read and write shared memory and global memory, enables the coupled SM 102 to improve overall application program kernel performance by outsourcing to the TMAU accesses to any type of data; [0064], lines 12-18, by delegating to hardware in the TMAU the numerous address calculations necessary for obtaining a large amount of data of a data structure or block and the associated coordination of the loads and stores of the respective subblocks of the large amount of data, the SM's power consumption may also be reduced), the addresses determined based on a base address ([0091], line 7, base global address, for example; also see, for example, [0089]-[0092]; [0132], line 12, base address) and an offset between addresses of data elements of the plurality of data elements ([0090], line 5, stride; [0091], line 4, stride; [0132], line 11, offset), wherein the addresses are addresses of data elements to be processed by a plurality of processor lanes of the execution circuitry ([0042], lines 16-17, the fetched data can be consumed by multiple threads executing on that SM or on multiple SMs; [0051], lines 8-9, single-instruction multiple-data (SIMD); [0180], lines 5-11, in an embodiment, the SM 1140 implements a SIMD (Single-Instruction, Multiple-Data) architecture where each thread in a group of threads (e.g., a warp) is configured to process a different set of data based on the same set of instructions. All threads in the group of threads execute the same instructions; [0051], line 12, single-instruction multiple-thread (SIMT); [0180], lines 11-15, in another embodiment, the SM 1140 implements a SIMT (Single-Instruction, Multiple Thread); and submit a memory access request to the memory device on behalf of the plurality of processor lanes to access the plurality of data elements at the addresses determined for the plurality of data elements ([0047], lines 6-9, in response to a single request for a data block from the requesting SM, the TMAU is capable of generating multiple requests each for a respective (different) portion of the requested block).
To any extent to which Minkin does not explicitly, implicitly, or inherently disclose that the addresses are byte addresses which are byte granularity addresses, Cepulis explicitly discloses an address that is a byte address that is a byte granularity address (col. 1, lines 14-19, modern microprocessors and/or computers address according to byte granularity. This means memory is generally organized and accessed as a sequence of bytes, and a byte address is used to address memory. Byte addressable memory encompasses what is generally referred to as memory address space). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Cepulis with the invention of Minkin in order to implement the capability of byte addressing in particular. Alternatively, this modification merely entails combining prior art elements (the prior art elements of Minkin as cited above, and the prior art element of byte addresses as explicitly disclosed by Cepulis) according to known methods (Examiner submits that byte addresses are known, as reflected by the disclosure of Cepulis) to yield predictable results (the invention of Minkin, wherein the addresses are byte addresses in particular), which is an example of a rationale that may support a conclusion of obviousness as per MPEP 2143.
Consider claim 38, the overall combination entails the data processing system of claim 37 (see above), wherein the offset between the addresses of the data elements is a first offset (Minkin, for example, [0090], line 5, stride), the memory access circuitry is additionally configured to add a second offset (for example, [0091], line 4, stride) to the base address (for example, [0090], lines 2-3, base global address), and the second offset is a global offset (for example, [0090], lines 2-4, base global address … is incremented by the tensor stride).
Consider claim 39, the overall combination entails the data processing system of claim 37 (see above), wherein the memory access request is a request to load or store a data element or a request to prefetch a data element to a cache memory (Minkin, [0057], line 9, TMAU 112 has read/write access; [0081], line 5, stores; [0081], line 5, loads; [0081], line 5, prefetch; [0144], lines 1-2, the TMAU supports data prefetch requests).
Consider claim 40, the overall combination entails the data processing system of claim 37 (see above), wherein each of the plurality of processor lanes is associated with a single instruction multiple data (SIMD) channel or a single instruction multiple thread (SIMT) thread (Minkin, [0051], lines 8-9, single-instruction multiple-data (SIMD); [0180], lines 5-11, in an embodiment, the SM 1140 implements a SIMD (Single-Instruction, Multiple-Data) architecture where each thread in a group of threads (e.g., a warp) is configured to process a different set of data based on the same set of instructions. All threads in the group of threads execute the same instructions; [0051], line 12, single-instruction multiple-thread (SIMT); [0180], lines 11-15, in another embodiment, the SM 1140 implements a SIMT (Single-Instruction, Multiple Thread) architecture where each thread in a group of threads is configured to process a different set of data based on the same set of instructions, but where individual threads in the group of threads are allowed to diverge during execution).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEITH E VICARY whose telephone number is (571)270-1314. The examiner can normally be reached Monday to Friday, 9:00 AM to 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached at (571)270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KEITH E VICARY/Primary Examiner, Art Unit 2183