DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 06/05/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 5, 10-11, 14, and 17-18 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zbiciak US 2020/0285470.
Regarding claim 1, Zbiciak teaches:
1. A processor comprising:
a plurality of cores to process instructions (Fig. 36 processor A and B are a plurality of cores);
a first core of the plurality of cores (Fig. 36 processor A (which may be processor 100 shown in Fig. 1, see [0278]) comprising:
decoder circuitry to decode a single instruction ([0057] decode unit 113 decodes instructions), the single instruction having a first field for an opcode to indicate a read-shared load operation to read data from a memory and a second field to indicate at least one memory address for a location of the data in the memory ([0302]-[0305]: the BLKPLD instruction has an opcode “BLKPLD” to indicate a preload operation to read/preload data from a memory and a source/second field to indicate the starting base address/location of the data in the memory);
execution circuitry to execute the single instruction to read the data from the location in the memory (any functional unit of the processor that performs the BLKPLD instruction is execution circuitry, see [0057] disclosing that the decoding identifies a functional unit for performing the instruction); and
cache controller circuitry (DMA 3660) to initially store the data in a shared state in at least a first cacheline of a cache associated with the first core based on the opcode indicating the read- shared load operation ([0303]-[0305] discloses that the BLKPLD instruction may preload data to L2 cache for read and that the streaming engine may request shared access to the data, which indicates that the data is initially stored in a shared state in a cacheline of L2 associated with processor A based on the BLKPLD opcode indicating the preload/read-shared load operation, see also [0278] and [0281] describing the DMA engine transferring data to L2, which indicates that the DMA engine may be cache controller circuitry).
Claim 10 is directed to a non-transitory machine-readable medium storing instructions to perform the steps of the processor of claim 1 and is rejected for the same reasons as claim 1.
Claim 17 is directed to a method corresponding to the steps performed by the processor of claim 1 and is rejected for the same reasons as claim 1.
Regarding claim 2, Zbiciak teaches:
2. The processor of claim 1 wherein the cache comprises a Level-2 (L2) and/or a Level-1 (L1) cache associated with the first core ([0303]: the cache that data is preloaded into for read may be a L2 cache of processor A, see also Fig. 36).
Claim 11 is directed to a non-transitory machine-readable medium storing instructions to perform the steps of the processor of claim 2 and is rejected for the same reasons as claim 2.
Claim 18 is directed to a method corresponding to the steps performed by the processor of claim 2 and is rejected for the same reasons as claim 2.
Regarding claim 5, Zbiciak teaches:
5. The processor of claim 1 wherein in response to changes to the data by the first core to produce modified data, the shared state of the first cacheline is to be changed to a modified state ([0295] discloses that the L2 cache has a tag field with a clean tag bit that indicates the data elements have not been modified by the processor since they were fetched from L3, which indicates that in response to a change by the processor the shared state of the cacheline in L2 is changed to a not clean/modified state).
Claim 14 is directed to a non-transitory machine-readable medium storing instructions to perform the steps of the processor of claim 5 and is rejected for the same reasons as claim 5.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3-4, 12-13, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Zbiciak US 2020/0285470 in view of Arimilli US 6,018,791.
Regarding claim 3, Zbiciak teaches:
3. The processor of claim 1
Zbiciak does not teach:
wherein, in response to the opcode indicating the read-shared load operation, the cache controller circuitry is to additionally store the data in at least a second cacheline of a Level-3 (L3) cache in a shared state
However, Arimilli teaches a cache inclusion property where if a block is present in the L1 cache of a given processing unit, it is also present in the L2 and L3 caches, see col 2 lines 45-48.
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the caches of Zbiciak to be inclusive as taught by Arimilli such that the preload to L2 in a shared state would additionally store the data in a cacheline of L3 in a shared state. One of ordinary skill in the art would have been motivated to make this modification because storing each line in L1/L2 in L3 would allow for simpler cache coherence since only L3 would need to be checked to know whether a line is present in L1/L2 or not.
Claim 12 is directed to a non-transitory machine-readable medium storing instructions to perform the steps of the processor of claim 3 and is rejected for the same reasons as claim 3.
Claim 19 is directed to a method corresponding to the steps performed by the processor of claim 3 and is rejected for the same reasons as claim 3.
Regarding claim 4, Zbiciak teaches:
4. The processor of claim 1
Zbiciak does not teach:
wherein, in response to a request from a second core of the plurality of cores, the cache controller circuitry is to store the data in at least a third cacheline of an L2 and/or L1 cache of the second core in a shared state
However, Arimilli teaches in response to a request from a second core, to store the data in at least a third cacheline of an L2 and/or L1 cache of the second core in a shared state (col 8 lines 43-52: in response to a request from processor 44b, the line/data in processor 44a is stored in the L2 cache of processor 44b in a recent state, which is a type of shared state for the most recently referenced block, see col 5 line 54-col 6 line 4 and col 6 lines 42-49)
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the cache system of Zbiciak to provide data to the L2 of a second core in a recent/shared state from the L2 of the first core in response to a request from the second core as taught by Arimilli. One of ordinary skill in the art would have been motivated to make this modification to improve latency and increase usable bus bandwidth (Arimilli col 4 lines 63-65).
Claim 13 is directed to a non-transitory machine-readable medium storing instructions to perform the steps of the processor of claim 4 and is rejected for the same reasons as claim 4.
Claim 20 is directed to a method corresponding to the steps performed by the processor of claim 4 and is rejected for the same reasons as claim 4.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Zbiciak US 2020/0285470 in view of Hienecke US 2022/0100502.
Regarding claim 8, Zbiciak teaches:
8. The processor of claim 1
Zbiciak does not teach:
wherein the opcode of the single instruction is to indicate that the load operation comprises a tile load operation, and wherein the data comprises multiple matrix data elements to be stored in a tile register
However, Hienecke teaches a tile load instruction that may load/store a source matrix or matrices (i.e., multiple matrix data elements) into a tile register, see [0180]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the BLKPLD instruction of Zbiciak to comprise a tile load operation for loading multiple matrix data elements into a tile register as taught by Hienecke. One of ordinary skill in the art would have been motivated to make this modification to support loading matrix data elements, which would improve flexibility of the BLKPLD instruction.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Zbiciak US 2020/0285470 in view of Vall US 2014/0195775.
Regarding claim 9, Zbiciak teaches:
9. The processor of claim 1
Zbiciak does not teach:
wherein the opcode of the single instruction is to indicate that the load operation comprises a vector load operation, and wherein the data comprises multiple vector data elements to be stored in a vector or packed data register
However, Vall teaches a vector load instruction that loads data elements from memory into a vector destination register, see [0138]).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the BLKPLD instruction of Zbiciak to comprise a vector load operation for loading multiple vector data elements into a vector register as taught by Vall. One of ordinary skill in the art would have been motivated to make this modification to support loading vector data elements, which would improve flexibility of the BLKPLD instruction.
Allowable Subject Matter
Claims 6-7 and 15-16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter:
The known prior art of record, taken alone or in combination, was not found to teach, in combination with other limitations in the claims, a cache controller circuitry that snoops a first cacheline in a cache associated with a first core and stores a copy of modified data (produced in response to changes to the data by the first core) in a cache associated with the second core in a shared state, as required by claims 6 and 15.
The closest prior art of record was found to be Arimilli US 6,018,791 and Pentkovski US 6,976,131. Arimilli teaches a first processor that stores a line to a second processor in a recent state in response to a request from the second processor, see col 8 lines 43-52. However, Arimilli does not teach the first processor modifying the line and then storing the modified line to the second processor in a shared state. Pentkovski teaches a shared state where the cache line may be stored in multiple caches but is not modified in any of them, see Table 1. Thus, while the prior art was found to generally teach snooping a cacheline in a cache of a first core and storing data in a cache in a shared state, the prior art was not found to teach snooping a first cacheline in a cache associated with a first core and storing a copy of modified data (produced in response to changes to the data by the first core) in a cache associated with the second core in a shared state as required by claims 6 and 15.
Claims 7 and 16 depend from claims 6 and 15, respectively, and thus contain the same allowable subject matter.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US7533242 teaches instructions conveying prefetch hint information which may provide any combination of parameters describing a memory reference traffic pattern to search for, when to begin searching, when to terminate prefetching, and how aggressively to prefetch
US 2008/0065832 teaches DCA logic that may transfer data into a shared cache before, instead of, or in parallel with placing the data into system memory, or by placing the data into system memory or an intermediate cache and using a hint to trigger the placement of the data into the shared cache, where the cache may be shared amongst multiple cores, see [0009]
US 2017/0357585 teaches an L3 cache that may indicate whether data is shared data or unshared when providing data to L2, see [0018]
US 2011/0276786 teaches prefetching data into memory that is shared by threads by distributing prefetch instructions across the threads to reduce execution skew, see Abstract
US 2021/0073129 teaches an instruction that demotes or copies cache lines to a shared cache as indicated by the opcode of the instruction, see [0141]
US 2012/0260056 teaches a "shared data prefetch instruction" that requests for prefetched data to be held in L2 in a shareable state ready to be shared with other processors, see [0136]
US 2005/0027941 teaches a helper core that pushes prefetched data to a main core, see Abstract
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KASIM ALLI whose telephone number is (571)270-1476. The examiner can normally be reached Monday - Friday 9am 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached at (571) 270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KASIM ALLI/Examiner, Art Unit 2183
/ANDREW CALDWELL/Supervisory Patent Examiner, Art Unit 2182