DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 03/31/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Objections
Claims 1 and 10 are objected to because of the following informalities:
Claim 1 line 10- insert –in the SLD circuit—after “entry” to clarify that the particular entry is in the SLD circuit.
Claim 10 line 6- insert –in the SLD circuit—after “entry” to clarify that the particular entry is in the SLD circuit.
Appropriate correction is required.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 10-14 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Meier US 2013/0298127.
Regarding claim 10, Meier teaches:
10. A method, comprising:
decoding, by a processor core, a particular store instruction for storing information in a memory hierarchy ([0040]: since decode unit 32 decodes instructions executed by the load-store unit, this indicates that the decode unit decodes store instructions which are instructions that store information to a memory hierarchy, see also [0006] description of store operation);
based on determining that no entry of a store-load dependency (SLD) circuit has been associated with the particular store instruction, storing, by the processor core in a store PC field of a particular entry, a store PC value corresponding to the particular store instruction (this limitation is not required under BRI as it is contingent on determining that no entry of a SLD circuit has been associated with the particular store instruction, which is not required by the claim);
storing, by the processor core in a load PC field of the particular entry, a default value indicative that a dependent load instruction is currently unidentified (this limitation is not required under BRI since it follows from the contingent limitation (i.e., the particular entry is introduced in the contingent limitation));
based on an initial decode of a subsequent load instruction, determining, by the processor core, whether an entry in the SLD circuit corresponds to the subsequent load instruction (this limitation is not required under BRI since it follows from the contingent limitation (i.e., the SLD circuit is introduced in the contingent limitation); further, this limitation is additionally contingent on an initial decode of a subsequent load instruction, which is not required by the claim); and
based on determining that no entry in the SLD circuit has been associated with the subsequent load instruction, predicting, by the processor core, that the subsequent load instruction is dependent on the particular store instruction (this limitation is not required under BRI as it is contingent on determining that no entry of a SLD circuit has been associated with the subsequent load instruction, which is not required by the claim).
Regarding claim 11, Meier teaches:
11. The method of claim 10, wherein predicting that the subsequent load instruction is dependent on the particular store instruction includes storing, by the processor core in a load PC field of the particular entry, a load PC value corresponding to the subsequent load instruction (this limitation is not required under BRI since it follows from a contingent limitation (i.e., the predicting is only performed contingently)).
Regarding claim 12, Meier teaches:
12. The method of claim 10, further comprising:
performing, by the processor core, the particular store instruction ([0042] and [0047]: the LSU performs store instructions);
subsequent to the performance of the particular store instruction, performing, by the processor core, the subsequent load instruction (this limitation is not required under BRI since the subsequent load instruction follows from a contingent limitation);
determining, by the processor core, that the subsequent load instruction is dependent on the particular store instruction (this limitation is not required under BRI since the subsequent load instruction follows from a contingent limitation); and
indicating, by the processor core, the dependency in a strength field in the particular entry (this limitation is not required under BRI since the particular entry follows from a contingent limitation).
Regarding claim 13, Meier teaches:
13. The method of claim 12, further comprising:
decoding, by the processor core after indicating the dependency, another occurrence of the particular store instruction (this limitation is not required under BRI since indicating the dependency follows from a contingent limitation);
storing, by the processor core, information associated with the particular store instruction in the SLD circuit (this limitation is not required under BRI since the SLD circuit follows from a contingent limitation);
decoding, by the processor core, another occurrence of the subsequent load instruction (this limitation is not required under BRI since the subsequent load instruction follows from a contingent limitation); and
performing, by the processor core, the subsequent load instruction using the stored information (this limitation is not required under BRI since the subsequent load instruction follows from a contingent limitation).
Regarding claim 14, Meier teaches:
14. The method of claim 10, further comprising:
performing, by the processor core, the particular store instruction ([0042] and [0047]: the LSU performs store instructions);
subsequent to the performance of the particular store instruction, performing, by the processor core, the subsequent load instruction (this limitation is not required under BRI since the subsequent load instruction follows from a contingent limitation);
determining, by the processor core, that the subsequent load instruction accesses a different memory location than the particular store instruction (this limitation is not required under BRI since the subsequent load instruction follows from a contingent limitation);; and
adjusting, by the processor core, a value in a misprediction field of the particular entry (this limitation is not required under BRI since the particular entry follows from a contingent limitation).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 7-9, and 15-20 are rejected under 35 U.S.C. 103 as being unpatentable over Meier US 2013/0298127 in view of Asanovic US 2020/0210197.
Regarding claim 1, Meier teaches:
1. An apparatus comprising:
a processor core (Fig. 2 core 30) that includes a store-load dependency (SLD) circuit that includes a plurality of entries, and wherein a given entry includes a plurality of fields including a store program counter (PC) field and a load PC field ([0068] and Fig. 4: LSD predictor table 90 is an SLD circuit that includes entries, the given entry shown in Fig. 4 includes fields including a store PC field and a load PC field), and wherein the processor core is configured to:
after an initial decode of a particular store instruction, determine whether an entry is associated with the particular store instruction in the SLD circuit ([0049] describes dispatching a store that searches the LSD predictor, this search determines whether an entry associated with the store exists in the LSD predictor, and this determination is made after decoding the store as Fig. 2 shows the decode unit before the LSD predictor);
based on a determination that no entry in the SLD circuit has been associated with the particular store instruction, write, to a store PC field of a particular entry, a store PC value corresponding to the particular store instruction ([0064] describes that the LSD predictor is trained by stores and loads that cause ordering violations to prevent such events in the future, thus when an ordering violation does occur (and causes an entry to be created/allocated in the table, see [0068]) this indicates that an entry for the particular store was determined to not exist (otherwise the predictor would have prevented the ordering violation); further [0044] describes allocating an entry for a load store pair when there is an ordering violation and then searching for the store of the pair on a next pass through, which indicates that the store PC is written to the LSD predictor when there is an ordering violation (i.e., based on a determination that no entry has been associated with the store));
in response to an initial decode of a subsequent load instruction, determine whether an entry in the SLD circuit corresponds to the subsequent load instruction ([0072]: the PC of a dispatched load (which may be subsequent to any instruction) is used to search the LSD predictor (in response to the load being initially decoded, see Fig. 2 showing the LSD predictor receiving instructions from the decode unit), which determines whether an entry in the LSD predictor corresponds to the load); and
based on a determination that no entry in the SLD circuit has been associated with the subsequent load instruction, write a load PC value corresponding to the subsequent load instruction to the load PC field of the particular entry ([0044] describes allocating an entry for a load store pair when there is an ordering violation and then searching for the load of the pair when its dispatched again, which indicates that the load PC is written to the LSD predictor when there is an ordering violation (i.e., based on a determination that no entry has been associated with the load)).
Meier does not teach:
write, to a load PC field of the particular entry, a default value indicative that a dependent load instruction is unidentified;
However, Asanovic teaches resetting predictor entries by setting entry values to a default value, see [0021].
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Meier to reset the LSD predictor entries as taught by Asanovic such that the combination would write a default value to the fields of the entries of the LDS predictor, including the load PC field, which would indicate that a dependent load is unidentified as the entry has been reset. One of ordinary skill in the art would have been motivated to make this modification to improve security, see Asanovic [0013].
Regarding claim 2, Meier in view of Asanovic teaches:
2. The apparatus of claim 1, wherein the particular entry further includes a strength field that indicates a strength of dependency between the subsequent load instruction and the particular store instruction (Meier Fig. 4 shows that the entry in the table includes a counter 102 which indicates the strength of the dependency prediction of the subsequent load instruction to the particular store instruction, see Meier [0078]), and wherein the processor core is further configured to:
perform the particular store instruction (Meier [0044] discloses dispatching the store on a next pass through the pipeline of the core, which indicates that the particular store instruction is performed by the pipeline);
subsequent to the performance of the particular store instruction, perform the subsequent load instruction (Meier [0044] discloses dispatching the load from the trained load-store pair (i.e., the subsequent load instruction), which indicates that the subsequent load instruction is performed by the pipeline); and
based on a determination that the subsequent load instruction is dependent on the particular store instruction, modify the strength field (Meier [0057] discloses increasing the strength indicator based on a determination of the data being forwarded to the load from the store queue, which is a determination that the subsequent load is dependent on the store, see also Meier Fig. 7 step 144).
Regarding claim 3, Meier in view of Asanovic teaches:
3. The apparatus of claim 2, wherein to modify the strength field, the processor core is further configured to increment a count value in the strength field (Meier Fig. 7 step 146 discloses incrementing the counter when the load data is in the store queue).
Regarding claim 4, Meier in view of Asanovic teaches:
4. The apparatus of claim 2, wherein the processor core is further configured to:
decode, after the modification of the strength field, another occurrence of the particular store instruction (Meier [0054] describes arming the entry in the LSD predictor when a store finds a valid entry and the strength indicator denotes the dependency prediction is above a threshold, any store that is decoded after the strength indicator is incremented above the threshold and causes the entry to be armed is another occurrence of the store);
based a current value of the strength field, preserve information associated with the particular store instruction in the SLD circuit (Meier [0049] discloses writing the store RNUM to the entry when the entry is armed (which is based on the current value of the strength field being above the threshold));
decode another occurrence of the subsequent load instruction (Meier [0050] discloses that a load may match on an armed entry, any load that is decoded after the entry is armed and matches on the armed entry is another occurrence of the subsequent load); and
based on the current value of the strength field, use the preserved information to perform the subsequent load instruction (Meier [0050] discloses that, based on the load matching on an armed entry (which is based on the current value of the strength indicator being above the threshold) the load may wait for the store RNUM to issue, which uses the preserved information/stored RNUM to perform the load).
Regarding claim 7, Meier in view of Asanovic teaches:
7. The apparatus of claim 1, wherein the particular entry further includes a misprediction field (Meier Fig. 4 counter 102 is a misprediction field since it may be decremented when the load data is not in the store queue (i.e., the dependence prediction is not correct)), and wherein the processor core is further configured to:
perform the particular store instruction (Meier [0044] discloses dispatching the store on a next pass through the pipeline of the core, which indicates that the particular store instruction is performed by the pipeline);
subsequent to the performance of the particular store instruction, perform the subsequent load instruction (Meier [0044] discloses dispatching the load from the trained load-store pair (i.e., the subsequent load instruction), which indicates that the subsequent load instruction is performed by the pipeline);
based on a determination that the subsequent load instruction accesses a different memory location than the particular store instruction, decrement a count in the misprediction field of the particular entry (Meier [0087]: in response to a miss in the store queue (which may be an indication that the load address changed/accesses a different memory location than the store, see [0057]) the counter is decremented at step 148 of Fig. 7).
While Meier in view of Asanovic does not explicitly teach incrementing the misprediction count, Meier further teaches that the counter values may have different representations, see [0079].
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the counters to represent strongly disabled with 11 such that the counter of an entry would be incremented in response to determining that the load accesses a different memory location. This modification would have been obvious to try as there are only a finite number of ways that strong/weakly enabled/disabled can be represented by counter values so that the counter may iterate between strongly disabled and strongly enabled.
Regarding claim 8, Meier in view of Asanovic teaches:
8. The apparatus of claim 7, wherein the processor core is further configured to:
in response to a determination that a different load instruction is associated with the particular store instruction, increment the count in the misprediction field of the particular entry (Meier [0057] and [0087] discloses that a miss in the store queue may be because the load address or store address changed, which is a determination that a different load instruction (different from the load instruction that missed) is associated with the store, and in the combination, the misprediction count would be incremented in response.
Regarding claim 9, Meier in view of Asanovic teaches:
9. The apparatus of claim 1, wherein to write the load PC value to the load PC field of the particular entry, the processor core is further configured to determine that a distance from the particular store instruction to the subsequent load instruction satisfies a distance limit (Meier [0048] describes detecting ordering violations by comparing the older store address against all younger loads in the load queue; the size of the load queue is a distance limit and, by determining a match with a younger load in the load queue, the processor determines that the distance from the older store to the subsequent load satisfies the distance limit to detect the ordering violation (to write the load PC to the load PC field of the LSD predictor, see [0044])).
Regarding claim 15, Meier teaches:
15. A system, comprising:
a memory hierarchy configured to store and access information ([0056]-[0057]: the store queue 54 and cache 48 are a memory hierarchy configured to store and access information): the; and
a processor core, coupled to the memory hierarchy (store queue 54 and cache 48 are coupled to the rest of core 30), wherein the processor core includes an instruction buffer circuit (Fig. 2 reservation station 50 and 52) and a store-load dependency (SLD) circuit (Fig. 2 LSD predictor), and wherein the processor core is configured to:
detect, in the instruction buffer circuit, a particular store instruction for storing information in the memory hierarchy ([0047] discloses dispatching store operations to the reservation stations and issuing operations out of the reservation stations, which involves detecting the store operations in the reservation station);
based on a determination that the SLD circuit does not have an entry associated with the particular store instruction, save, in a store PC field of a particular entry of the SLD circuit, a store PC value corresponding to the particular store instruction ([0064] describes that the LSD predictor is trained by stores and loads that cause ordering violations to prevent such events in the future, thus when an ordering violation does occur (and causes an entry to be created/allocated in the table, see [0068]) this indicates that an entry for the particular store was determined to not exist (otherwise the predictor would have prevented the ordering violation); further [0044] describes allocating an entry for a load store pair when there is an ordering violation and then searching for the store of the pair on a next pass through, which indicates that the store PC is written to the LSD predictor when there is an ordering violation (i.e., based on a determination that no entry has been associated with the store));
detect, in the instruction buffer circuit, a subsequent load instruction for reading information from the memory hierarchy ([0047] discloses dispatching load operations to the reservation stations and issuing operations out of the reservation stations, which involves detecting the load operations in the reservation station, including a load subsequent to the store that may read information from the store queue or cache, see [0057]); and
based on a determination that the SLD circuit does not have an entry associated with the subsequent load instruction, predict that the subsequent load instruction is dependent on the particular store instruction ([0044] describes allocating an entry for a load store pair when there is an ordering violation and then searching for the load of the pair when its dispatched again, which indicates that the load PC is written to the LSD predictor when there is an ordering violation (i.e., based on a determination that no entry has been associated with the load)).
Meier does not teach:
store, in a load PC field of the particular entry, a default value for a load PC value, wherein the default value indicates that a dependent load instruction has not been detected;
However, Asanovic teaches resetting predictor entries by setting entry values to a default value, see [0021].
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify Meier to reset the LSD predictor entries as taught by Asanovic such that the combination would store a default value to the fields of the entries of the LDS predictor, including the load PC field, which would indicate that a dependent load is unidentified as the entry has been reset. One of ordinary skill in the art would have been motivated to make this modification to improve security, see Asanovic [0013].
Regarding claim 16, Meier in view of Asanovic teaches:
16. The system off claim 15, wherein to predict that the subsequent load instruction is dependent on the particular store instruction, the processor core is further configured to write a load PC value corresponding to the subsequent load instruction to the load PC field of the particular entry (Meier [0044] discloses allocating an entry in the LSD predictor for a load-store pair and Fig. 4 shows the entry including a load PC, which indicates that the processor writes the load PC corresponding to the load to the load PC field of an entry corresponding to the pair to predict that the load is dependent on the store of the pair).
Regarding claim 17, Meier in view of Asanovic teaches:
17. The system off claim 15, wherein the processor core is further configured to:
use the memory hierarchy to perform the particular store instruction (Meier [0047]: the store queue stores data corresponding to store operations, i.e., the store queue is used to perform the store); and
subsequent to the performance of the particular store instruction, use the memory hierarchy to perform the subsequent load instruction (Meier [0056]: the load may wait for the store to issue from the reservation station (i.e., subsequent to performance of the store) so that the load data may be forwarded from the store queue, which uses the store queue to perform the subsequent load).
Regarding claim 18, Meier in view of Asanovic teaches:
18. The system off claim 17, wherein the processor core is further configured to:
based on a determination that the particular store instruction and the subsequent load instruction access a same location in the memory hierarchy, modify a strength field of the particular entry, wherein the strength field indicates a dependency between the subsequent load instruction and the particular store instruction (Meier [0057] discloses increasing the strength indicator based on a determination of the data being forwarded to the load from the store queue, which is a determination that the subsequent load and the store access the same location in the store queue, see also Meier Fig. 7 step 144).
Regarding claim 19, Meier in view of Asanovic teaches:
19. The system off claim 17, wherein the processor core is further configured to:
based on a determination that the particular store instruction and the subsequent load instruction access different locations in the memory hierarchy, modify a misprediction field of the particular entry (Meier [0057] discloses decreasing the strength indicator (i.e., a misprediction field) based on a determination of the load data coming from the cache instead of the store queue, which is a determination that the subsequent load and the store access different locations in the memory hierarchy, see also Meier Fig. 7 step 148).
Regarding claim 20, Meier in view of Asanovic teaches:
20. The system off claim 19, wherein the processor core is further configured to:
determine that the misprediction field of the particular entry satisfies a threshold value (Meier [0083] discloses that the counter field of the entry (which is the strength indicator/misprediction field) may be below a threshold (i.e., determined to satisfy a threshold value by being below the threshold value)); and
based on the determination that the misprediction field satisfies the threshold value, invalidate the particular entry (Meier [0083] discloses that if the load matches on an armed entry but the counter is below the threshold, it may not constitute a real match, which invalidates the particular entry by not treating it as a real match).
Allowable Subject Matter
The following is a statement of reasons for the indication of allowable subject matter:
The known prior art of record, taken alone or in combination, was not found to teach, in combination with other limitations in the claims, accessing a particular entry based on the store PC value, selecting a different entry in the SLD circuit based on the load PC field of the particular entry, and preserving information associated with the store in the different entry, as required by claims 5-6.
The closest prior art of record was found to be:
US 2013/0298127 (hereinafter, Meier) which teaches an LSD predictor that stores the store PC, load PC and store RNUM, see Fig. 4, and that the different fields of the predictor table may be separate tables, see [0069]. However, Meier does not teach using the load PC of a particular entry accessed by the store PC to select a different entry in the predictor and preserving information associated with the store in the different entry as required by claim 5.
US 6,622,237 (hereinafter, Keller) which teaches using a load PC to lookup a store PC in a load/store dependency table and using the store PC to lookup an entry in a store PC table, see Fig. 4. However, this is different from using the store PC to lookup a load PC and using the load PC to access a different entry for preserving information associated with the store as required by claim 5.
No other prior art was found to cure this deficiency.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 2019/0310845 teaches a memory dependence predictor table storing a load PC, store PC, store valid, and confidence counter, see Fig. 7
US 2017/0286119 teaches initializing a load address prediction table entry using a predictor tag and actual memory address for the load instruction when the confidence of the entry is zero, see [0052]
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KASIM ALLI whose telephone number is (571)270-1476. The examiner can normally be reached Monday - Friday 9am 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached at (571) 270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KASIM ALLI/Examiner, Art Unit 2183 /JYOTI MEHTA/Supervisory Patent Examiner, Art Unit 2183