Prosecution Insights
Last updated: October 02, 2026
Application No. 18/540,663

THROTTLING KERNEL SCHEDULING TO MINIMIZE CACHE CONTENTION

Final Rejection §103
Filed
Dec 14, 2023
Examiner
WU, BENJAMIN C
Art Unit
2195
Tech Center
2100 — Computer Architecture & Software
Assignee
Advanced Micro Devices Inc.
OA Round
2 (Final)
87%
Grant Probability
Favorable
3-4
OA Rounds
1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
472 granted / 540 resolved
+32.4% vs TC avg
Strong +16% interview lift
Without
With
+16.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
21 currently pending
Career history
559
Total Applications
across all art units

Statute-Specific Performance

§101
19.2%
-20.8% vs TC avg
§103
51.4%
+11.4% vs TC avg
§102
0.8%
-39.2% vs TC avg
§112
14.5%
-25.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 540 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . 2. Claims 1–20 are pending for examination in the reply filed on 06/03/2026. Examiner’s Remarks 3. Examiner refers to and explicitly cites particular pages, sections, figures, paragraphs or columns and lines in the references as applied to Applicant’s claims to the extent practicable to streamline prosecution. Although the cited portions of the references are representative of the best teachings in the art and are applied to meet the specific limitations of the claims, other uncited but related teachings of the references may be equally applicable as well. It is respectfully requested that, in preparing responses to the rejections, the Applicant fully considers not only the cited portions of the references, but also the references in their entirety, as potentially teaching, suggesting or rendering obvious all or one or more aspects of the claimed invention. Abbreviations 4. Where appropriate, the following abbreviations will be used when referencing Applicant’s submissions and specific teachings of the reference(s): i. figure / figures: Fig. / Figs. ii. column / columns: Col. / Cols. iii. page / pages: p. / pp. References Cited 5. (A) Ashbaugh et al., US 2020/0293380 A1 (“Ashbaugh”). (B) Alexander et al., US 2011/0161734 A1 (“Alexander”). Ashbaugh and Alexander were cited in the previous Office action. Notice re prior art available under both pre-AIA and AIA 6. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. A. 7. Claims 1, 8 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over (A) Ashbaugh in view of (B) Alexander. See “References Cited” section, above, for full citations of references. 8. Regarding claim 1, (A) Ashbaugh teaches/suggests the invention substantially as claimed, including: “An integrated circuit comprising: a plurality of compute circuits; and” (Fig. 17 and ¶ 231: processing platform incorporated within a system-on-a-chip (SoC) integrated circuit; Figs. 28–29 and ¶ 340: exemplary integrated circuits and associated graphics processors that may be fabricated using one or more IP cores); “a scheduler comprising circuitry configured to: group kernels of a plurality of kernels into a plurality of scheduling groups including at least a first scheduling group that accesses a first data set and a second scheduling group that accesses a second data set” (Fig. 15 and ¶ 222: scheduling of thread groups to processors. Certain sub-sets of thread assignments, such as sub-set 1510, are cached together. In some embodiments, in contrast with the conventional scheduling of thread groups illustrated in FIG. 14, thread groups are assigned utilizing cache locality. For example, in a particular instance thread groups 0 to 3 are assigned in a manner to follow the cache locality established for thread group assignments; ¶ 224: the bias is a hint regarding cache locality that may be utilized in thread group assignment, such as for KERNELS with regular access patterns; ¶ 52: a scheduler 210, which is configured to distribute commands or other work items to a processing cluster array; Fig. 14 and ¶ 219: scheduling of thread groups for graphics processing. In FIG. 14, the illustrated grid 1400 represents the processor operation and caching of data. Each block in the grid 1400 represents the scheduling of a thread group to a processor. However, data caching is based on a set of thread group assignments); “dispatch the first scheduling group for execution by one or more of the plurality of compute circuits” (¶ 223: a certain bias is provided for synchronous scheduling, the bias (using hint or assertion) being that a group of thread groups/warps are to EXECUTE TOGETHER on a single microprocessor process or cache domain to improve cache locality; ¶ 346: the one or more processors are to schedule a plurality of groups of threads for processing by the plurality of graphics processors, the scheduling of the plurality of groups of threads including the plurality of processors to apply a bias for scheduling the plurality of groups of threads according to a cache locality for the one or more caches; ¶ 220: thread groups 0-5 when being scheduling will each be scheduled to the first available processor; ¶ 74: instructions are cached in the instruction cache 252 and dispatched for execution by the instruction unit 254. The instruction unit 254 can dispatch instructions as thread groups (e.g., warps), with each thread of the thread group assigned to a different execution unit within GPGPU core 262; ¶ 86: A scheduler/dispatcher 368 schedules and dispatches the graphics threads for execution on the various cores 370, 371, 372). “based at least in part on whether the first data set is stored in a cache selectively: dispatch the second scheduling group for execution using compute resources of the plurality of compute circuits that are available to execute the second scheduling group” (Fig. 15 and ¶ 222: scheduling of thread groups to processors. Certain sub-sets of thread assignments, such as sub-set 1510, are cached together. In some embodiments, in contrast with the conventional scheduling of thread groups illustrated in FIG. 14, thread groups are assigned utilizing cache locality. For example, in a particular instance thread groups 0 to 3 are assigned in a manner to follow the cache locality established for thread group assignments; ¶ 224: the bias is a hint regarding cache locality that may be utilized in thread group assignment, such as for KERNELS with regular access patterns; ¶ 225: the scheduling is required to be executed on basis of similar cache domain; ¶ 346: one or more caches for storage of data for the plurality of graphics processors, wherein the one or more processors are to schedule a plurality of groups of threads for processing by the plurality of graphics processors, the scheduling of the plurality of groups of threads including the plurality of processors to apply a bias for scheduling the plurality of groups of threads according to a cache locality for the one or more caches; ¶ 282: one or more data caches (e.g., 2212) are included to cache thread data during thread execution). Ashbaugh do not teach “a first scheduling group that accesses a first data set and a second scheduling group that accesses a second data set different from the first data set.” (B) Alexander however teaches or suggests: “a first scheduling group that accesses a first data set and a second scheduling group that accesses a second data set different from the first data set” (¶ 32: to receive one or more kernels from a compiler and schedule the work (e.g., one or more kernels and/ or data sets) for dispatch to/by one or more of the multiple processor cores; ¶ 39: provide work divided into one or more work items 2040-2042, each associated with a kernel ( e.g., a kernel of kernels 2010-2014). The kernels 2010-2014 are forwarded to scheduler 1335. In one or more embodiments, scheduler 1350 includes a scheduler that performs the functions of: (1) scheduling (placing) work elements … (2) selectively allocating the work items to selected processor cores; Fig. 3 and ¶ 43: work items can be grouped with a respective work counter and a respective kernel that can be used to process the work items. As illustrated, work groups 3010-3013 can include respective work items 3040-3043 and respective WIR counters 3050-3053. As shown, work groups 3010 and 3011 can include kernel 2010, and work groups 3012 and 3013 can include kernel 2011; ¶ 20: (1) Work Item: a base element of a data set ( e.g., a byte, a string, an integer number, an floating point number, a pixel, an array, a data structure, etc.)). It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of (B) Alexander with those of (A) Ashbaugh to implement and schedule multiple thread/kernel groups to process different data sets (work items). The motivation or advantage to do so is to provide for balanced and optimal division and distribution of work, i.e. data sets. 9. Regarding claims 8, it is the corresponding method claim reciting similar limitations of commensurate scope as the apparatus (integrated circuit) of claim 1. Therefore, it is rejected on the same basis as claim 1 above. 10. Regarding claims 15 , it is the corresponding system claim reciting similar limitations of commensurate scope as the apparatus (integrated circuit) of claim 1. Therefore, it is rejected on the same basis as claim 1 above, including the following rationale: Ashbaugh teaches or suggests “a cache configured to store a copy of data stored in a memory” (Ashbaugh, Fig. 15 and ¶ 222: scheduling of thread groups to processors. Certain sub-sets of thread assignments, such as sub-set 1510, are cached together. In some embodiments, in contrast with the conventional scheduling of thread groups illustrated in FIG. 14, thread groups are assigned utilizing cache locality. For example, in a particular instance thread groups 0 to 3 are assigned in a manner to follow the cache locality established for thread group assignments; ¶ 224: the bias is a hint regarding cache locality that may be utilized in thread group assignment, such as for KERNELS with regular access patterns; ¶ 225: the scheduling is required to be executed on basis of similar cache domain; ¶ 346: one or more caches for storage of data for the plurality of graphics processors, wherein the one or more processors are to schedule a plurality of groups of threads for processing by the plurality of graphics processors, the scheduling of the plurality of groups of threads including the plurality of processors to apply a bias for scheduling the plurality of groups of threads according to a cache locality for the one or more caches; ¶ 282: one or more data caches (e.g., 2212) are included to cache thread data during thread execution; “a processing circuit comprising: a plurality of chiplets;” (¶ 47: the one or more parallel processor(s) 112, memory hub 105, processor(s) 102, and I/O hub 107 can be integrated into a system on chip (SoC) integrated circuit; Fig. 13 and ¶ 213: FIG. 13 illustrates an exemplary inferencing system on a chip (SOC) 1300 suitable for performing inferencing using a trained model. The SOC 1300 can integrate processing components including a media processor 1302, a vision processor 1304, a GPGPU 1306 and a multi-core processor 1308. The SOC 1300 can additionally include on-chip memory 1305 that can enable a shared on-chip data pool that is accessible by each of the processing components; Fig. 17 and ¶ 231: processing platform incorporated within a system-on-a-chip (SoC) integrated circuit; Fig. 28 and ¶ 341: FIG. 28 is a block diagram illustrating an exemplary system on a chip integrated circuit 2800 that may be fabricated using one or more IP cores, according to an embodiment. Exemplary integrated circuit 2800 includes one or more application processor(s) 2805 (e.g., CPUs), at least one graphics processor 2810, and may additionally include an image processor 2815 and/or a video processor 2820; the Examiner notes that a “chiplet” is a small, modular integrated circuit that can be combined together to create a more complex system-on-chip (SoC) of a single chip). “a scheduler comprising circuitry” (Fig. 15 and ¶ 222: scheduling of thread groups to processors. Certain sub-sets of thread assignments, such as sub-set 1510, are cached together. In some embodiments, in contrast with the conventional scheduling of thread groups illustrated in FIG. 14, thread groups are assigned utilizing cache locality. For example, in a particular instance thread groups 0 to 3 are assigned in a manner to follow the cache locality established for thread group assignments; ¶ 224: the bias is a hint regarding cache locality that may be utilized in thread group assignment, such as for KERNELS with regular access patterns; ¶ 52: a scheduler 210, which is configured to distribute commands or other work items to a processing cluster array … the scheduler 210 is implemented via firmware logic executing on a microcontroller). Allowable Subject Matter 11. Claims 2–7, 9–14, and 16–20 are objected to as being dependent upon a rejected base claim, but would be allowable if 1) rewritten in independent form including all of the limitations of the base claim and any intervening claims. Response to Arguments 12. Applicant’s arguments with respect to the claims have been considered but are moot because the arguments do not apply to any of the newly applied teachings or references being used in the current rejection. In the Remarks, the Applicant also argues that cited references do not disclose or suggest a scheduler that refrains from dispatching a pending second scheduling group when compute resources are available to execute the second scheduling group. The Examiner disagrees. The limitation of “based at least in part on whether the first data set is stored in a cache selectively: dispatch the second scheduling group for execution using compute resources … OR withhold dispatch of the second scheduling group despite compute resources being available to execute the second scheduling group” is recited in the alternative, thus requiring only one of the alternatives to be met by the references’ teachings. As applied in the rejection, Ashbaugh teaches or suggests: “based at least in part on whether the first data set is stored in a cache selectively: dispatch the second scheduling group for execution using compute resources of the plurality of compute circuits that are available to execute the second scheduling group” in at least - Fig. 15 and ¶ 222: scheduling of thread groups to processors. Certain sub-sets of thread assignments, such as sub-set 1510, are cached together. In some embodiments, in contrast with the conventional scheduling of thread groups illustrated in FIG. 14, thread groups are assigned utilizing cache locality. For example, in a particular instance thread groups 0 to 3 are assigned in a manner to follow the cache locality established for thread group assignments; - ¶ 224: the bias is a hint regarding cache locality that may be utilized in thread group assignment, such as for KERNELS with regular access patterns; -- ¶ 225: the scheduling is required to be executed on basis of similar cache domain, wherein Ashbaugh teaches the assignment of thread groups based on cache locality to improve data access. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. (a) Puthoor et al., US 2019/0370059 A1, teaching a multi-kernel wavefront scheduler. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN C WU whose telephone number is (571)270-5906. The examiner can normally be reached Monday through Friday, 8:30 A.M. to 5:00 P.M.. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee J. Li can be reached on (571)272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BENJAMIN C WU/Primary Examiner, Art Unit 2195 August 25, 2026
Read full office action

Prosecution Timeline

Dec 14, 2023
Application Filed
Mar 19, 2026
Non-Final Rejection mailed — §103
Jun 03, 2026
Response Filed
Aug 27, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12717638
LOAD MANAGEMENT SYSTEM FOR DEVICE TO OPTIMIZE USER EXPERIENCE
3y 4m to grant Granted Aug 25, 2026
Patent 12717879
TARGETED CLUSTERING SYSTEM AND METHOD
2y 8m to grant Granted Aug 25, 2026
Patent 12699611
Statistics and Feedback-Based Scan Framework for Cluster Nodes
2y 3m to grant Granted Aug 04, 2026
Patent 12688073
ADAPTABLE RESPONSE TIME PREDICTION FOR STORAGE SYSTEMS UNDER VARIABLE WORKLOADS
3y 10m to grant Granted Jul 21, 2026
Patent 12688067
GRAPHICS PROCESSING UNIT RESOURCE MANAGEMENT METHOD, APPARATUS, AND DEVICE, STORAGE MEDIUM, AND PROGRAM PRODUCT
3y 0m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
87%
Grant Probability
99%
With Interview (+16.4%)
2y 11m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 540 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month