Prosecution Insights
Last updated: August 14, 2026
Application No. 19/024,411

Group Thread Dispatch for Graph Streaming Processor

Non-Final OA §103
Filed
Jan 16, 2025
Priority
Jun 12, 2023 — continuation of 12/236,245
Examiner
PETRANEK, JACOB ANDREW
Art Unit
Tech Center
Assignee
Blaize Inc.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
2y 2m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
620 granted / 776 resolved
+19.9% vs TC avg
Moderate +9% lift
Without
With
+8.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
25 currently pending
Career history
808
Total Applications
across all art units

Statute-Specific Performance

§101
4.2%
-35.8% vs TC avg
§103
57.3%
+17.3% vs TC avg
§102
16.2%
-23.8% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 776 resolved cases

Office Action

§103
DETAILED ACTION Claims 1-20 are pending. The office acknowledges the following papers: ADS filed on 2/14/2025 and 2/12/2025. Allowable Subject Matter Claims 2-5, 10, 13-16, and 18-20 would be allowable if rewritten to overcome the double patenting rejections, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims. As allowable subject matter has been indicated, applicant's reply must either comply with all formal requirements or specifically traverse each requirement not complied with. See 37 CFR 1.111(b) and MPEP § 707.07(a). Priority The effective filing date for the subject matter defined in the pending claims in this application is 6/12/2023. Drawings The Examiner contends that the drawings submitted on 1/16/2025 are acceptable for examination proceedings. Specification The disclosure is objected to because of the following informalities: The lengthy specification has not been checked to the extent necessary to determine the presence of all possible minor errors. The Applicant’s cooperation is requested in correcting any errors of which the Applicant may become aware. Appropriate correction is required. Claim Objections Claim 18 is objected to because of the following informalities: Claim 18 recites “The graph streaming processor,” at line 1 that should be changed to “The graph streaming processor of claim 17,” for proper antecedent basis and to provide consistent claim language throughout the claims. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the "right to exclude" granted by a patent and to prevent possible harassment by multiple assignees. See In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970);and, In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent is shown to be commonly owned with this application. See 37 CFR 1.130(b). Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b). Applicants can file an eTerminal Disclaimer (eTD) in utility applications filed under 35 U.S.C. 111(a) or in compliance with 35 U.S.C. 371, and design applications. Filing an eTD via EFS-Web is highly recommended due to an extensive backlog for processing paper TDs. However, applicants may still file a TD for manual review. Claims 1-20 are rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over claims 1-15 of U.S. Patent No. 12,236,245. Although the conflicting claims are not identical, they are not patentably distinct from each other because U.S. Patent No. 12,236,245 contains every element of claims 1-20 of the instant application and thus anticipates the claims of the instant application. Claims of the instant application therefore are not patently distinct from earlier patent claims and as such are unpatentable over obvious-type double patenting. A later application claim is not patently distinct from an earlier claim if the later claim is anticipated by the earlier claim. Instant Application Patent 12,236,245 1. A method of group thread dispatch for a graph streaming processor, comprising 1. A method of group thread dispatch for a graph streaming processor, comprising receiving, by a thread scheduler of the graph streaming processor, a group of threads, wherein the group of threads comprises a plurality of threads which operate on an input tensor, wherein each of the plurality of threads operates on inputs of the input tensor and a subset of a weight tensor to generate a subset of an output tensor; receiving, by a thread scheduler of the graph streaming processor, a group of threads, wherein the group of threads comprises a plurality of threads which operate on an input tensor, wherein each of the plurality of threads operates on the inputs of the input tensor and a subset of a weight tensor to generate a subset of an output tensor; calculating by the thread scheduler, a resource requirement for execution of the group of threads; calculating by the thread scheduler, a resource requirement for execution of the group of threads; calculating, by the thread scheduler, resource availability in a plurality of processors of each of a plurality of processor arrays; and calculating, by the thread scheduler, resource availability in a plurality of processors of each of a plurality of processor arrays; dispatching the group of threads to a selected one of the plurality of processors of the plurality of processor arrays that has a resource availability that meets or exceeds the resource requirement for execution of the group of threads. dispatching the group of threads to a selected one of the plurality of processors of the plurality of processor arrays that has a resource availability that meets or exceeds the resource requirement for execution of the group of threads; and scheduling a group load instruction for all threads of the group of threads, comprising: loading into a group load register a subset of inputs of the input tensor for processing of each thread of the group of threads, wherein the group load register provides the subset of the inputs of the input tensor to the group of threads of the selected one of the plurality of processors; wherein all threads of the group of threads are synchronized when executing the group load instruction; wherein all threads of the group of threads are processed independently on the selected one of the plurality of processors when not executing the group load instruction; wherein the processing of each thread of the group of threads comprises generating a subset of outputs of the output tensor for each thread of the plurality of threads based on the subset of weights of the weight tensor and the inputs of the input tensor. Claims 17 is similar to claim 1 and is rejected for the same reasons. Dependent claims 2-16 and 18-20 are read upon by the claims 1-15 of U.S. Patent No. 12,236,245. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 6-9, 11-12, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Beckman et al. (U.S. 2022/0206876), in view of Official Notice. As per claim 1: Beckman disclosed a method of group thread dispatch for a graph streaming processor, comprising receiving, by a thread scheduler of the graph streaming processor, a group of threads, wherein the group of threads comprises a plurality of threads which operate on an input tensor, wherein each of the plurality of threads operates on inputs of the input tensor and a subset of a weight tensor to generate a subset of an output tensor (Beckman: Figures 2-3 elements 250 and 305, paragraphs 2, 37, and 39)(The dispatch unit (i.e. thread scheduler) receives wavefronts (i.e. thread groups) to schedule for execution. Official notice is given that GPUs can execute matrix convolutions using input features/activations/tensor matrices (i.e. input tensor) and weight matrices (i.e. weight tensor) for the advantage of increased parallel performance. Thus, it would have been obvious to one of ordinary skill in the art that the specific workloads performed by GPU wavefronts includes matrix convolutions on input and weight tensors.); calculating by the thread scheduler, a resource requirement for execution of the group of threads (Beckman: Figure 5 element 510, paragraphs 39-40 and 55)(The dispatch unit determines how many registers (i.e. resource requirement) are needed for executing the wavefront.); calculating, by the thread scheduler, resource availability in a plurality of processors of each of a plurality of processor arrays (Beckman: Figures 2-3 and 5 elements 255A-N, 350A-N, and 515, paragraphs 36, 39-40, and 56)(The dispatch unit determines if there are enough available registers (i.e. resource availability) to meet the register requirement of the wavefront to be dispatched. Subsets of compute units/SIMD units read upon the processor arrays.); and dispatching the group of threads to a selected one of the plurality of processors of the plurality of processor arrays that has a resource availability that meets or exceeds the resource requirement for execution of the group of threads (Beckman: Figures 2-3 elements 250 and 305, paragraphs 2, 36-37, 39-40, and 55-56)(The dispatch unit (i.e. thread scheduler) receives wavefronts (i.e. thread groups) to schedule for execution. Wavefronts are dispatched to subsets of compute units/SIMD units with sufficient available registers for execution.). As per claim 6: Beckman disclosed the method of claim 1, wherein each thread of the group of threads is scheduled on one processor of the selected one of the plurality of processors (Beckman: Figures 2-3 elements 250 and 305, paragraphs 2, 36-37, 39-40, and 55-56)(The dispatch unit (i.e. thread scheduler) receives wavefronts (i.e. thread groups) to schedule for execution. Wavefronts are dispatched to subsets of compute units/SIMD units with sufficient available registers for execution.). As per claim 7: Beckman disclosed the method of claim 1, wherein calculating resource requirement for execution of the group of threads comprises determining the number of threads in the group of threads and determining the number of registers required by each thread of the group of threads (Beckman: Figure 5 element 510, paragraphs 39-40 and 55)(The dispatch unit determines how many registers (i.e. resource requirement) are needed for executing the wavefront. The width of the wavefront corresponds to the compute unit/SIMD unit width.). As per claim 8: Beckman disclosed the method of claim 1, wherein calculating resource availability in the plurality of processors comprises determining the available thread slots and determining the number of available registers in each processor of the plurality of processors of each of the plurality of processor arrays (Beckman: Figures 3 and 5 elements 330A-N and 510, paragraphs 39-46 and 55)(The dispatch unit determines how many registers (i.e. resource requirement) are needed for executing the wavefront. The width of the wavefront corresponds to the compute unit/SIMD unit width. Determining available threads can be done by determining unused registers corresponding to compute units/SIMD units (i.e. compute unit A hasn’t been allocated a wavefront using registers, it has SIMD lane X threads available). The stack frame table tracks all allocated register ranges from compute units.). As per claim 9: Beckman disclosed the method of claim 8, wherein determining the available thread slots and determining the number of available registers in each processor of the plurality of processors of each of the plurality of processor arrays comprises: decrementing, by a counter of the thread scheduler corresponding with the processor, the number of available registers when starting a thread (Beckman: Figure 3 elements 315-317, paragraph 46)(Official notice is given that resource availability can be tracked by incrementing/decrementing counters for the advantage of determining the exact number of resources available at a time. Thus, it would have been obvious to one of ordinary skill in the art to decrement a register counter several times equal to the number of registers allocated for a wavefront.); and incrementing, by the counter of the thread scheduler corresponding with the processor, the number of available when a thread is completed (Beckman: Figure 3 elements 315-317, paragraph 46)(In view of the above official notice, it would have been obvious to one of ordinary skill in the art to increment a register counter several times equal to the number of registers retired for a wavefront.). As per claim 11: Beckman disclosed The method of claim 1, wherein each thread includes an instance of a set of instructions running on a processor of the graph streaming processor (Beckman: Figure 2 element 205, paragraphs 2 and 36)(The wavefronts are executed on GPUs (i.e. graph streaming processor).). As per claim 12: Beckman disclosed the method of claim 1, wherein dispatching the group of threads to the plurality of processors of a one of the plurality of processor arrays comprising loading attributes of each thread of the group of threads to each of the available thread slots of a processor, wherein the attributes include a program pointer (Beckman: Figures 2-3 elements 250 and 305, paragraphs 2, 36-37, 39-40, and 55-56)(Wavefronts are dispatched to subsets of compute units/SIMD units with sufficient available registers for execution. Official notice is given that wavefronts use program counters for the advantage of executing a large set of parallel program code. Thus, it would have been obvious to one of ordinary skill in the art to implement sending a program counter to the allocated compute units/SIMD units.), pointer to the input tensor, pointer to the weight tensor, and pointer to the output tensor (Beckman: Figures 2-3 elements 250 and 305, paragraphs 2, 36-37, 39-40, and 55-56)(Wavefronts are dispatched to subsets of compute units/SIMD units with sufficient available registers for execution. Official notice is given that matrix convolutions can be performed on GPUs using pointers to input features, input weights, and output results in registers/memory for the advantage of reading the required inputs for convolution operations and storing the result. Thus, it would have been obvious to one of ordinary skill in the art to implement matrix convolutions in Beckman that use input and output pointers.). As per claim 17: Claim 17 essentially recites the same limitations of claim 1. Therefore, claim 17 is rejected for the same reasons as claim 1. Conclusion The following is text cited from 37 CFR 1.111(c): In amending in reply to a rejection of claims in an application or patent under reexamination, the applicant or patent owner must clearly point out the patentable novelty which he or she thinks the claims present in view of the state of the art disclosed by the references cited or the objections made. The applicant or patent owner must also show how the amendments avoid such references or objections. The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Krashinsky (U.S. 2013/0042090), taught a WARP scheduler. Webber (U.S. 2017/0192779), taught a thread instruction scheduler. Chen et al. (U.S. 2019/0171448), taught a matrix stream processor. Koneru et al. (U.S. 2019/0235917), taught a graph streaming processor. Fleischer et al. (U.S. 2021/0049230), taught execution of matrix convolutions. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB A. PETRANEK whose telephone number is (571)272-5988. The examiner can normally be reached on M-F 8:00-4:30. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached on (571) 270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JACOB PETRANEK/Primary Examiner, Art Unit 2183
Read full office action

Prosecution Timeline

Jan 16, 2025
Application Filed
Aug 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699565
Processor Employing Instruction That Performs A Bitwise Majority Vote Operation
2y 1m to grant Granted Aug 04, 2026
Patent 12675314
STREAMING ENGINE WITH SHORT CUT START INSTRUCTIONS
2y 1m to grant Granted Jul 07, 2026
Patent 12675367
CYCLE ACCURATE TRACING OF VECTOR INSTRUCTIONS
2y 0m to grant Granted Jul 07, 2026
Patent 12657156
MULTIPLE MULTITHREADED PROCESSORS WITH SHARED DATA CACHE
1y 10m to grant Granted Jun 16, 2026
Patent 12651306
LOW LATENCY STREAMING REMAPPING ENGINE
1y 9m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
88%
With Interview (+8.6%)
3y 9m (~2y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 776 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month