Prosecution Insights
Last updated: October 02, 2026
Application No. 18/758,280

DYNAMICALLY POOLED ALLOCATIONS OF MEMORY BUFFERS ON SPATIAL COMPUTE ARCHITECTURES

Non-Final OA §102§103§DOUBLEPATENT
Filed
Jun 28, 2024
Examiner
MACASIANO, JOANNE GONZALES
Art Unit
2197
Tech Center
2100 — Computer Architecture & Software
Assignee
Amd
OA Round
1 (Non-Final)
67%
Grant Probability
Favorable
1-2
OA Rounds
1y 3m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
217 granted / 323 resolved
+12.2% vs TC avg
Strong +42% interview lift
Without
With
+41.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 6m
Avg Prosecution
18 currently pending
Career history
353
Total Applications
across all art units

Statute-Specific Performance

§101
12.7%
-27.3% vs TC avg
§103
62.3%
+22.3% vs TC avg
§102
15.1%
-24.9% vs TC avg
§112
8.5%
-31.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 323 resolved cases

Office Action

§102 §103 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 2-4, 7-9, 11-13 and 15-20 are objected to because of the following informalities: Claims 2, 8, 11 and 18 recite on Lines 13, 6, 12 and 14: “the first compute tile” which should be “[[the]] a first compute tile”. Claims 7-8 recite on Lines 6 and 9: “the data memory” which should be “the local data memory”. Claim 15 recites on Line 1: “The method of Claim 10” which should likely be “The method of Claim [[10]] 14”, since Claim 15 recites “the first compute tile further comprises” and Claim 14 recites “a first one of the compute tiles comprises”. Claim 16 recites on Lines 5 and 12: “the data memory” which should be “the local data memory”. Claim 17 recites on Line 6: “in memory of a first one of the compute tiles” which should likely be “in memory ”, see Claims 1 and 10 which recite similar language. Claims 3-4, 8-9, 12-13 and 18-20 are also objected to since they depend from at least one of objected Claims 2, 7, 11 and 17-18, and as such inherit the same deficiencies. Appropriate correction is required. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-25 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over Claims 9-10 of copending Application No. 18/758,300 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other because every limitation of instant Claims 1-25 is anticipated by the Claims 9-10 of copending Application No. 18/758,300, in view of the associated specification. See the table below for further details: Current Application 18/758,280 Copending Application 18/758,300 Claim 1 Claims 10 + 9 Claim 2 Claim 9 Claims 3 and 12 Claim 9 + Spec. [0075] “determine the depth based on access/fire patterns of the producer and consumer processes… under the specified execution order.” Claims 4, 13 and 19 Claim 9 + Spec. [0085] “when P3 produces 1 token and C3 consumes 3 tokens, a depth of 4 tokens is needed,” see also [0080]-[0082]. Claims 5-7 Claims 10 + 9 + Spec. [0053] “compute tile 102-1 includes… a data movement accelerator (DMA).” Claim 8 Claim 9 + Spec. [0050] “the DMA may be referred to as a token producer, and the core may be referred to as a token consumer.” Claim 9 Claim 9 + Spec. [0050] “the DMA may be referred to as a token producer, and the core may be referred to as a token consumer.” + Spec. [0056] “configuring DMAs 108 of compute tiles 102, DMAs 116 of shared memory tiles 112” Claim 10 Claims 10 + 9 Claim 11 Claim 9 Claims 14-15 Claims 10 + 9 + Spec. [0053] “compute tile 102-1 includes… a data movement accelerator (DMA).” Claim 16 Claims 10 + 9 + Spec. [0050] “the DMA may be referred to as a token producer, and the core may be referred to as a token consumer” + Spec. [0053] “compute tile 102-1 includes… a data movement accelerator (DMA).” Claim 17 Claims 10 + 9 + Spec. [0076] “generate schedules 214… to enforce the specified execution order.” Claim 18 Claim 9 Claims 20-25 Claims 10+9 This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 10 and 17 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Prabhakar et al. (US PGPUB 2022/0261364; hereinafter “Prabhakar”). Claim 1: (Currently Amended) Prabhakar teaches a non-transitory computer readable medium encoded with a computer program that comprises instructions to cause a processor to ([0270] “One or more implementations of the technology disclosed, or elements thereof can be implemented in the form of a computer product including a non-transitory computer readable storage medium with computer usable program code for performing the method steps indicated.” See also Claim 11 of Prabhakar: “A non-transitory computer readable storage medium impressed with computer program instructions, the instructions, when executed on a processor, implement a method.”): compile an application to execute on a computing platform such that, when the application executes on the computing platform ([0079] “The applications 102 are executed on the reconfigurable processors 152 in a distributed fashion by programming the individual compute and memory components to asynchronously receive, process, and send data and control information.” [0093] “The compile time logic 132 translates the applications 102 developed with commonly used open-source packages such as Keras and PyTorch into reconfigurable processor specifications… The compile time logic 132 loads the configuration files on the reconfigurable processors 152 and causes the configuration files to implement the dataflow graphs 122.”), the application allocates a pooled synchronized buffer in memory ([0073] “the compiler can insert additional synchronization across the read operations of all input buffers, or ‘input siblings,’ of a stage. The additional ‘sibling synchronization’ can be implemented as a control barrier that produces a new control event only when all dependencies of all input buffers are met. The synchronization limits the skew between the start times of input buffers of a stage.” [0088] “Memory allocations represent the creation of logical memory spaces in on-chip and/or off-chip memories for data required to implement the dataflow graphs 122, and these memory allocations are specified in the execution file.” [0103] “The compile time logic 132 is configured to map each of the stage buffers to one or more pattern memory units (PMUs) of the reconfigurable processors 152.” [0104] “Runtime logic 142 parses the execution file and determines configurations of virtual data flow resources required to execute the applications 102. The runtime logic 142 allocates physical configurable units and memory in the pool of reconfigurable data flow resources to the virtual data flow resources.”), a first token producer of the application and a first token consumer of the application synchronously exchange tokens via the pooled synchronized buffer, and a second token producer of the application and a second token consumer of the application synchronously exchange tokens via the pooled synchronized buffer ([0108] “the dataflow graph 400 can comprise a plurality of producers, a plurality of compute nodes, and a plurality of consumers, such that a compute node can receive input from multiple producers and can provide output to multiple consumers.” See also Fig. 8, wherein [0138] “Stage 1.0 has two producers D and E and one consumer F,” i.e. a first producer and consumer, and [0140] “Stage 1.1 has one producer F and one consumer G,” i.e. a second producer and consumer.). Claims 10 and 17: With regard to Claims 10 and 17, these claims are equivalent in scope to Claim 1 rejected above, merely having a different independent claim type, and as such Claims 10 and 17 are rejected under the same grounds and for the same reasons as discussed above with regard to Claim 1. With further regard to Claim 17, the claim recites additional elements not specifically addressed in the rejection of Claim 1. The Prabhakar reference also anticipates these additional elements of Claim 17, for example, Prabhakar teaches: An apparatus, comprising: a processor and memory comprising instructions ([0270] “one or more implementations of the technology disclosed, or elements thereof can be implemented in the form of an apparatus including a memory and at least one processor that is coupled to the memory and operative to perform exemplary method steps”), and the computing platform enforces a specified sequence of token exchanges amongst the first token producer, the first token consumer, the second token producer, and the second token consumer ([0163] “FIG. 16 is a simplified diagram of a tile comprising an array of configurable units with associated instrumentation units… Each of these configurable units contains a configuration store comprising a set of registers or flip-flops storing configuration data that represent either the setup or the sequence to run a program, and can include the number of nested loops, the limits of each loop iterator, the instructions to be executed for each stage, the source of the operands, and the network parameters for the input and output interfaces,” wherein the “array of configurable units” comprise the claimed first and second token producers and consumers.). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 5-9, 14-16 and 20-25 are rejected under 35 U.S.C. 103 as being unpatentable over Prabhakar as applied to Claims 1 and 10 above, and further in view of Ashbaugh et al. (US PGPUB 2025/0291746; hereinafter “Ashbaugh”). Claim 5: (Currently Amended) Prabhakar teaches all the limitations of claim 1 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches: wherein the computing platform comprises multiple compute tiles that include respective compute cores, data movement accelerators (DMAs), and local data memory ([0279] “FIG. 19 illustrates a tile 1900 of a multi-tile processor… the tile 1900 is representative of one of the graphics engine tiles 1710A-1710D of FIG. 17A or compute engine tiles 1740A-1740D of FIG. 17B. The tile 1900 of the multi-tile graphics processor includes an array of graphics core clusters… with each graphics core cluster having an array of graphics cores 515A-515N,” wherein the “graphics cores” are “compute cores”. [0391] “The GPGPU engine 3544 can include GPU tiles 3545… GPU tiles 3545 may also include local volatile memory… the GPU tiles 3545 can include a tensor data movement accelerator circuitry (TDMA circuitry 3547), which can perform tensor data movement operations,” wherein “local volatile memory” is the “local data memory”, see also Fig. 19 showing Tile 1900 comprising “L2 Cache” 1904.), and wherein the computer program further comprises instructions to cause the processor to compile the application such that: a first one of the compute tiles comprises the first token producer ([0357] “The tensor data movement accelerator 2929 enables the offload of tensor data movement for input and output data of the tensor accelerator 2923… A tensor store 3017 can be performed on output generated by the tensor accelerator 2923,” wherein the “output” indicates that the associated “GPU tile,” discussed above, operates as a data producer, i.e. a “first token producer”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tiles as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 6: (Currently Amended) Prabhakar in view of Ashbaugh teaches all the limitations of claim 5 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches wherein the computer program further comprises instructions to cause the processor to compile the application such that: the first compute tile further comprises the first token consumer ([0357] “The tensor data movement accelerator 2929 enables the offload of tensor data movement for input and output data of the tensor accelerator 2923, by facilitating a tensor load 3016 from global memory 3010 to local memory 3020. A tensor load 3016 can be performed on data to be processed by the tensor accelerator 2923.,” wherein the “input data” indicates that the associated “GPU tile” operates as a data consumer, i.e. a “first token producer”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tile consuming data as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 7: (Currently Amended) Prabhakar teaches all the limitations of claim 1 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches: wherein the computing platform comprises multiple compute tiles that include respective compute cores, data movement accelerators (DMA), and local data memory ([0279] “FIG. 19 illustrates a tile 1900 of a multi-tile processor… the tile 1900 is representative of one of the graphics engine tiles 1710A-1710D of FIG. 17A or compute engine tiles 1740A-1740D of FIG. 17B. The tile 1900 of the multi-tile graphics processor includes an array of graphics core clusters… with each graphics core cluster having an array of graphics cores 515A-515N,” wherein the “graphics cores” are “compute cores”. [0391] “The GPGPU engine 3544 can include GPU tiles 3545… GPU tiles 3545 may also include local volatile memory… the GPU tiles 3545 can include a tensor data movement accelerator circuitry (TDMA circuitry 3547), which can perform tensor data movement operations,” wherein “local volatile memory” is the “local data memory, see also Fig. 19 showing Tile 1900 comprising “L2 Cache” 1904.), and wherein the computer program further comprises instructions to cause the processor to compile the application such that: the application allocates the pooled synchronized buffer in the data memory of one of the compute tiles ([0124] “a virtualized graphics execution environment is provided in which the resources of the graphics processing engines… are shared with multiple applications… resources may be subdivided into ‘slices’ which are allocated to different VMs and/or applications based on the processing requirements and priorities.” [0248] “The cache/SLM 1528A-1528F can be configured as cache memory or as a pool of shared memory that is local to each of the respective graphics cores 1521A-1521F.”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tiles as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 8: (Currently Amended) Prabhakar in view of Ashbaugh teaches all the limitations of claim 7 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches wherein the computer program further comprises instructions to cause the processor to compile the application such that: the first token producer corresponds to the DMA of a first one of the compute tiles ([0357] “The tensor data movement accelerator 2929 enables the offload of tensor data movement for input and output data of the tensor accelerator 2923… A tensor store 3017 can be performed on output generated by the tensor accelerator 2923,” wherein the output of data by the “tensor data movement accelerator (TDMA)” indicates that the associated “GPU tile,” discussed above, operates as a data producer, i.e. the “first token producer”.); the first token consumer corresponds to the compute core of one of the first compute tile and a second one of the compute tile ([0357] “The tensor data movement accelerator 2929 enables the offload of tensor data movement for input and output data of the tensor accelerator 2923, by facilitating a tensor load 3016 from global memory 3010 to local memory 3020. A tensor load 3016 can be performed on data to be processed by the tensor accelerator 2923.,” wherein the “input data” indicates that the associated “GPU tile” operates as a data consumer, i.e. “the first token consumer corresponds to the compute core of… the first compute tile”.); and the application allocates the pooled synchronized buffer in the data memory of one of the first and second compute tiles ([0391] “The GPGPU engine 3544 can include GPU tiles 3545… GPU tiles 3545 may also include local volatile memory… the GPU tiles 3545 can include a tensor data movement accelerator circuitry (TDMA circuitry 3547), which can perform tensor data movement operations,” wherein “local volatile memory” comprises the “buffer in the data memory of… the first.. compute tiles”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tiles as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 9: (Currently Amended) Prabhakar in view of Ashbaugh teaches all the limitations of claim 7 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches: wherein the computing platform further comprises a memory tile that comprises memory that is accessible to the multiple compute tiles via the DMAs of the respective compute tiles, wherein the memory tile further comprises a DMA configured to access the local memory of the compute tiles ([0391] “The GPU tiles 3545 may also include local volatile memory or can be coupled with one or more memory tiles. One or more of the GPU tiles 3545 can include a tensor data movement accelerator circuitry (TDMA circuitry 3547), which can perform tensor data movement operations”), and wherein the computer program further comprises instructions to cause the processor to compile the application such that: the first token producer corresponds to the DMA of the shared memory tile ([0357] “The tensor data movement accelerator 2929 enables the offload of tensor data movement for input and output data of the tensor accelerator 2923… A tensor store 3017 can be performed on output generated by the tensor accelerator 2923,” wherein the output of data by the “tensor data movement accelerator (TDMA)” indicates that the associated “memory tile,” discussed above, operates as a data producer, i.e. the “first token producer”.); the first token consumer and the second token producer correspond to the DMA of a first one of the compute tiles ([0357] “The tensor data movement accelerator 2929 enables the offload of tensor data movement for input and output data of the tensor accelerator 2923, by facilitating a tensor load 3016 from global memory 3010 to local memory 3020. A tensor load 3016 can be performed on data to be processed by the tensor accelerator 2923.,” wherein the “input data” indicates that the associated “memory tile” operates as a data consumer, as well as a data producer for providing data to another of the plurality of tiles.); and the second token consumer corresponds to the compute core of one of the first compute tile and a second one of the compute tiles ([0391] “One or more of the GPU tiles 3545 can include a tensor data movement accelerator circuitry (TDMA circuitry 3547), which can perform tensor data movement operations described herein. The TDMA circuitry 3547 is configured to perform complex tensor data transfers.”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the memory tiles as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claims 14-15: With regard to Claims 14-15, these claims are equivalent in scope to Claims 5-6 rejected above, merely having a different independent claim type, and as such Claims 14-15 are rejected under the same grounds and for the same reasons as discussed above with regard to Claims 5-6. Claim 16: With regard to Claim 16, this claim is equivalent in scope to Claims 7 and 8 rejected above, merely having a different independent claim type, and as such Claim 16 is rejected under the same grounds and for the same reasons as discussed above with regard to Claim 7 and 8. Claim 20: (Currently Amended) Prabhakar teaches all the limitations of claim 17 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches: wherein the computing platform comprises multiple compute tiles ([0279] “FIG. 19 illustrates a tile 1900 of a multi-tile processor… the tile 1900 is representative of one of the graphics engine tiles 1710A-1710D of FIG. 17A or compute engine tiles 1740A-1740D of FIG. 17B. The tile 1900 of the multi-tile graphics processor includes an array of graphics core clusters… with each graphics core cluster having an array of graphics cores 515A-515N,” wherein the “graphics cores” are “compute cores”.), and wherein the instructions further cause the processor to compile the application such that: the first token producer corresponds to a compute core of a first one of the first compute tile ([0391] “The GPGPU engine 3544 can include GPU tiles 3545… GPU tiles 3545 may also include local volatile memory… the GPU tiles 3545 can include a tensor data movement accelerator circuitry (TDMA circuitry 3547), which can perform tensor data movement operations.” [0357] “The tensor data movement accelerator 2929 enables the offload of tensor data movement for input and output data of the tensor accelerator 2923… A tensor store 3017 can be performed on output generated by the tensor accelerator 2923,” wherein the “output” indicates that a first associated “GPU tile,” i.e. a first one of the plurality of “GPU tiles 3545”, operates as a data producer, i.e. a “first token producer”.).; and the second token producer corresponds to a compute core of a second one of the compute tiles ([0357] “The tensor data movement accelerator 2929 enables the offload of tensor data movement for input and output data of the tensor accelerator 2923… A tensor store 3017 can be performed on output generated by the tensor accelerator 2923,” wherein the “output” indicates that a second associated “GPU tile,” i.e. a second one of the plurality of “GPU tiles 3545”, also operates as a data producer, i.e. a “second token producer”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tiles as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 21: (New) Prabhakar in view of Ashbaugh teaches all the limitations of claim 5 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches wherein the computer program further comprises instructions to cause the processor to compile the application such that: a second one of the compute tiles comprises the first token consumer ([0255] “Each graphics engine tile 1710A-1710D can be interconnected via a set of tile interconnects 1723A-1723F. Each graphics engine tile 1710A-1710D can also be connected to a memory module or memory devices 1726A-1726D via memory interconnects 1725A-1725D.” [0256] “a cache coherent NUMA (ccNUMA) system is enabled that uses the tile interconnects 1723A-1723F to enable communication between cache controllers within the graphics engine tiles 1710A-1710D.” [0257] “The interconnect fabric 1724 can enable communication between graphics engine tiles 1710A-1710D,” wherein each of the “graphics engine tiles,” also referred to as “GPU tiles” above, are capable of performing data input and output operations and as such a second one of the “graphic engine tiles” is the claimed “second one of the compute tiles”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tile configuration as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 22: (New) Prabhakar in view of Ashbaugh teaches all the limitations of claim 5 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches wherein the computer program further comprises instructions to cause the processor to compile the application such that: the first compute tile further comprises the second token producer ([0255] “Each graphics engine tile 1710A-1710D can be interconnected via a set of tile interconnects 1723A-1723F. Each graphics engine tile 1710A-1710D can also be connected to a memory module or memory devices 1726A-1726D via memory interconnects 1725A-1725D.” [0256] “a cache coherent NUMA (ccNUMA) system is enabled that uses the tile interconnects 1723A-1723F to enable communication between cache controllers within the graphics engine tiles 1710A-1710D.” [0257] “The interconnect fabric 1724 can enable communication between graphics engine tiles 1710A-1710D,” wherein each of the “graphics engine tiles,” also referred to as “GPU tiles” above, are capable of performing data input and output operations and as such a first one of the “graphic engine tiles” is the claimed “first compute tile”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tile configuration as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 23: (New) Prabhakar in view of Ashbaugh teaches all the limitations of claim 5 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches wherein the computer program further comprises instructions to cause the processor to compile the application such that: a second one of the compute tiles comprises the second token producer ([0255] “Each graphics engine tile 1710A-1710D can be interconnected via a set of tile interconnects 1723A-1723F. Each graphics engine tile 1710A-1710D can also be connected to a memory module or memory devices 1726A-1726D via memory interconnects 1725A-1725D.” [0256] “a cache coherent NUMA (ccNUMA) system is enabled that uses the tile interconnects 1723A-1723F to enable communication between cache controllers within the graphics engine tiles 1710A-1710D.” [0257] “The interconnect fabric 1724 can enable communication between graphics engine tiles 1710A-1710D,” wherein each of the “graphics engine tiles,” also referred to as “GPU tiles” above, are capable of performing data input and output operations and as such a second one of the “graphic engine tiles” is the claimed “second one of the compute tiles”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tile configuration as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 24: (New) Prabhakar in view of Ashbaugh teaches all the limitations of claim 5 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches wherein the computer program further comprises instructions to cause the processor to compile the application such that: the first compute tile further comprises the second token consumer ([0255] “Each graphics engine tile 1710A-1710D can be interconnected via a set of tile interconnects 1723A-1723F. Each graphics engine tile 1710A-1710D can also be connected to a memory module or memory devices 1726A-1726D via memory interconnects 1725A-1725D.” [0256] “a cache coherent NUMA (ccNUMA) system is enabled that uses the tile interconnects 1723A-1723F to enable communication between cache controllers within the graphics engine tiles 1710A-1710D.” [0257] “The interconnect fabric 1724 can enable communication between graphics engine tiles 1710A-1710D,” wherein each of the “graphics engine tiles,” also referred to as “GPU tiles” above, are capable of performing data input and output operations and as such a first one of the “graphic engine tiles” is the claimed “first compute tile”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tile configuration as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Claim 25: (New) Prabhakar in view of Ashbaugh teaches all the limitations of claim 5 as described above. Prabhakar does not teach the following, however, Ashbaugh teaches wherein the computer program further comprises instructions to cause the processor to compile the application such that: a second one of the compute tiles comprises the second token consumer ([0255] “Each graphics engine tile 1710A-1710D can be interconnected via a set of tile interconnects 1723A-1723F. Each graphics engine tile 1710A-1710D can also be connected to a memory module or memory devices 1726A-1726D via memory interconnects 1725A-1725D.” [0256] “a cache coherent NUMA (ccNUMA) system is enabled that uses the tile interconnects 1723A-1723F to enable communication between cache controllers within the graphics engine tiles 1710A-1710D.” [0257] “The interconnect fabric 1724 can enable communication between graphics engine tiles 1710A-1710D,” wherein each of the “graphics engine tiles,” also referred to as “GPU tiles” above, are capable of performing data input and output operations and as such a second one of the “graphic engine tiles” is the claimed “second one of the compute tiles”.). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the computer readable medium as disclosed by Prabhakar with the compute tile configuration as taught by Ashbaugh in order to “improve performance or the programmability of accelerator devices when performing tensor-based AI operations” (Ashbaugh [0354]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure is as follows: Weaver et al. (US PGPUB 2022/0237137) discloses a system comprising a pipeline depth determination circuit and a buffer depth determination circuit, wherein the buffer depth determination circuit analyzes input-output connections between processing nodes and assigns a depth value to each of a plurality of buffer memories. Pellauer et al. (“Symphony: Orchestrating Sparse and Dense Tensors with Hierarchical Heterogeneous Processing,” 2023) discusses a hybrid programmable/specialized architecture that focuses on the orchestration of data throughout the memory hierarchy to reduce both the movement of unnecessary data and data movement distances, including discussion regarding the use of Compute Tiles (CTs) and producer-consumer pipelines. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Joanne G. Macasiano whose telephone number is (571)270-7749. The examiner can normally be reached Monday to Thursday, 10:30 AM to 6:00 PM Eastern Standard Time. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at (571) 272-3338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JOANNE G MACASIANO/Examiner, Art Unit 2197
Read full office action

Prosecution Timeline

Jun 28, 2024
Application Filed
Jan 22, 2025
Response after Non-Final Action
Sep 04, 2026
Non-Final Rejection mailed — §102, §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12657119
SELF-CONTAINED MOBILE APPLICATION PROCESSING AND INTEGRATION
3y 9m to grant Granted Jun 16, 2026
Patent 12657076
SYSTEM AND METHOD FOR PROCESSING DATA OF ANY EXTERNAL SERVICES THROUGH API CONTROLLED UNIVERSAL COMPUTING ELEMENTS
2y 2m to grant Granted Jun 16, 2026
Patent 12650682
INDUSTRIAL AUTOMATION PROJECT DESIGN TELEMETRY
4y 8m to grant Granted Jun 09, 2026
Patent 12639193
SYSTEMS AND METHODS FOR RETRIEVAL-AUGMENTED PATCH GENERATION FOR AUTOMATIC PROGRAM REPAIR
3y 9m to grant Granted May 26, 2026
Patent 12613689
ELECTRONIC CONTROL DEVICE, REPROGRAM EXECUTION METHOD, AND NON-TRANSITORY COMPUTER READABLE STORAGE MEDIUM
3y 0m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
67%
Grant Probability
99%
With Interview (+41.8%)
3y 6m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 323 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month