Prosecution Insights
Last updated: October 02, 2026
Application No. 18/394,797

CONTROLLER FOR AN ARRAY OF DATA PROCESSING ENGINES

Final Rejection §103§DOUBLEPATENT
Filed
Dec 22, 2023
Examiner
TRUONG, LECHI
Art Unit
2194
Tech Center
2100 — Computer Architecture & Software
Assignee
Advanced Micro Devices Inc.
OA Round
2 (Final)
87%
Grant Probability
Favorable
3-4
OA Rounds
2m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
776 granted / 889 resolved
+32.3% vs TC avg
Strong +36% interview lift
Without
With
+36.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
25 currently pending
Career history
921
Total Applications
across all art units

Statute-Specific Performance

§101
18.1%
-21.9% vs TC avg
§103
63.8%
+23.8% vs TC avg
§102
4.1%
-35.9% vs TC avg
§112
8.1%
-31.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 889 resolved cases

Office Action

§103 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-9, 11-21 are presented for the examination. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the "right to exclude" granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applyingonline/ terminal-disclaimer. Claims 1, 5, 9, and 18 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of copending Application No. 18/394,859 (see claims filed 3/4/2026) in view of "CHARM: A Composable Heterogeneous Accelerator-Rich Microprocessor" (2012-Cong). Claims 11 and 13-17 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of copending Application No. 18/394,859 (see claims filed 3/4/2026). With respect to claim 1, 18/394,859 teaches A system on a chip (SoC), comprising ([claim 1 line 1]): at least one central processing unit (CPU) ([claim 1 line 2]); an accelerator comprising an array of data processing engines (DPEs) ([claim 1 lines 3-4]); and an interface communicatively coupling the CPU to the controller and the accelerator. Claim 1 of 18/394,859 does not teach receive a task from the CPU; and control data movement into and out of the array of DPEs in the accelerator to perform the task; and inform the CPU when the task is complete. However, 2012-Cong teaches a controller comprising circuitry configured to (in FIG. 2, this is the tile labeled "ABC", [page 380], which stands for accelerator block composer, see [Abstract] lines 4-5; FIG. 3, shows what the ABC is configured to perform, [page 318]): receive a task from the CPU (in FIG. 3(A), the arrow from the "core" to the "ABC", [page 381]; see also the caption "A core sends a request for an LCA to the ABC;", see also "The core sends a data flow graph (DFG) of the desired LCA to the ABC (Figure 3A), [page 382 col 1 paragraph 1 lines 5-6], where the graph is a graph of tasks); and control data movement into and out of the array of DPEs in the accelerator to perform the task (shown in FIG. 3(B), and FIG. 3(C), where the ABC allocates ABBs, [page 381]; the actual algorithm is discussed in the 3.2.1 ABC design section, "The ABC uses a two-tiered allocation policy to decide which ABBs to compose into a given LCA. First, the ABC will attempt to balance the concentration of memory-accessing ABBs across the entire system. The purpose of this is to limit contention in the DMA associated with each node. Second, the ABC will employ a simple greedy approach to select ABBs that are local to other ABBs they communicate with. This is done in order to minimize the cost of communication between ABBs.", [page 381 col 1 paragraph 5 lines 7-15]); and inform the CPU when the task is complete (shown in FIG.3(D), "The ABC signals completion to the core.", [page 381]). It would have been obvious to one skilled in the art before the effective filing date to combine 18/394,859 with 2012-Cong because a teaching, suggestion, or motivation in the prior art would have led one skilled in the art to combine prior art teaching to arrive at the claimed invention. Claim 1 along with claim 7 of the '859 application discloses a system that teaches all of the claimed features except for how the controller is configured. 2012-Cong teaches: Running medical imaging benchmarks, our experimental results show an average speed up of 2.1X (best case 3.7X) compared to approaches that use LCAs together with a hardware resource manager. We also gain in terms of energy consumption (average 2.4X; best case 4.7X). (2012-Cong [Abstract] lines 14-18]). A person having skill in the art would have a reasonable expectation of successfully speeding up the system in the system of 18/394,859 by modifying the '859 Application with the steps performed by the controller shown in FIG. 3, (2012-Cong [page 381]). Therefore, it would have been obvious to combine 18/394,859 with 2012-Cong to a person having ordinary skill in the art. With respect to claim 5, '859 Application in view 2012-Cong teaches all of the limitations of claim 1, as noted above. Claim 5 of the '859 Application further teaches: wherein the interface is a second NoC, wherein the second NoC is larger than the NoC in the Al accelerator ([see claim 5]). With respect to claim 9, '859 Application in view 2012-Cong teaches all of the limitations of claim 1, as noted above. Claim 8 of the '859 Application further teaches: wherein each of the DPEs comprises a core, a memory module, and an interconnect, wherein the interconnects in the DPEs are interconnected so that the DPEs are able to transmit data between each other ([see claim 8]). With respect to claim 11, 18/394,859 teaches A method, comprising ([claim 10 line 1]) : receiving, from a CPU, an instruction at a controller to perform a hardware acceleration task using an accelerator, wherein the CPU, the controller, and accelerator are disposed on a same integrated circuit (IC) ([claim 10 lines 2-4]; controlling, using the controller, data movement into and out of an array of DPEs in the accelerator to perform the hardware acceleration task ([claim 10 lines 5-10]); and informing the CPU that the hardware acceleration task is complete using the controller ([claim 10 lines 11-12]). With respect to claim 13, '859 Application teaches all of the limitations of claim 11, as noted above. Claim 10 of the '859 Application further teaches: transmitting data generated by the DPEs when performing the hardware acceleration task to a NoC in the accelerator ([claim 10 In 7-8]); performing, at an IOMMU in the accelerator, an address translation on the data received from the NoC ([claim 10 In 9-10]); and transmitting the address translated data to the CPU ([claim 10 In 11-12]). With respect to claim 14, '859 Application teaches all of the limitations of claim 13, as noted above. Claim 11 of the '859 Application further teaches: performing the address translation comprises: translating virtual addresses used by the accelerator to physical addresses used to store the address translated data ([claim 11]). With respect to claim 15, '859 Application teaches all of the limitations of claim 14, as noted above. Claim 12 of the '859 Application further teaches: wherein the virtual addresses are memory mapped virtual addresses, wherein the memory mapped virtual addresses are used to transmit the data from the DPEs, through the NoC, and to the IOMMU. ([claim 12]). With respect to claim 16, '859 Application teaches all of the limitations of claim 13, as noted above. Claim 14 of the '859 Application further teaches: wherein the controller communicates with the CPU only through a second NoC, wherein the second NoC is larger than the NoC in the accelerator, wherein the second NoC also communicatively couples the CPU to the accelerator ([claim 14]). With respect to claim 17, '859 Application teaches all of the limitations of claim 16, as noted above. Claim 17 of the '859 Application further teaches: wherein the controller communicates with the CPU only through a second NoC, wherein the second NoC is larger than the NoC in the accelerator, wherein the second NoC also communicatively couples the CPU to the accelerator ([claim 17]). With respect to claim 18, '859 Application teaches A system, comprising ([claim 18 line 1]: an IC, comprising ([claim 18 line 2]): at least one CPU ([claim 18 line 3]), an accelerator comprising DPEs ([claim 18 lines 4-5]) a memory controller ([claim 18 line 10]), and an interface communicatively coupling the CPU to the accelerator, the controller, and the memory controller ([claim 18 In 11-12]); and at least one memory coupled to the memory controller in the IC ([claim 18 In 13]). Claim 18 of the '859 Application does not teach a controller configured to: receive a task from the CPU; control data movement into and out of the DPEs in the accelerator to perform the task; and inform the CPU when the task is complete. However, 2012-Cong teaches a controller comprising circuitry configured to (in FIG. 2, this is the tile labeled "ABC", [page 380], which stands for accelerator block composer, see [Abstract] lines 4-5; FIG. 3, shows what the ABC is configured to perform, [page 318]): receive a task from the CPU (in FIG. 3(A), the arrow from the "core" to the "ABC", [page 381]; see also the caption "A core sends a request for an LCA to the ABC;", see also "The core sends a data flow graph (DFG) of the desired LCA to the ABC (Figure 3A), [page 382 col 1 paragraph 1 lines 5-6], where the graph is a graph of tasks); and control data movement into and out of the array of DPEs in the accelerator to perform the task (shown in FIG. 3(B), and FIG. 3(C), where the ABC allocates ABBs, [page 381]; the actual algorithm is discussed 3.2.1 ABC design section, "The ABC uses a two-tiered allocation policy to decide which ABBs to compose into a given LCA. First, the ABC will attempt to balance the concentration of memory-accessing ABBs across the entire system. The purpose of this is to limit contention in the DMA associated with each node. Second, the ABC will employ a simple greedy approach to select ABBs that are local to other ABBs they communicate with. This is done in order to minimize the cost of communication between ABBs.", [page 381 col 1 paragraph 5 lines 7-15]); and inform the CPU when the task is complete (shown in FIG. 3(D), "The ABC signals completion to the core.", [page 381]). It would have been obvious to one skilled in the art before the effective filing date to combine 18/394,859 with 2012-Cong because a teaching, suggestion, or motivation in the prior art would have led one skilled in the art to combine prior art teaching to arrive at the claimed invention. Claim 1 along with claim 7 of the '859 application discloses a system that teaches all of the claimed features except for how the controller is configured. 2012-Cong teaches: Running medical imaging benchmarks, our experimental results show an average speed up of 2.1X (best case 3.7X) compared to approaches that use LCAs together with a hardware resource manager. We also gain in terms of energy consumption (average 2.4X; best case 4.7X). (2012-Cong [Abstract] lines 14-18]). A person having skill in the art would have a reasonable expectation of successfully speeding up the system in the system of 18/394,859 by modifying the '859 Application with the steps performed by the controller shown in FIG. 3, (2012-Cong [page 381]). Therefore, it would have been obvious to combine 18/394,859 with 2012-Cong to a person having ordinary skill in the art. This is a provisional nonstatutory double patenting rejection. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 11, 18, 21 are rejected under 35 U.S.C. 103 as being unpatentable over Lichtenau ( US 20230305818 A1 ) in view of Wang( US 20200050457 A1) and further in view of Raha( US 20210271960 A1). As to claim 1, Lintenau teaches A system on a chip (SoC), comprising: at least one central processing unit (CPU); an artificial intelligence (A)accelerator ( FIG. 1 shows an embodiment of the disclosure, which is a block diagram representing a processor chip 100 including some components. The term “processor chip” as used herein refers to any semiconductor chip or die having one or more processors. The processor chip 100 includes a plurality of CPUs 101, 102, 103 (indicated respectively as CPU.sub.0, CPU.sub.1 . . . CPUn, where n is any positive integer), a Level 3 (L3) memory 105 (i.e., Level 3 cache memory), and an AI accelerator 110, para[0028], ln 1-15); comprising an array of data processing engines (DPEs)( The AI accelerator 110 can implement an AI primitive on a hardware systolic array architecture, para[0029], ln 11-14/ the AI accelerator 110 can contain more or less than eight (8) engines, para[0039], ln 1-3/ The “pipeline” includes a plurality of engines/stages through which init flit of the temporal accelerator code 108 is moved through stages 0-3 in which any variables are replaced with actual values, within the AI accelerator 110 itself. In the disclosed example, there are four (4) stages (stage 0-3) of the pipeline taking place in the four (4) variable replacement engines 118A-D, para[0037], ln 6-14/ Fig.2 a controller comprising circuitry configured to: receive a task from the CPU to execute one or more layers in an Al model, and an interface communicatively coupling the CPU to the controller and the Al accelerator (the firmware 104 (e.g., on one or more CPUs, like those in FIG. 1) can reside in a restricted area of memory. The firmware 104 can push a bootstrap program 106, e.g., embedded in a large command, to the AI accelerator 110…….The firmware 104 pushes the variable lookup table 107 into the AI accelerator 1010 through the DMA interface 111. ……..A template accelerator code (or program) 108 and tensor data 109 from a user, both shown in the L3 memory 105, can be fetched by DMA interface 111 of the AI accelerator 110 through a gateway 112 to the cache system. The template accelerator code 108 and the tensor data 109 are moved, e.g. through engine 0 115A to be stored in a static random access memory (SRAM)-based scratchpad 117, para[0033], ln 1-28 to para[0034]/ ) control data movement into and out of the array of DPEs in the accelerator to perform the task, ( in order to start to perform, or trigger, variable replacement of the template accelerator code 108, the AI accelerator 110 starts initialization by reading initialization data flits (“init flits” or “init data flits”) of the template accelerator code 108 from the scratchpad 117. One init flit can be one 128 byte data line of the template accelerator code 108. Controls can be generated by the FSM 113, which can interpret the type of the init flit….. ach init data flit can enter into a first of a plurality of variable replacement engines/stages that make up the variable replacement hardware (such as variable replacement engines 118A-D) and can then enter into subsequent engines/stages. The disclosed process of variable replacement is described as being operable in a “pipeline” because each init flit enters the first stage, stage 0 118A, or variable replacement engine 0, where it may possibly undergo variable replacement and then moves to the stage 1 118B or variable replacement engine 1 to possibly undergo variable replacement there, and so on. Each init flit continues to move through each of the variable replacement engines 0-3 118A-118D (or stages 0-3) until the variables are replaced with actual values. Each time an init flit moves to the next stage, or variable replacement engine, it leaves an opening for another init flit to enter the variable replacement engine. This process continues until all of the init flit of the template accelerator code 108 are processed and have had variables replaced with actual values, as necessary., para[0035], ln 1-10 / para[0036]. Wang teaches a controller comprising circuitry configured to: receive a task from the CPU to execute one or more layers in an Al model; inform the CPU when the task is complete( various special-purpose processors for neural network are developed, such as GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), or ASIC (Application Specific Integrated Circuits), para[0004], ln 6-11/ the special-purpose execution component executing the operation instruction can be determined as a convolution engine (an integrated circuit with which the convolution operation is integrated), para[0049], ln 6-49/ The CPU 11 may interact with the artificial intelligence chip 12 via the bus 13 to send and receive messages. The CPU 11 can send descriptive information of a neural network model to the artificial intelligence chip 12, and receive a processing result returned by the artificial intelligence chip 12. The artificial intelligence chip 12, also referred to as an AI accelerator card or computing card, is specially used for processing a large amount of compute-intensive computational tasks in artificial intelligence applications. The artificial intelligence chip 12 may include at least one general-purpose execution component and at least one special-purpose execution component. The general-purpose execution component is communicatively connected to the special-purpose execution components respectively. The general-purpose execution component can receive and analyze the descriptive information of the neural network model sent by the CPU 11, and then send analyzed operation instruction to a specific special-purpose execution component (e.g., a special-purpose execution component locked by the general-purpose execution component). The special-purpose execution component can execute the operation instruction sent by the general-purpose execution component, para[0038] to para[0039]/ The artificial intelligence chip may be communicatively connected to the CPU. An executing body (e.g., the general-purpose execution component of the artificial intelligence chip 12 in FIG. 1) of the method for executing an instruction for an artificial intelligence chip can receive descriptive information for describing the neural network model sent by the CPU. The descriptive information may include at least one operation instruction. Here, the operation instruction may be an instruction that can be executed by the special-purpose execution component of the artificial intelligence chip, e.g., a matrix calculation instruction, or a vector operation instruction, para[0044], ln 3-15/ The special-purpose execution component is configured to: execute, in response to receiving the operation instruction sent by the general-purpose execution component, the received operation instruction; and return the notification for instructing the operation instruction being completely executed after the received operation instruction is completely executed, para[0021], ln 19-26/ end a notification for instructing the at least one operation instruction being completely executed to the CPU after determining the at least one operation instruction being completely executed, para[0106]). It would have been obvious to one of the ordinary skill in the art before the effective filling date of claimed invention was made to modify the above teaching to incorporate the above feature because this avoids frequent interaction between the artificial intelligence chip and the CPU, and improving the performance of the artificial intelligence chip. Raha teaches array of data processing engines( As examples, these processor(s) or accelerators 124 may be a cluster of artificial intelligence (AI) GPUs, para[0099], ln 12-18/ FIG. 18 and/or accelerator(s) 124 of FIG. 1. As illustrated by both FIGS. 2a and 2b, the architecture 200 includes a spatial array 210 of processing elements (PEs) 230, para[0027], ln 6-14). It would have been obvious to one of the ordinary skill in the art before the effective filling date of claimed invention was made to modify the above teaching to incorporate the above feature because this improving resource efficiency in HW accelerators. As to claim 11, it is rejected for the same reason as to claim 1 above. In additional, Lintenau teaches IC( para[0028], ln 1-15), Raha teaches a plurality of ICs (e.g., such as system-on-Chips (SoCs)), para[0034] ,ln 6-9) for the same reason as to claim 1 above. As to claim 18, it is rejected for the same reason as to claim 1 above. In additional, Lintenau teaches and the memory controller; and at least one memory coupled to the memory controller in the IC( The processor chip 100 shown includes a plurality of CPUs 101, 102, 103 (indicated respectively as CPU.sub.0, CPU.sub.1 . . . CPUn, where n is any positive integer), a Level 3 (L3) memory 105 (i.e., Level 3 cache memory), and an AI accelerator 110. It should be noted that the CPU can alternatively be any suitable processor. The AI accelerator 110 can be connected to the L3 memory 105 via a direct memory access (DMA) interface 111. “Cache memory” is a faster and smaller segment of memory with access time that can be slower than registers but faster than a main memory. “L3 cache memory” is a third level of cache memory that is present outside a CPU and shared by all cores of the CPU. Other possible components of the processor chip 100 are contemplated but are not shown or discussed, para[0028], ln 5-19). As to claim 21, Lintenau teaches the task comprises executing a plurality of layers in the Al model, and wherein each layer corresponds to moving data into and out of the DPEs;wherein, for each layer of the plurality of layers except a final layer, the controller is configured to:control data movement into and out of the array of DPEs in the accelerator to perform the layer, determine that the layer is complete( Each init data flit can enter into a first of a plurality of variable replacement engines/stages that make up the variable replacement hardware, based on determining that the layer is complete, control data movement into and out of the array of DPEs to perform a next layer of the plurality of layers (such as variable replacement engines 118A-D) and can then enter into subsequent engines/stages. The disclosed process of variable replacement is described as being operable in a “pipeline” because each init flit enters the first stage, stage 0 118A, or variable replacement engine 0, where it may possibly undergo variable replacement and then moves to the stage 1 118B or variable replacement engine 1 to possibly undergo variable replacement there, and so on. Each init flit continues to move through each of the variable replacement engines 0-3 118A-118D (or stages 0-3) until the variables are replaced with actual values. Each time an init flit moves to the next stage, or variable replacement engine, it leaves an opening for another init flit to enter the variable replacement engine. This process continues until all of the init flit of the template accelerator code 108 are processed and have had variables replaced with actual values, as necessary., para[0036]/ Once the variable replacements are made in all of the init flit of the template accelerator code 108, the code is considered usable accelerator code, and can be referred to as “variable replaced machine code” or “variable replaced AI accelerator code.” Each 128 byte init flit of variable replaced accelerator code can be forwarded to a distribution network 119 to be distributed to the engines 0-7 (115A-115C), or specifically to instruction buffers (IBuffs) 116A-C in the engine 0-7 (115A-C), respectively. Each 128 byte init flit can be divided or sliced into 8×16 byte slices, and the eight (8) slices can then be distributed to the IBuffs 116A-C of the appropriate or corresponding engine of engines 0-7 (115A-115C), para[0038], ln 1-16), and Raha teaches the plurality of layers( IG. 7 depicts multiple PE arrangements including a standard PE array 701, a standard scaled-up PE array 702, and a static MAC scaling PE array 700 according to various embodiments. Each of the PE arrays 700, 701, and 702 include at least one MAC 606 and RF 608 (although not all MACs 606 and RF 608 are labelled in FIG. 7). The standard PE array 701 is an N×N PE 230 array where each PE 230 includes a single MAC 606 and a single RF 608. The scaled-up PE array 702 is a scaled-up version of PE array 701, which has been scaled-up according to conventional techniques. The scaled-up PE array 702 is an N×N×M PE 230 array and includes four times the duplication of PE array 701 (e.g., M=4). The scaled-up PE array 702 increases the TOPS by four in comparison with PE array 701, but the TOPS/W and TOPS/mm.sup.2 remains same as the TOPS/W and TOPS/mm.sup.2 of PE array 701, para[0049], ln 1-8 to para[0050]) for the same reason as to claim 1 above. Claim(s) 2, 3, 4, 5, 6, 7, 13, 14, 15, 16 are rejected under 35 U.S.C. 103 as being unpatentable over Lichtenau ( US 20230305818 A1 ) in view of Wang( US 20200050457 A1) in view of Raha( US 20210271960 A1) and further in view of SHRIVASTAVA( US 20240396844 A1). As to claim 2, SHRIVASTAVA teaches the Al accelerator further comprises: a network on chip (NoC); and an Input-Output Memory Management Unit (IOMMU)( accelerators 1242 can include a single or multi-core processor, graphics processing unit, logical execution unit single or multi-level cache, functional units usable to independently execute programs or threads, application specific integrated circuits (ASICs), neural network processors (NNPs), programmable control logic, and programmable processing elements such as field programmable gate arrays (FPGAs). Accelerators 1242 can provide multiple neural networks, CPUs, processor cores, general purpose graphics processing units, or graphics processing units can be made available for use by artificial intelligence (AI) or machine learning (ML) models, para[0110], ln 10-22/ Packet processing device 1110 can be implemented as one or more of: a microprocessor, processor, accelerator, field programmable gate array (FPGA), application specific integrated circuit (ASIC) or circuitry described herein., para[0084], ln 1-6/ ACC 1120 can execute a virtual switch such as vSwitch or Open vSwitch (OVS), Stratum, or Vector Packet Processing (VPP) that provides communications between virtual machines executed by host 1100 or with other devices connected to a network. For example, ACC 1120 can configure packet processing pipeline circuitry 1140 as to which VM is to receive traffic and what kind of traffic a VM can transmit. For example, packet processing pipeline circuitry 1140 can execute a virtual switch such as vSwitch or Open vSwitch that provides communications between virtual machines executed by host 1100 and packet processing device 1110, para[0089]/ Packet processing device 1110 can include multiple compute complexes, such as an Acceleration Compute Complex (ACC) 1120 and Management Compute Complex (MCC) 1130, as well as packet processing circuitry 1140 and network interface technologies for communication with other devices via a network. ACC 1120 can be implemented as one or more of: a microprocessor, processor, accelerator, field programmable gate array (FPGA), application specific integrated circuit (ASIC) or circuitry described at least with respect to herein, para[0083], ln 1-12/ Examples of operations of packet processing circuitry 1140 include issuance of non-volatile memory express (NVMe) reads or writes, issuance of Non-volatile Memory Express over Fabrics (NVMe-oF™) reads or writes, lookaside crypto Engine (LCE) (e.g., compression or decompression), Address Translation Engine (ATE) (e.g., input output memory management unit (IOMMU) to provide virtual-to-physical address translation), para[0093], ln 5-14/ For example, ACC 1120 can execute a virtual switch such as vSwitch or Open vSwitch (OVS), Stratum, or Vector Packet Processing (VPP) that provides communications between virtual machines executed by host 1100 or with other devices connected to a network. For example, ACC 1120 can configure packet processing pipeline circuitry 1140 as to which VM is to receive traffic and what kind of traffic a VM can transmit. For example, packet processing pipeline circuitry 1140 can execute a virtual switch such as vSwitch or Open vSwitch that provides communications between virtual machines executed by host 1100 and packet processing device 1110, para[0089]/ omponents of examples of switches described herein can be implemented in a switch system on chip (SoC) that includes at least one interface to other circuitry in a switch system, para[0107]). It would have been obvious to one of the ordinary skill in the art before the effective filling date of claimed invention was made to modify the above teaching to incorporate the above feature because this provides secure access, firewall, and per-tenant network isolation at edge and core data center. As to claim 3, SHRIVASTAVA teaches the IOMMU is configured to translate virtual addresses used by the AI accelerator to physical addresses used to store data before transmitting the data from the Al accelerator to the interface(Para[0096]/ para[0055], ln 10-20/ para[0023], ln 1-5/ para[0028]/ para[0090]/ para[0091]) for the same reason as to claim 1 above. As to claim 4, Raha teaches the controller communicates with the array of DPEs through the NoC( para[0029], ln 1-15/Fig.2b) for the same reason as to claim 1 above. As to claim 19, it is rejected for the same reason as to claim 4 above. As to claim 5, Raha teaches the controller communicates with the CPU only through the interface, wherein the interface is a second NoC, wherein the second NoC is larger than the NoC in the AI accelerator( para[0026], ln 1-1-20/ para[0029], ln 10-21) for the same reason as to claim 1 above. As to claim 6, Raha teaches the controller does not contain any programmable logic( para[0033], ln 1-20/ para[089]/ para[0098] to para[0099]) for the same reason as to claim 2 above. As to claim 7, Wang teaches the controller comprises circuitry that is separate from the CPU ( para[0038] to para[0039) and Raha teaches the controller comprises circuitry that is separate from the CPU, wherein the controller is configured to execute software code or firmware for orchestrating the DPEs to perform the task( para[0027], ln 1-19) for the same reason as to claims 1 amd 2. As to claims 13, 14, 15, 16, 20, they are rejected for the same reason as to claims 2-5 above. Claim(s) 8, 9, 17 are rejected under 35 U.S.C. 103 as being unpatentable over Lichtenau (US 20230305818 A1 ) in view of Wang( US 20200050457 A1) in view of Raha( US 20210271960 A1) and further in view of HAN( US 20210375681 A1). As to claim 8, Han teaches controller is configured to control data movement into the memory tiles and interface tiles such that data flows from the interface tiles into the memory tiles, and then from the memory tiles into the DPEs( In some embodiments as shown in FIG. 4, logic wafer 400 includes an array of three by three logic tiles including an artificial intelligence (AI) accelerator logic tile (e.g., logic tile 410) placed in the center of the array, and alternating AI logic tiles (e.g., tiles 401, 403, 406, and 408) and video logic tiles (e.g., tiles 402, 404, 405, and 407) surrounding the AI accelerator logic tile. The array of the three by three logic tiles may be communicatively interconnected by a plurality of global interconnects. IA some embodiments, the AI accelerator logic tile includes a plurality of deep learning processing elements (DPEs), a central processing unit (CPU), and one or more memory controllers interconnected by a first local network on chip (NoC). The one or more memory controllers may be connected to one or more memory tiles on the memory wafer. The AI accelerator logic tile may include a connectivity unit (e.g., including a Peripheral Component Interconnect Express (PCIE) card) configured to be pluggable via a connection to a host system. For example, an IC including the AI accelerator logic tile can be plugged to the host system via the PCIE card for data communication with the host system. In some embodiments, a respective AI logic tile (e.g., substantially similar to AI tile 310) includes a plurality of DPEs, a CPU, and one or more memory controllers interconnected by a local (NoC). The one or more memory controllers may be connected to one or more memory tiles on the memory wafer. In some embodiments, a respective video logic tile (e.g., substantially similar to video tile 330) includes one or more video processing units, a CPU, and one or more memory controllers interconnected by a local NoC. The one or more memory controllers may be connected to one or more memory tiles on the memory wafer, para[0074]). It would have been obvious to one of the ordinary skill in the art before the effective filling date of claimed invention was made to modify the above teaching to incorporate the above feature because this supports certain computing functions including, but not limited to, general computing, machine learning, artificial intelligence (AI) accelerator, edge computing, cloud computing, video codec (e.g., compression or decompression), or video transcoding. As to claim 9, Han teaches wherein each of the DPEs comprises a core, a memory module, and an interconnect, wherein the interconnects in the DPEs are interconnected so that the DPEs are able to transmit data between each other( para[0053], ln 1-16/ para[0070],ln 1-26) for the same reason as to claims 1-30) for the same reason as to claim 8 above. As to claim 17, it is rejected for the same reason as to claim 9 above. Claim(s) 12 is rejected under 35 U.S.C. 103 as being unpatentable over Lichtenau ( US 20230305818 A1 ) in view of Wang( US 20200050457 A1) in view of Raha( US 20210271960 A1) and further in view of Rotem(ACCE). As to claim 12, Rotem teaches controlling the array of DPEs comprises: configuring, using the controller, direct memory access (DMA) circuitry in the DPEs to complete the hardware acceleration task received from the CPU( performance of hardware-based AI accelerators based on an analysis of a substantially static (i.e., predictable) incoming instruction stream. In one example, a computing device capable of performing such a task may include a plurality of special-purpose, hardware-based functional units configured to perform AI-specific computing tasks, para[0004], ln 3-10/ an AI accelerator capable of optimizing its power usage may include a plurality of special-purpose, hardware-based functional units configured to perform AI-specific computing tasks, para[0009], ln 1-6/ of such functional units include, without limitation, matrix multipliers or general matrix-to-matrix multiplication (GEMM) units (which may be used, e.g., to apply filter weights to input data in a neural network, convolve layers in a convolutional neural network, etc.), DMA engines or other memory management units (which may be used, e.g., to access system memory independent of a processing unit), MAC units or other logical operator units (which may perform, e.g., the multiplication and summation operations used when performing convolution operations), caches or other memory devices (implemented using, e.g., SRAM devices, dynamic random access memory (DRAM) devices, etc.) designed to store incoming instruction streams, calculation results, data models, etc., among many others. Special-purpose functional units may be implemented in a variety of ways, including via hard-wiring and/or using application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), para[0035], ln 11- 27). It would have been obvious to one of the ordinary skill in the art before the effective filling date of claimed invention was made to modify the above teaching to incorporate the above feature because this optimizes the power usage and/or performance of AI and ML systems. Response to Arguments Applicant’s arguments with respect to claim(s) 1-21 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Conclusion US 20230305818 A1 teaches intelligence (AI) accelerator code. The system includes: at least one memory ; at least one processor communicatively coupled to the at least one memory, and configured for computing at least one table of variables from a template AI accelerator code; and an AI accelerator including a plurality of engines, and communicatively coupled to the at least one processor and the at least one memory. The AI accelerator is configured to create a variable replaced AI accelerator code for the plurality of engines of the AI accelerator from the template AI accelerator code by replacing variables in the template US 20210375681 A1 teaches In some embodiments as shown in FIG. 4, logic wafer 400 includes an array of three by three logic tiles including an artificial intelligence (AI) accelerator logic tile (e.g., logic tile 410) placed in the center of the array, and alternating AI logic tiles (e.g., tiles 401, 403, 406, and 408) US 20200050457 A1 teaches shown in FIG. 1, the system architecture 100 may include a CPU 11, an artificial intelligence chip 12, and a bus 13. The bus 13 serves as a medium providing a communication link between the CPU 11 and the artificial intelligence chip 12, e.g., a PCIE (Peripheral Component Interconnect Express) bus. CN 108345555 B teaches The AI acceleration processing chip 30 is used for performing AI operation acceleration processing based on the data to be calculated sent by the host CPU, and returning the operation result data to the interface bridge circuit 20. AI acceleration processing chip 30 respectively configured with two Serdes interface, a Serdes interface for data communication with the upper level AI acceleration processing chip or interface bridge circuit, and the other Serdes interface is used for data communication with the next level AI acceleration processing chip. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LECHI TRUONG whose telephone number is (571)272-3767. The examiner can normally be reached 10-8 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Young Kevin can be reached on (571)270-3180. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /LECHI TRUONG/ Primary Examiner, Art Unit 2194
Read full office action

Prosecution Timeline

Dec 22, 2023
Application Filed
Apr 02, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT
Jul 01, 2026
Applicant Interview (Telephonic)
Jul 02, 2026
Examiner Interview Summary
Jul 05, 2026
Response Filed
Sep 08, 2026
Final Rejection mailed — §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743300
DOUBLE-FLASH SWITCHING DEVICE AND SERVER
2y 9m to grant Granted Sep 22, 2026
Patent 12724638
MANAGEMENT APPARATUS, MANAGEMENT METHOD AND MANAGEMENT PROGRAM
2y 10m to grant Granted Sep 01, 2026
Patent 12717621
TRANSPARENTLY EXECUTING ACTIONS WITHIN A CONTAINERIZED CLOUD ENVIRONMENT
3y 8m to grant Granted Aug 25, 2026
Patent 12705088
Task Repacking
4y 5m to grant Granted Aug 11, 2026
Patent 12675339
WORKLOAD MEASURES BASED ON ACCESS LOCALITY
4y 2m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
87%
Grant Probability
99%
With Interview (+36.4%)
3y 0m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 889 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month