Prosecution Insights
Last updated: August 06, 2026
Application No. 19/076,153

ACCELERATOR SYSTEM USING DIGITAL IN-MEMORY COMPUTE CHIPLET DEVICES FOR COMPUTATIONAL WORKLOADS

Non-Final OA §103
Filed
Mar 11, 2025
Priority
Nov 30, 2021 — continuation of 11/847,072 +1 more
Examiner
BARTELS, CHRISTOPHER A.
Art Unit
Tech Center
Assignee
D-Matrix Corporation
OA Round
1 (Non-Final)
68%
Grant Probability
Favorable
1-2
OA Rounds
1y 10m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 68% — above average
68%
Career Allowance Rate
380 granted / 563 resolved
+7.5% vs TC avg
Moderate +12% lift
Without
With
+12.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
22 currently pending
Career history
596
Total Applications
across all art units

Statute-Specific Performance

§101
2.4%
-37.6% vs TC avg
§103
66.1%
+26.1% vs TC avg
§102
24.6%
-15.4% vs TC avg
§112
4.1%
-35.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 563 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This office action is in response to the claim listing filed on March 11th, 2025. Claims 1-20 are currently pending. Information Disclosure Statement The information disclosure statement (IDS) submitted on 03/11/2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Ju (USPGPUB No. 2022/0342666 A1) in view of Hornung et al (USPGPUB No. 2023/0058355 A1, hereinafter referred to as Hornung) in view of CHRYSOS et al. (USPGPUB No. 2022/0100680 A1, hereinafter referred to as Chrysos). Referring to claim 1, Ju discloses an accelerator system {“data center infrastructure layer 1010” (see Fig. 10 [0181]) that includes “node C.R.s 1016(1)-1016(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators” ([0181])}, the system comprising: one or more data gathering devices {“program executed by host processor encodes a command stream in a [gather device] buffer that provides workloads to PPU 3300” (see Fig. 33, [0531], 1st sentence) where PPU 3300 executes a plurality of command streams ([0531], last sentence), a stream per gathering device “a host processor writes a command stream to a buffer” ([0531], last sentence)} comprising: a host computing device coupled {“host processor encodes a command stream in a buffer”, see Fig. 33 [0531], 1st sentence} to the one or more data gathering devices {“that provides workloads to PPU 3300” comprising one or more data gathering device “buffer” ([0531], 1st sentence} and a data for a target application {target application executing on “host processor and PPU 3300 a host interface” (see Fig. 33 [0531])}; wherein each of the chiplets comprises a plurality of tiles {“tiling unit 1858 to accelerate tiling operations for tile-based rendering”, see Fig. 18B [0373], last sentence}; wherein each of the tiles comprises a plurality of slices {“graphics core 1900 can include multiple slices 1901A-1901n or a partition for each core”, see Figs. 19A-19B [0377]}, a CPU coupled to the plurality of slices {“featuring graphics cores 2380A-2380N, which can be modular and are sometimes referred to as core slices”, see Fig. 23 [0433], 1st sentence}; and wherein each of the plurality of slices includes a DIMC device coupled to a clock {“graphics multiprocessor 2134, processing can be performed over consecutive clock cycles”, see Fig. 23, [0431], last two sentences}; Ju does not appear to explicitly disclose wherein the accelerator system is a digital in-memory compute (DIMC) accelerator system; wherein the one or more gather devices configured to obtain computational workload; and wherein each accelerator apparatus includes a global CPU coupled to one or more chiplets; wherein each of the CPUs of the plurality of tiles is configured to receive a portion of the plurality of matrix inputs from the global CPU; However, Hornung discloses wherein the accelerator system is a digital in-memory compute (DIMC) accelerator system {“includes data for the requested memory address, the presence of the [digital in memory] in-flight memory request” (see Fig. 3 [0087]) said memory requests sent by accelerator “HTP and HTF accelerators of the CNM system 102” (see Figs. 2 and 3, [0075], 1st sentence)}; wherein the one or more gather devices configured to obtain computational workload {“maintain workload balance across the HTP 140 module and the HTF 142 module”, see Figs. 1 and 2, [0072], last sentence}; and wherein each accelerator apparatus includes a global CPU coupled {per “memory-compute device 112” comprises a global CPU “internal host processor, see Fig. 1 [0066], 1st and 2nd sentence} to one or more chiplets {a plurality of “memory-compute device 112 can comprise a chiplet-based architecture”, see Fig. 1, [0058], 2nd sentence}; wherein each of the CPUs of the plurality of tiles {CPU tile “tiling unit 1858”, see Figs. 18a and 18b, [0373], last sentence} is configured to receive a portion {“scene are subdivided in image space”, see Figs. 18a-18b [0373], last sentence} of the plurality of matrix inputs from the global CPU {“accelerate tiling operations for [matrix inputs] tile-based rendering, in which rendering operations for a scene”, see Figs. 18 and 18b [0373], last sentence}; Ju and Hornung are analogous art because they are from the same problem-solving area, method and systems for parallel computing tasks. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Ju and Hornung before him or her, to modify Ju’s “data center infrastructure layer 1010” (see Fig. 10 [0181]) incorporating Hornung’s “HTP and HTF accelerators of the CNM system 102” (see Figs. 2 and 3, [0075], 1st sentence). The suggestion/motivation for doing so would have been to provide optimizing connectivity and leveraging the optimized connectivity by controlling compute resource usage in a coarse grained reconfigurable architecture to thereby conserve power and reduce latency throughout the architecture (Hornung ([0025], last sentence). Therefore, it would have been obvious to combine Hornung with Ju to obtain the invention as specified in the instant claim(s). Neither Ju or Hornung appears to explicitly disclose wherein the host computing device is configured to compile the computational workload data in an instruction set architecture (ISA) graph and to execute the ISA graph using the plurality of accelerator apparatuses; wherein each of the DIMC devices is configured to perform a throughput of one portion of the plurality of matrix inputs from the global CPU; and or more matrix computations according to the ISA graph using one or more of the plurality of matrix inputs such that the throughput is characterized by a plurality of multiply accumulates per a clock cycle; Furthermore, CHRYSOS discloses wherein the host computing device is configured to compile the computational workload data {“second compilation phase starts with the minimal dataflow graph [virtual ISA] VISA representation and produces an executable” ([0326], last two sentences) performed by following “a minimal dataflow graph is produced prior to binding to the specific DFE implementation” ([0326], 3rd sentence) DFE via accelerator apparatuses “a dataflow execution circuit accelerator 3400 including a plurality of dataflow execution circuits 3200A-D” ([0320], 1st line)} in an instruction set architecture (ISA) graph {“base set of functionality and optional instruction set classes,” by the example of RISC-V, see Fig. 6a [0117], last sentence} and to execute the ISA graph using the plurality of accelerator apparatuses {“original native (e.g., x86) ISA binary for the region is executed on the main core” (see Fig. 35 [0326], last sentence}; wherein each of the DIMC devices is configured to perform a throughput {“latency may not be a concern so long as fabric throughput is maintained[/performed]”, see Fig. 62 [0528, 2nd sentence]} of one portion of the plurality of matrix inputs from the global CPU {“a snapshot 6200 of an in-flight, pipelined [matrix] extraction according to embodiments of the disclosure” (see Fig. 62 [0528], 1st sentence) pipelined routing/networking “Several local networks may be ganged together to form routing channels, e.g., which are interspersed (as a grid) between rows and columns of PEs” (see Figs. 8 and 9, [0202])}; and one or more matrix computations {“CSA dataflow execution may depend (e.g., only) on highly localized status, for example, resulting in a highly scalable architecture with a distributed, asynchronous execution model. Dataflow operators may include arithmetic dataflow operators, for example, one or more of floating point addition and multiplication” ([0155], 4th and 5th sentences} according to the ISA graph {“multiplier node 308 of [ISA] dataflow graph 300”, see Fig. 3a [0157], 3rd sentence} using one or more of the plurality of matrix inputs {“operation of selecting input X with pick node 304, multiplying X by Y e.g., multiplication node 308” (see Fig. 3a [0156], 5th sentence) each input including matrix “include dataflow operators for vectorized, low precision arithmetic” ([0155], last sentence) in a CSA data flow execution vectors by LLVM by a “sequential unit may also provide a model for handling code that does not fit in the [matrix] spatial array” (see Fig. 64, [0533]).} such that the throughput is characterized {“ energy efficient [CSA] dataflow processing elements (and/or communications network (e.g., a network dataflow endpoint circuit thereof)) to form [characterized by] a high-throughput, low-latency, energy-efficient HPC fabric”, [0177], 1st sentence} by a plurality of multiply accumulates {“switch node 406 (e.g., to provide its input out of port “0” [accumulates] to a destination (e.g., a downstream processing element)” (see Fig. 4, [0160]) as the parallel function of switches used to take intermediate results to a final output/classification when formed “latency-insensitive channels provide a critical abstraction layer” ([0159], last two sentences) where “Several local networks may be ganged together to form routing channels, e.g., which are interspersed (as a [matrix] grid) between rows and columns of PEs” (see Fig. 8, [0202], 8th sentence)} per a clock cycle {“in a pipelined fashion with no more than one [clock] cycle of latency.”, see Figs. 3 and 4 [0159], last three sentences; another example of clock cycle “clock frequency to ensure digital timing discipline”, see Figs. 61 and 70 [0583], last two sentences}.` Ju/Hornung and CHRYSOS are analogous art because they are from the same problem-solving area, method and systems for parallel computing tasks. Before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art, having the teachings of Ju/Hornung and CHRYSOS before him or her, to modify Ju/Hornung’s system incorporating Chrysos’ “a dataflow execution circuit accelerator 3400 including a plurality of dataflow execution circuits 3200A-D” ([0320], 1st line). The suggestion/motivation for doing so would have been to provide CSA architecture have been co-designed with a compilation tool chain used to mitigate programmability risks surrounding revolutionary architecture of spatial array of processing elements (Chrysos ([0135] paraphrased). Therefore, it would have been obvious to combine CHRYSOS with Ju/Hornung to obtain the invention as specified in the instant claim(s). As per claim 2, the rejection of claim 1 is incorporated and Hornung discloses wherein the one or more chiplets of each accelerator apparatus are coupled to one or more double data rate (DDR) {“AIB I/O cells support three clocking modes: asynchronous (i.e. non-clocked), SDR, and [double data rate] DDR”, [0127], last 4th sentences} dynamic random access memory (DRAM) devices {“associated with one or more DRAM devices”, see Fig. 6a, [0136], last sentence} using a DRAM interface {“Double Data Rate (DDR) interface connecting the memory controller chiplet 614 to a dynamic random access memory (DRAM) memory device chiplet 616”, see Fig. 6, [0135], last sentence}; and wherein the DDR DRAM devices are configured to store the plurality of matrix inputs {“from vector length information. The mask can be used to specify a number of bytes to write for a particular synchronous flow, including for conditions where a vector length is other than a multiple of the data path width”, see Fig. 10a, [0176], 1st and 2nd sentences} using the DRAM interface {“compute operations [including vectors] for the tile: a Multiply and Shift Operation block” that includes “mask information can be used to control downstream operations” (both citations in [0116]) sentence) the memory mask operation via “memory interface 1020” performed by memory device chiplet 616 can be, or include any combination of, volatile memory devices or non-volatile memories. Examples of volatile memory devices include, but are not limited to, random access memory (RAM)—such as DRAM) synchronous DRAM (SDRAM)” (see Fig. 6a, [0139], 1st sentence}. As per claim 3, the rejection of claim 1 is incorporated and Hornung discloses wherein each of the CPUs in each of the tiles is coupled {“[coupling] host interface 724 can be configured to command individual [each of the tiles within] HTF tile clusters” (see Fig. 7 [0146])} to a peripheral component interconnect express (PCIe) bus {such clusters/each of the tiles couple within “first CNM package 700 can include a PCIe scale fabric interface (PCIe/SFI)” (see Fig. 7 [0145], last two sentences}; wherein a main bus device {“[main bus device] dispatch interface or memory interface can be used to initialize or prepare a mask field for data to be communicated throughout the HTF cluster 502”, see Fig. 5, [0101], 2nd sentence} is coupled to each PCIe bus in each chiplet using a master chiplet device {“ A [[master] first or base tile (e.g., Tile N−2, in the example of FIG. 5) of a synchronous flow can initiate a thread of work through the pipelined tiles”, see Fig. 5, [0112], 3rd sentence}; wherein the master chiplet device is coupled to each of the other chiplet devices using at least a plurality of die-to-die (D2D) interconnects {“a greater number of [die-to-die] low-latency paths can be available to realize flows”, [0027], last there sentence} coupled to each of the CPUs in each of the tiles {“ passthrough channel that can provide a low-latency communication datapath”, see Fig. 5 [0108], 1st sentence}. As per claim 4, the rejection of claim 1 is incorporated and Ju discloses wherein each of the plurality of slices is coupled to a network on chip (NoC) device {“network-on-chip, or with dedicated connections”, see Fig. 26 [0456], last two sentences} configured to perform a multicast process {“allows synapses to be allocated to [multicast process] different neurons 2602 as needed based on neural network topology and neuron fan-in/out”, see Fig. 26 [0456]}. As per claim 5, the rejection of claim 1 is incorporated and Chrysos discloses wherein each of the DIMC devices is configured to wherein each the DIMC devices is configured to support a block structured support {“[block structured support] aliased the MMX packed integer flat register file 9850—in the embodiment illustrated, the x87 stack is an eight-element stack used to perform scalar floating-point operations on 32/64/80-bit floating point data”, [0833]} one or more block floating point data types {more than one type “scalar floating point, packed integer, packed floating point, vector integer, vector floating point”, see Figs. 99a, 99b, [0843]} using a shared exponent {“the [shared] addend exponent may be observed in advance of multiplication ”, [0525], last two sentences}; and sparsity {“CSA embodiments herein may provide for more computational density and energy efficiency” where the lack of density becomes sparse, see Figs. 73 and 74, [05921], last sentence}. As per claim 6, the rejection of claim 1 is incorporated and Chrysos discloses wherein the one or more data gathering devices includes a web-scraping device {Examiner’s note: recitation “or” renders this dependent claim as a Markush claim, thereby the reference needs only disclose one member in the group to address the claim}, a dataset reader device, a crowdsourcing device, a sensor device, a simulation device, or an Internet of Things (IoT) network {“characteristics relevant to all forms of computing ranging from supercomputing and datacenter to the internet-of-things”, [0723], last sentence}. As per claim 7, the rejection of claim 1 is incorporated and Chrysos discloses wherein the host computing device includes a compiler stack {“there are two phases of compilation. In one example of a first phase, a minimal dataflow graph is produced prior to binding to the specific DFE implementation”, see Figs. 35 and 36, [0326], 1st three sentences} configured to determine the ISA graph using the computational workload data {“how much of the [computational] workload might be successfully offloaded to a DFE.” And thereby determined as a dataflow graph to a DFE, [0303], last sentence}. As per claim 8, the rejection of claim 1 is incorporated and Chrysos discloses wherein the host computing device includes a workload preprocessor configured to determine a plurality of workload parameters {“minimal dataflow graph expressed in a virtual ISA (VISA) and comes with [workload parameter] metadata that”, see Fig. 35, [0326], 3rd sentence} using the ISA graph {“second compilation phase starts with the minimal dataflow graph VISA representation”, see Figs. 35 [0326], last two sentences}. As per claim 9, the rejection of claim 1 is incorporated and Chrysos discloses wherein the host computing device includes an execution stack {“produces an executable [stack] (e.g., the dataflow operation entries) specific to the DFE target implementation”, see Figs. 35 and 36 [0326], last two sentences} configured to transfer the ISA graph to the plurality of accelerator apparatuses {“the original native (e.g., x86) [transferring] ISA binary for the region is executed on the main core” if the DFE doesn’t exist or compilation fails, see Figs. [0326], last sentence}. As per claim 10, the rejection of claim 1 is incorporated and Chrysos discloses wherein the target application includes natural language processing (NLP) {Examiner’s note: recitation “or” renders this dependent claim as a Markush claim, thereby the reference needs only disclose one member in the group to address the claim}, autonomous reasoning/decision-making, video/image processing, cybersecurity/fraud detection, manufacturing/industrial processes, agentic artificial intelligence (AI), or smart cities/Internet of Things (IoT) {“characteristics relevant to all forms of computing ranging from supercomputing and datacenter to the internet-of-things”, [0723], last sentence}. Referring to claims 11-17 are system claims reciting claim functionality corresponding to the system claim of claims 1-10, respectively, thereby rejected under the same rationale as claims 1-10 recited above, inter alia, Chrysos discloses wherein the compiler stack includes a handles layer {“there are two phases of compilation. In one example of a first phase, a minimal dataflow graph is produced prior to binding to the specific DFE implementation”, see Figs. 35 and 36, [0326], 1st three sentences} configured to determine references to resources for the program {“such [resources] functions may be configured (e.g., by a user and not a manufacturer) into the fabric based on the requirement of each application”, [0180]}, workload, or model of the target application {“based on the requirement of each [target] application”, [0180]}; and wherein the host computing device includes an execution stack configured {“produces an executable [stack] (e.g., the dataflow operation entries) specific to the DFE target implementation”, see Figs. 35 and 36 [0326], last two sentences} to transfer the ISA graph {“multiplier node 308 of [ISA] dataflow graph 300”, see Fig. 3a [0157], 3rd sentence} to the plurality of accelerator apparatuses {“the original native (e.g., x86) [transferring] ISA binary for the region is executed on the main core” if the [accelerator] DFE doesn’t exist or compilation fails, see Figs. [0326], last sentence}; and Hornung discloses a hardware dispatch device coupled to the CPU {“dispatch interface or memory interface can be used to initialize or prepare a mask field for data to be communicated throughout the HTF cluster 502” and respective claimed CPUs, see Fig. 5, [0101], 2nd sentence}; The 103 motivation for independent claim 11 relied upon as recited in claim 1 above. Referring to claims 18-20 are system claims reciting claim functionality corresponding to the system claim of claims 1-10, respectively, thereby rejected under the same rationale as claims 1-10 recited above, inter alia, Chrysos discloses wherein each of the tiles comprises a RISC CPU coupled to receive a plurality of matrix inputs {“operation of selecting input X with pick node 304, multiplying X by Y e.g., multiplication node 308” (see Fig. 3a [0156], 5th sentence) each input including matrix “include dataflow operators for vectorized, low precision arithmetic” ([0155], last sentence) in a CSA data flow execution vectors by LLVM by a “sequential unit may also provide a model for handling code that does not fit in the [matrix] spatial array” (see Fig. 64, [0533])} using a global reduced instruction set computer (RISC) interface {“most CSA operations combined with a [global] traditional RISC-like control-flow architecture”, see Fig. 64 [0533]}; wherein each of the RISC CPUs of the plurality of tiles {“Accelerator tile 100 may be a portion of a larger tile”, see Fig. 1, [0137], 1st and 2nd sentences} is configured to receive a portion of the plurality of matrix inputs {“a snapshot 6200 of an in-flight, pipelined [matrix] extraction according to embodiments of the disclosure” (see Fig. 62 [0528], 1st sentence) pipelined routing/networking “Several local networks may be ganged together to form routing channels, e.g., which are interspersed (as a grid) between rows and columns of PEs” (see Figs. 8 and 9, [0202])}; wherein the global RISC interface is configured to map each attention layer {“configures the [attention layer subcomponents] PEs and network for execution”, see Fig. 63, [0531]} on to one of the plurality of slices to communicate with the RISC CPU associated {plurality “variety of ways including time sliced multithreading”} with the tile of the slice to process the portion of the workload associated with the attention layer {“CSA captures most vector-parallel workloads such that most vector-style workloads run directly on the CSA”, [0168], last sentence}. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The following references are indicative the current state of the art regarding claim 1’s “accelerator system”, “matrix input”, or “die to die interconnects” (see claim 3): US 20250117223 A1, US 20230316075 A1, US 20230282156 A1, US 20220342666 A1, US 20220100680 A1, US 20210133123 A1, and US 6434634 B1. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHRISTOPHER A. BARTELS whose telephone number is (571)270-3182. The examiner can normally be reached on Monday-Friday 9:00a-5:30pm EST. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dr. Henry Tsai can be reached on 571-272-4176. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C. B./ Examiner, Art Unit 2184 /HENRY TSAI/ Supervisory Patent Examiner, Art Unit 2184
Read full office action

Prosecution Timeline

Mar 11, 2025
Application Filed
Jul 23, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699666
HYBRID BUS COMMUNICATION CIRCUIT
3y 5m to grant Granted Aug 04, 2026
Patent 12699663
A/D CONVERTER CONTROL CIRCUIT
2y 2m to grant Granted Aug 04, 2026
Patent 12693985
ACCESS FOR COMPUTE NODES TO A STORAGE SERVER THROUGH A PCI EXPRESS FABRIC
2y 1m to grant Granted Jul 28, 2026
Patent 12681877
METHOD AND APPARATUS FOR CACHE TIERING
2y 3m to grant Granted Jul 14, 2026
Patent 12670108
SYSTEM AND METHOD FOR LOW LATENCY PACKET PROCESSING
2y 11m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
68%
Grant Probability
80%
With Interview (+12.2%)
3y 3m (~1y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 563 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month