DETAILED ACTION
This Office Action is in response to claims filed on 04/08/2026.
Claims 1-22 and 27-30 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see page 7 of applicant's remarks, filed 04/08/2026, with respect to 35 U.S.C. 112(a) and 112(b) rejection of claims 23-26 have been fully considered and are persuasive. The rejection of 02/12/2026 has been withdrawn.
Applicant’s arguments, see page 7 and 8 of applicant's remarks, filed 04/08/2026, with respect to 35 U.S.C. 112(b) rejection of claims 11-22 and 27-30 have been fully considered and are persuasive. The rejection of 02/12/2026 has been withdrawn.
Applicant’s arguments with respect to claim(s) 1-22 and 27-30 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-5 and 9-10 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou et al. Patent No. US 11,093,276 B2 (hereinafter Zhou) in view of Chen et al. Patent No. US 11,709,664 B2 (hereinafter Chen) in view of Temam et al. Patent No. US 11,501,144 B2 (hereinafter Temam).
With regard to claim 1, Zhou teaches an apparatus comprising:
a plurality of slices, wherein each slice of the plurality of slices is configured for distributed information processing (Col. 2, Chip communication system 102 can include a global manager 1022 and a plurality of cores 1024. Global manager 1022 can include at least one task manager to coordinate with one or more cores 1024); and
However, Zhou does not explicitly teach a plurality of dedicated databuses coupled to slices configured for coordination of distributed information processing.
a plurality of dedicated databuses, wherein each slice of the plurality of slices is coupled to one of the plurality of dedicated databuses (FIG. 13a, Processor array network coupled to databuses in a grid; Col. 23, lines 4-6, In this example, the array of configurable units 1330 includes a plurality of types of configurable units, which are configured with the anti-congestion logic 232; Col. 23, lines 42-51, The array level network includes links interconnecting configurable units in the array. The links in the array level network include one or more and, in this case, three kinds of physical buses: a chunk-level vector bus (e.g., 128 bits of data), a word-level scalar bus (e.g., 32 bits of data), and a multiple bit-level control bus. For instance, interconnect 1321 between switch units 1311 and 1312 includes a vector interconnect with a vector bus width of 128 bits, a scalar bus interconnect with scalar bus width of 32 bits, and a control interconnect) and each slice of the plurality of slices is configured to coordinate locally the distributed information processing (Col. 9, Coarse-grained reconfigurable architectures (CGRAs) comprise distributed compute and memory components in a programmable interconnect fabric. Applications 102 are executed on CGRAs in a distributed fashion by programming the individual compute and memory components to asynchronously receive, process, and send data and control information)
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Chen with the teachings of Zhou in order to provide an apparatus that teaches a plurality of databases coupled to computing slices configured for local coordination of distributed processing. The motivation for applying Chen teaching with Zhou teaching is to provide an apparatus that allows for efficient propagation of control events throughout an array of computing slices through the plurality of associated databuses, such that the system can manage the rate of execution of pipeline stages and prevent buffer overflows and processing bottlenecks (Chen, Col. 3). Zhou and Chen are analogous art directed towards task management and arrangements of distributed systems. Therefore, it would have been obvious for one of ordinary skill in the art to combine Chen with Zhou to teach the claimed invention in order to provide a system which enables efficient control signaling and communication among an array of computing slices through associated databuses.
However, Zhou and Chen do not explicitly teach locally coordinating each slice independently using an asynchronous local output sync signal indicating execution completion of a workload batch.
In a similar field of endeavor, Temam teaches each slice of the plurality of slice is configured to coordinate locally the distributed information processing (Col. 10, FIG. 2 illustrates an example … compute tile 200 … In various implementations, compute tile 200 may also be referenced or referred to as computing unit 200. Each compute tile 200 is self-contained computational unit configured to execute instructions independently relative other corresponding tiles within tile sets 112, 114; Col. 10, lines 63-66, the different instruction types are executed by independent control units within compute tile 200 that synchronize on data through sync flag controls that are managed within compute tile 200 (Examiner notes: per-tile local coordination).) using an asynchronous local output sync signal (Col. 10-Col. 11, The sync flag controls manage concurrency between executions of different instruction types within compute tile 200. Each compute operation associated with each instruction type will be executed in strict order of issuance (i.e., First-In, First-Out).) to indicate completion of the execution of a workload batch (Col. 21, every write by a tile to a ring output will trigger an increment of the corresponding sync flag count. Controller 702 may examine the payload data to determine the number of data chunks or segments that comprise the payload (Examiner notes: workload batch). Controller 702 then monitors execution by the tile to ensure the expected number of data segments are forwarded and/or consumed by the tile before another tile executes in master mode) in each slice independently (Col. 4, Each computing unit of the hardware computing system is self-contained and can independently execute computations required).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Temam with the teachings of Zhou and Chen in order to provide an apparatus that teaches a local coordination mechanism of a processing tile indicating execution completion of a workload batch per-tile. The motivation for applying Temam teaching with Zhou and Chen teaching is to provide an apparatus that allows for an individual tile to maintain its own thread of control independent from other tiles, such that enables independent execution of instructions per-tile while maintaining data synchronization (Temam, Col. 10 and Col. 21). Zhou and Chen and Temam are analogous art directed towards arrangements of distributed systems. Therefore, it would have been obvious for one of ordinary skill in the art to combine Temam with Zhou and Chen to teach the claimed invention in order to provide per-tile local execution completion synchronization signal.
With regard to claim 2, Zhou teaches the apparatus of claim 1, wherein the each slice of the plurality of slices includes a memory unit (Col. 4, lines 22-25, CGRA 200 can includes cores 210-217 and buffers 220-227. Each of cores 210-217 can be connected to its own memory buffer (e.g., buffers 220-227).
With regard to claim 3, Zhou teaches the apparatus of claim 2, wherein the each slice of plurality of slices is a processing unit (Col. 2, lines 57-62, Cores 1024 can include one or more processing elements that each include single instruction, multiple data (SIMD) architecture including one or more processing units configured to perform one or more operations).
With regard to claim 4, Zhou teaches the apparatus of claim 2, further comprising a plurality of current workload batches (Col. 5, lines 4-9, Buffer controller 402 can generate instructions for performing a plurality of buffer transactions on at some of buffers 410-417. For example, on-chip system 100 can receive a computation task (e.g., deep learning), and buffer controller 402 can generate instructions based on the received task).
With regard to claim 5, Zhou teaches the apparatus of claim 4, further comprising a plurality of current workload batches is stored within the memory unit the each slice of the plurality of slices (Col. 5, lines 9-11, These instructions can be assigned to data managers corresponding to the at least some of buffers 410-417 for execution; Col. 5, lines 37-40, As shown in FIG. 4, system 400 includes buffers 410-417 and each buffer can further include a plurality of data units indicated by blocks in the buffer).
With regard to claim 9, Zhou reasonably teaches the plurality of cores executing buffer transactions that include read and write operations (Col. 5). However, Zhou does not explicitly teach the apparatus comprising data producers configured to execute write requests.
Chen teaches the apparatus of claim 1, further comprising a plurality of data producers (Col. 12, lines 45-48, The buffer classification logic 212 is configured to classify the stage buffers as producers and consumers on a stage-by-stage basis by classifying those stage buffers that provide input data to a particular stage as the producers), wherein each of the plurality of data producers is configured to execute a write request (Col. 12, lines 51-57, The control connections creation logic 222 is configured to create control connection between the stage buffers by extending the control connections from the consumers in the particular stage. The control connections extend from a particular consumer to one or more corresponding producers that write data into the particular consumer) in each of the plurality of slices (Col. 14, lines 16-20, One skilled in the art will appreciate that the data processing pipeline can comprise a plurality of producers, a plurality of compute nodes, and a plurality of consumers such that that a compute node can receive input from multiple producers and can provide output to multiple consumers).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Chen with the teachings of Zhou and Temam in order to provide an apparatus that teaches a plurality of data producers configured to perform write requests in each of the plurality of slices. The motivation for applying Chen teaching with Zhou and Temam teaching is to provide an apparatus that allows for classification of stage buffers such that the appropriate control connections between stage buffers can be applied for the particular stage. For example, write buffer designations control when output data can be written by associated data producers to maintain consumer data saturation (Chen, Col. 7). Zhou, Temam, and Chen are analogous art directed towards task management and arrangements of distributed systems. Therefore, it would have been obvious for one of ordinary skill in the art to combine Chen with Zhou and Temam to teach the claimed invention in order to provide the appropriate stage buffer classification which controls data transmission between compute slices.
With regard to claim 10, Zhou reasonably teaches the plurality of cores executing buffer transactions that include read and write operations (Col. 5). However, Zhou does not explicitly teach the apparatus comprising data consumers configured to execute read requests.
Chen teaches the apparatus of claim 1, further comprising a plurality of data consumers (Col. 12, lines 45-50, The buffer classification logic 212 is configured to classify the stage buffers as producers and consumers on a stage-by-stage basis by … classifying those stage buffers that store output data from the particular stage as the consumers), wherein each of the plurality of data consumers is configured to execute a read request (Col. 12, lines 61-67, The anti-congestion logic 232 is configured to configure each of the producers with a ready-to-read credit counter, such that the ready-to-read credit counter of a particular producer is initialized with as many read credits as a buffer depth of a corresponding consumer that reads data from the particular producer) in each of the plurality of slices (Col. 14, lines 16-20, One skilled in the art will appreciate that the data processing pipeline can comprise a plurality of producers, a plurality of compute nodes, and a plurality of consumers such that that a compute node can receive input from multiple producers and can provide output to multiple consumers).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Chen with the teachings of Zhou and Temam in order to provide an apparatus that teaches a plurality of data consumers configured to perform read requests in each of the plurality of slices. The motivation for applying Chen teaching with Zhou and Temam teaching is to provide an apparatus that, analogous to data producers, allows for classification of stage buffers such that the appropriate control connections between stage buffers can be applied for the particular stage. For example, read buffer designations control the amount of input data needed to fully saturate an associated data consumer and signals buffer availability (Chen, Col. 7). Zhou, Temam, and Chen are analogous art directed towards task management and arrangements of distributed systems. Therefore, it would have been obvious for one of ordinary skill in the art to combine Chen with Zhou and Temam to teach the claimed invention in order to provide appropriate stage buffer classification which controls data transmission between compute slices.
Claims 6, 7, and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Zhou in view of Chen in view of Temam as applied to claim 2 and 5 above, and further in view of Vembu et al. Pub. No. US 2023/0039853 A1 (hereinafter Vembu).
With regard to claim 6, Vembu teaches the apparatus of claim 5, wherein each of plurality of current workload batches is different from another of the plurality of current workload batches ([0159], Instead, each tile executes a subset of the submitted workloads. In one embodiment a given tile can execute a specific subset of the workload (Examiner notes: A workload is divided and distributed into different batches) based on an identifier provided to a hardware context associated with the tile; [0153], As shown in FIG. 16A, the graphics processing system 1600 includes an application and/or graphics driver (app/driver 1601) that can send workloads 1604A-1604D to one or more engine block tiles 1605A-1605D, which can be similar to or variants of the engine block tiles 1524A-1524N of FIG. 15. The workloads 1604A-1604D can be part of the same workload and/or separate workloads).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Vembu with the teachings of Zhou, Chen, and Temam in order to provide an apparatus that teaches a plurality of workload batches different from each other. The motivation for applying Vembu teaching with Zhou, Chen, and Temam teaching is to provide an apparatus that allows for the use of known techniques of multi-tasking in multi-core systems such that different workloads can be executed simultaneously, leading to improved efficiency of the system. Zhou, Chen, Temam, and Vembu are analogous art directed towards task management and arrangements of distributed systems. Therefore, it would have been obvious for one of ordinary skill in the art to combine Vembu with Zhou, Chen, and Temam to teach the claimed invention in order to provide efficient resource allocation by dividing different workload batches to different processors.
With regard to claim 7, Vembu teaches the apparatus of claim 5, wherein at least two of the plurality of current workload batches are the same ([0159], In one embodiment, to enable cross-tile workloads, the same batch buffer containing the superset of work items to be performed is submitted to each tile that is to be included within tile work-group. All commands are submitted to all tiles (Examiner notes: A single workload is distributed as identical batches) that are to execute the commands, even if a given tile is not intended to execute all submitted workloads; [0153], As shown in FIG. 16A, the graphics processing system 1600 includes an application and/or graphics driver (app/driver 1601) that can send workloads 1604A-1604D to one or more engine block tiles 1605A-1605D, which can be similar to or variants of the engine block tiles 1524A-1524N of FIG. 15. The workloads 1604A-1604D can be part of the same workload and/or separate workloads).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Vembu with the teachings of Zhou, Chen, and Temam in order to provide an apparatus that teaches a plurality of workload batches can be associated with the same workload. The motivation for applying Vembu teaching with Zhou, Chen, and Temam teaching is to provide an apparatus that allows for the use of known techniques of parallel execution in order to speed up the execution of a workload by increasing the processing throughput of the plurality of identical workloads. Zhou, Chen, Temam, and Vembu are analogous art directed towards task management and arrangements of distributed systems. Therefore, it would have been obvious for one of ordinary skill in the art to combine Vembu with Zhou, Chen, and Temam to teach the claimed invention in order to provide parallel execution of identical workloads.
With regard to claim 8, Vembu teaches the apparatus of claim 2, further comprising a plurality of external memory units configured to store a plurality of future workload batches (FIG. 16B, Workload 1605 held in Batch Buffer Submitter 1618; [0152], FIG. 16B shows an example of a system graphics interface 1602; [0157], The doorbell 1603 is one of the multiple doorbell interface through which a workload 1605 can be submitted, where the workload 1604 can be anyone one of workloads 1604A-1604D of FIG. 16A … In one embodiment the work request is provided in the form of a buffer of batched commands (e.g., a batch buffer). The batch buffer can be processed via the batch buffer submitter 1618 (Examiner notes: Batch Buffer Submitter is an intermediary memory device to store workloads for scheduling). In one embodiment, the system/device address translator 1616 to translate system addresses to device local address for the engine block tile. The commands of the batch buffer can then be submitted to the associated engine block tile), wherein each of the plurality of slices is associated with one of the plurality of future workload batches (FIG. 16A, Plurality of Engine Block Tiles 1605A-D coupled to System Graphics Interface 1602A-D holding Workloads 1604A-D; [0152], FIG. 16A shows an overview of the graphics processing system 1600, according to an embodiment; [0153], As shown in FIG. 16A, the graphics processing system 1600 includes an application and/or graphics driver (app/driver 1601) that can send workloads 1604A-1604D to one or more engine block tiles 1605A-1605D).
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Vembu with the teachings of Zhou, Chen, and Temam in order to provide an apparatus that teaches memory units configured to queue workloads for execution. The motivation for applying Vembu teaching with Zhou, Chen, and Temam teaching is to provide an apparatus that allows for pre-processing of future workload batches during the execution of a current batch such that enables efficient scheduling and execution upon batch dispatching (Vembu, [0157]). Zhou, Chen, Temam and Vembu are analogous art directed towards task management and arrangements of distributed systems. Therefore, it would have been obvious for one of ordinary skill in the art to combine Vembu with Zhou, Chen, and Temam to teach the claimed invention in order to provide workload queuing for associated computing slices.
Claims 11-22 and 27-30 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. Patent No. US 11,709,664 B2 (hereinafter Chen) in view of Chandra et al. Patent No. US 5,978,936 (hereinafter Chandra) in view of Temam et al. Patent No. US 11,501,144 B2 (hereinafter Temam).
With regard to claim 11, Chen teaches a method for implementing slice (Col. 14, lines 6-10, A data processing pipeline/operation comprises at least a producer, a compute node, and a consumer; Col. 14, lines 30-35, In the context of this application, a producer can be referred to as an upstream buffer or upstream memory node/unit, a compute node can be referred to as an intermediary compute node/unit or an intermediary processing node/unit, and a consumer can be referred to as a downstream buffer or downstream memory node/unit) coordination, the method comprising (Col. 7, lines 1-6 and lines 19-22, A computer-implemented method is described that includes accessing a dataflow graph with compute nodes that asynchronously transmit data along data connections. The data flow graph includes a loop nest in which loops are arranged in a hierarchy of levels, such that a loop at a second level is within a loop at a first level. The method includes … controlling data transmission between the compute nodes along the data connections by using control connections to control writing of the data by the producers into the consumers):
executing a write request with local coordination (Col. 15, lines 10-12 and lines 22-28, The write credit counter is configured to decrement when the producer begins writing the buffer data unit into the consumer … The write credit counter is configured to increment when the producer receives from the consumer a write done token. The write done token indicates to the producer that the writing of the buffer data unit into the producer that the writing of the buffer data unit into the consumer has completed. The producer resumes writing data into the consumer when the producer receives the write done token from the consumer. The producer stops writing data into the consumer when the write credit counter has zero write credits (Examiner notes: Where write request coordination is managed by the credit counters of the associated slice) in a first batch of a plurality of batches in each slice of a plurality of slices (Col. 15, lines 48-49, At a first timestep, the producer begins writing a first buffer unit data into the consumer…;
setting a local event tracker (Col. 3, lines 56-58, The compiler is further configured to configure each of the producers (Examiner notes: each of the slice) with a ready-to-read credit counter; Col. 4, lines 16-18, In some embodiments, the compiler is further configured to configure each of the producers with a write credit counter that is initialized with one or more write credits; Col. 15, lines 36-38, As discussed above, the ready-to-read credit counter is initialized with three read credits and the write credit counter is initialized with two write credits) in the first batch of the plurality of batches in the each slice of the plurality of slices (Col. 15, lines 49-52, In response, the ready-to-read counter is decremented by one (read credit = 2) and the write credit counter is also decremented by one (write credit = 1);
monitoring the write request in the first batch of the plurality of batches in the each slice of the plurality of slices (Col. 15, lines 53-64, At a second timestep, the producer begins writing a second buffer unit data into the consumer. In response, the ready-to-read credit counter is further decremented by one (read credit=1) and the write credit counter is further decremented by one (write credit = 0). The write credit counter expires and therefore the producer stops writing data into the consumer (writing stopped). At a third timestep, the writing of the first buffer unit data into the consumer is complete. In response, the consumer sends a first write done token to the produce. In response, the write credit counter is incremented by one and therefore reactivated (write credit = 1) (Examiner notes: Where slice request execution is monitored by the credits available);
executing a read request with the local coordination (Col. 15, lines 7-11 and lines 12-22, The ready-to-read counter is configured to decrement when the producer begins writing a buffer data unit into the consumer. The size of the buffer data unit is s (e.g., 16 bytes, 64 bytes, 512 bytes … The ready-to-read credit counter is configured to increment when the producer receives from the consumer a read ready token. The read ready token indicates to the producer that the consumer has freed a buffer data unit and is ready to receive an additional buffer data unit. The producer stops writing data into the consumer when the ready-to-read credit counter has zero read credits. The producer resumes writing data into the consumer when the producer receives read ready token from the consumer) in the first batch of the plurality of batches in the each slice of the plurality of slices after monitoring the write request is completed (Col. 16, lines 31-40, At the fifth timestep, the producer begins writing a third buffer unit data into the consumer. In response, the ready-to-read credit counter is further decremented by one (read credit = 0) and the write credit counter is decremented by one (write credit = 1). The ready-to-read credit counter expires and therefore the producer stops writing data into the consumer (writing stopped). At a sixth timestep, the writing of the third buffer unit consumer is complete. In response, the consumer sends a third write done token to the producer); and
monitoring the read request in the first batch of the plurality of batches in each slice of the plurality of slices (Col. 16, lines 43-51, At a seventh timestep, a downstream consumer reads a buffer unit data from the consumer, i.e., the consumer is an upstream producer from the perspective of the downstream consumer and write the buffer unit data into the downstream consumer. This frees up space in the consumer equaling to a buffer unit data and therefore the consumer is ready to receive an additional buffer unit data from the producer. In response, the consumer sends a read ready token to the producer).
Chen reasonably teaches executing requests with local coordination in a slice (Col. 15). However, Chen does not explicitly teach a plurality of batches in each slice of the plurality of slices. Further, Chen reasonably teaches monitoring of requests through counters (Col. 15), however, does not explicitly teach local event tags set in the batches of the plurality of batches.
Chandra teaches coordination in a first batch of a plurality of batches (Col. 1, lines 46-50, According to the present invention, the foregoing and other objects are attained by providing corresponding sets of test instructions (Examiner notes: a batch of instructions) for a number of nodes in a computer network. The test instructions are partitioned into test modules; Col. 1, lines 64-67, The test modules have an ordered sequence. That is, there is a first module, a second module, a third module, and so on (Examiner notes: a workload partitioned into a plurality of ordered batches), in each set of modules on each one of the nodes) in each slice of a plurality of slice a plurality of batches in each slice of a plurality of slices (Col. 1, lines 51-56, When a node completes processing of one of its test modules, the node stores a test result corresponding to the test module. Meanwhile, another node processing it set of the test modules, stores test results for each one of the modules when it completes. This the two nodes process the test modules asynchronously (Examiner notes: the slice apparatus of Chen); Col. 2, lines 1-2, The nodes process the corresponding test modules in the same ordered sequence (Examiner notes: each node receives the same copy of the first batch)
setting a local event tag in the first batch of the plurality of batches in each of the plurality of slices (Col. 4, lines 54-67, If the comparison at 160 of node Nx and node Ny test module processing status indicates that node Nx has processed at least as many test modules as node Ny, then, at 178, node Nx sends the result of the highest order test module processed by node Nx to node Ny. Then, at 180 node Nx waits for node Ny to process additional test modules until node Ny catches up with node Nx. Once node Ny has processed a test module of the same order completed by node Nx, the at 185, node Nx gets the result of that test module from node Ny. Then, at 190, the results of the highest order test modules processed by both node Nx and node Ny are compared)
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Chandra with the teachings of Chen in order to provide a method that teaches execution of request with local coordination and setting local event tags in a distributed manner by means of batch processing of a workload to a computing slices. The motivation for applying Chandra teaching with Chen teaching is to provide a method that allows for the known benefits of concurrent execution coupled with proper execution management of operations distributed across a multi-processing hardware, such that coordinating execution though events can enable processing remedies, which is advantageous for ensuring proper functioning of a node to be confirmed with certainty prior to an operation (Chandra, Col. 2). Chen and Chandra are analogous art directed towards workload management in distributed environments. Therefore, it would have been obvious for one of ordinary skill in the art to combine Chandra with Chen to teach the claimed invention in order to provide a method that utilizes distributed computing with proper operation management across the distributed system.
However, Chen and Chandra do not explicitly teach coordinate distributed information processing locally using an asynchronous local output sync signal.
Temam teaches using an asynchronous local output sync signal (Col. 10-Col. 11, The sync flag controls manage concurrency between executions of different instruction types within compute tile 200. Each compute operation associated with each instruction type will be executed in strict order of issuance (i.e., First-In, First-Out).) to indicate completion of the execution of a workload batch (Col. 21, every write by a tile to a ring output will trigger an increment of the corresponding sync flag count. Controller 702 may examine the payload data to determine the number of data chunks or segments that comprise the payload (Examiner notes: workload batch). Controller 702 then monitors execution by the tile to ensure the expected number of data segments are forwarded and/or consumed by the tile before another tile executes in master mode) in each slice independently (Col. 4, Each computing unit of the hardware computing system is self-contained and can independently execute computations required) which is substantially similar to claim 1 and therefore rejected with similar rationale.
Examiner notes: It would be obvious for one of ordinary skill in the art to recognize that the apparatus of claim 1 is being substantially recited again for the method of claim 11.
With regard to claim 12, Chen teaches the method of claim 11, wherein the write request in the each slice of the plurality of slices is executed asynchronously (Col. 13, lines 20-23, The anti-congestion logic 232 is configured to configure each of the producers with a write credit counter that is initialized with one or more write credits; Col. 14, lines 35-40, Additionally, the producers, the compute nodes, and the consumers operate asynchronously and therefore use the anti-congestion flow control described herein to handle backpressure and avoid processing bottlenecks and buffer overflows between the producers and the consumers)
However, Chen does not explicitly teach that the asynchronous execution is with respect to other slices of the plurality of slices
Chandra teaches with respect to other slices of the plurality of slices (Col. 3, lines 51-56, However, by partitioning signatures analysis programs into test modules producing a signature for each of the modules, as provided in the present invention, and coordinating the processing of the test modules concurrently but in an asynchronous mode among a number of nodes)
It would have been obvious to one of ordinary skill in the art at the time the invention was filed to apply the teachings of Chandra with the teachings of Chen and Temam in order to provide a method that teaches asynchronous execution of an operation for a slice independent of the executions of the other slices in the system. The motivation for applying Chandra teaching with Chen and Temam teaching is to provide a method that allows for the known benefits of continuous processing and throughput of data in a distributed system, such that reduces the idleness of computing nodes and mitigates system stall due to failure of any single computing node (Chandra, Col. 3 - Col. 4). Chen, Temam, and Chandra are analogous art directed towards workload management in distributed environments. Therefore, it would have been obvious for one of ordinary skill in the art to combine Chandra with Chen and Temam to teach the claimed invention in order to provide asynchronous execution of operations of all computing slice in a system.
With regard to claim 13, Chen teaches the method of claim 12, further comprising the each slice of the plurality of slices suppling a local output coordination signal to indicate completion of the execution of a batch in the each slice of the plurality of slices (Col. 2, lines 54-63, In CGRAs and other processing systems that comprise a plurality of processing units that participate in a data processing operation, part of the data processing operation to be executed in one processing unit may need to be synchronized with other parts being executed in processing units distributed across the system. For example, several parts of the data processing operation may need to complete before a next part can safely begin. Thus, techniques for distributed control signals among elements of the processing system are required; Col. 15, lines 36-38, The read ready tokens and the write done tokens are pulse signals that emanate from the consumer and terminate at the producer)
With regard to claim 14, Chen teaches the method of claim 13, wherein monitoring the write request is performed until all previous write requests in a path are completed (FIG. 4, Dataflow graph 400 illustrating the path for General Matrix Multiplication (GeMM) operations where GeMM 0, 1, and 2 (402, 412, 422) perform write request to GeMM 4 (407) until completion; Col. 17, lines 27-29, FIG. 4 is an example of a dataflow graph 400 with compute nodes that asynchronously transmit data along data connections; Col. 17, lines 41-59, In the outer loop 410 each of the first three matrix multiplication nodes 402, 412, 422 receives a respective input (e.g., a respective tensor), executes a general matrix multiply (GeMM) operation on the respective input using a respective set of weights, and produces a respective output. The outputs from the first three matrix multiplication nodes 402, 412, 422 are piecewise processed by the innermost loop 409 over multiple iterations).
With regard to claim 15, Chen teaches the method of claim 13, wherein monitoring the write request is performed until a path has a verified receipt of the local event tag (Col. 13, lines 28-31, The write done token indicates to the particular producer that the writing of the buffer data unit into the corresponding consumer has completed; Col. 16, lines 6-10, The write done token ensures that at most K samples are processed through the compute node(s) between the producer and the consumer without receiving an acknowledgement from the consumer, where K is the initialization value of the write credit counter).
With regard to claim 16, Chen teaches the method of claim 13, wherein monitoring the write request is performed until the write request in the first batch of the plurality of batches is completed (Col. 13, lines 20-25 and lines 31-33, The anti-congestion logic 232 is configured to configure each of the producers with a write credit counter that is initialized with one or more write credits. The write credit counter is configured to decrement when the particular producer begins writing the buffer data unit into the corresponding consumer along the data connection … The particular producer stops writing data into the corresponding consumer when the write credit counter has zero write credits; Col. 16, lines 14-20, A zero write credit means that the consumer is still collecting the results from a previous sample and processing of the previous sample is not yet finished. The producer sends the next sample to the consumer only when the consumer finishes writing the result of the previous sample to a downstream consumer and is ready to start processing the next sample).
With regard to claim 17, Chen teaches the method of claim 13, wherein monitoring the write request is performed until two or more of the following occur: a) all previous write requests in a path are completed; b) the path has a verified receipt of the local event tag; c) the write request in the first batch of the plurality of batch is completed (Col. 20, lines 34-45, At action 1062, the method includes (i) configuring each of the producers with a write credit counter, (ii) initializing the write credit counter with one or more write credits, (iii) decrementing the write credit counter when the particular producer begins writing the buffer data unit into the corresponding consumer along the data connection (Examiner notes: Condition C, monitoring disabled after K batches), and (iv) incrementing the write credit counter when the particular producer receives from the corresponding consumer a write done token along the control connection. The write done token indicates to the particular producer that the writing of the buffer data unit into the corresponding consumer has completed (Examiner notes: Condition B, receipt of local event completion tag)
With regard to claim 18, Chen teaches the method of claim 11, wherein the read request in the each slice of the plurality of slices is executed asynchronously (Col. 12, lines 61-67, The anti-congestion logic 232 is configured to configure each of the producer with a ready-to-read credit counter, such that the ready-to-read credit counter of a particular producer is initialized with as many read credits as a buffer depth of a corresponding consumer that reads data from the particular producer; Col. 14, lines 35-40, Additionally, the producers, the compute nodes, and the consumers operate asynchronously and therefore use the anti-congestion flow control described herein to handle backpressure and avoid processing bottlenecks and buffer overflows between the producers and the consumers)
However, Chen does not explicitly teach that the asynchronous execution is with respect to other slices of the plurality of slices
Chandra teaches with respect to other slices of the plurality of slices (Col. 3, lines 51-56, However, by partitioning signatures analysis programs into test modules producing a signature for each of the modules, as provided in the present invention, and coordinating the processing of the test modules concurrently but in an asynchronous mode among a number of nodes) which is substantially similar to claim 12 and therefore reject with similar rationale.
Examiner notes: It would be obvious for one of ordinary skill in the art to recognize that the limitation of claim 12 is being substantially recited again as limitations for claim 17.
With regard to claim 19, Chen teaches the method of claim 18, wherein monitoring the read request is performed until all previous read requests in a path are completed (FIG. 4, Dataflow graph 400 illustrating the path for General Matrix Multiplication (GeMM) operations where GeMM (407) perform read request to aggregate results for the write request to GeMM 5 (408) until completion; Col. 17, lines 27-29, FIG. 4 is an example of a dataflow graph 400 with compute nodes that asynchronously transmit data along data connections; Col. 17, lines 54-59, The outputs from the multiple iterations are combined (e.g., concatenated) to generate an input for the matrix multiplication node 408).
With regard to claim 20, Chen teaches the method of claim 18, wherein monitoring the read request is performed until a path has a verified receipt of the local event tag (Col. 13, lines 6-9, The read ready token indicates to the particular producer that the corresponding consumer has freed a buffer data unit and is ready to receive an additional buffer data unit; Col. 16, lines 54-55, The consumer sends the read ready token to the producer when the consumer is ready to receive a new sample).
With regard to claim 21, Chen teaches the method of claim 18, wherein monitoring the read request is performed until the read request in the first batch of the plurality of batches is completed (Col. 12-Col. 13, lines 61-67 and lines 1-3 and lines 9-11, The anti-congestion logic 232 is configured to configure each of the producers with a ready-to-read counter, such that the ready-to-read credit counter of a particular producer is initialized with as many read credits as a buffer depth of a corresponding consumer that reads data from the particular producer. The ready-to-read credit counter is configured to decrement when the particular producer begins writing buffer data unit into the corresponding consumer along a data connection … The particular producer stops writing data into the corresponding consumer when the ready-to-read credit counter has zero read credits).
With regard to claim 22, Chen teaches the method of claim 18, wherein monitoring the read request is performed until two or more of the following occur: a) all previous read requests in a path are completed; b) the path has a verified receipt of the local event tag; c) the read request in the first batch of the plurality of batches is completed (Col. 20, lines 19-32, At action 1052, the method further includes (i) configuring each of the producers with a read-to-read credit counter, (ii) initializing the ready-to-read credit counter of a particular producer with as many read credits as a buffer depth of a corresponding consumer that reads data from the particular producer, (iii) decrementing the ready-to-read credit counter when the particular producer begins writing a buffer data unit into the corresponding consumer along a data connection, and (iv) incrementing the ready-to-read credit counter when the particular producer receives from the corresponding consumer a read ready token along a control connection. The read ready token indicates to the particular producer that the corresponding consumer has freed a buffer data unit and is ready to receive an additional buffer data unit (Examiner notes: Condition B, receipt of local event completion tag and Condition C, awaiting saturation of read buffer after processing first data batch)
With regard to claim 27, Chen teaches a non-transitory computer-readable medium storing computer executable code, operable on a device comprising at least one processor and at least one memory coupled to the at least one processor, wherein the at least one processor is configured to implement slice (Col. 14, lines 6-10, A data processing pipeline/operation comprises at least a producer, a compute node, and a consumer; Col. 14, lines 30-35, In the context of this application, a producer can be referred to as an upstream buffer or upstream memory node/unit, a compute node can be referred to as an intermediary compute node/unit or an intermediary processing node/unit, and a consumer can be referred to as a downstream buffer or downstream memory node/unit) coordination, the computer executable code comprising (Col. 20, lines 46-54, Other implementations of the method described in this section can include a non-transitory computer readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another implementation of the method described in this section include a system including memory and one or more processors operable to execute instructions, stored in the memory, to perform any of the methods described above):
instructions for causing a computer to execute a write request with local coordination (Col. 15, lines 10-12 and lines 22-28, The write credit counter is configured to decrement when the producer begins writing the buffer data unit into the consumer … The write credit counter is configured to increment when the producer receives from the consumer a write done token. The write done token indicates to the producer that the writing of the buffer data unit into the producer that the writing of the buffer data unit into the consumer has completed. The producer resumes writing data into the consumer when the producer receives the write done token from the consumer. The producer stops writing data into the consumer when the write credit counter has zero write credits) in a first batch of a plurality of batches in each slice of a plurality of slices (Col. 15, lines 48-49, At a first timestep, the producer begins writing a first buffer unit data into the consumer);
instructions for causing the computer to set a local event tracker (Col. 3, lines 56-58, The compiler is further configured to configure each of the producers (Examiner notes: each of the slice) with a ready-to-read credit counter; Col. 4, lines 16-18, In some embodiments, the compiler is further configured to configure each of the producers with a write credit counter that is initialized with one or more write credits; Col. 15, lines 36-38, As discussed above, the ready-to-read credit counter is initialized with three read credits and the write credit counter is initialized with two write credits) in the first batch of the plurality of batches in the each slice of the plurality of slices (Col. 15, lines 49-52, In response, the ready-to-read counter is decremented by one (read credit = 2) and the write credit counter is also decremented by one (write credit = 1) …;
instructions for causing the computer to monitor the write request in the first batch of the plurality of batches in the each slice of the plurality of slices (Col. 15, lines 53-64, At a second timestep, the producer begins writing a second buffer unit data into the consumer. In response, the ready-to-read credit counter is further decremented by one (read credit=1) and the write credit counter is further decremented by one (write credit = 0). The write credit counter expires and therefore the producer stops writing data into the consumer (writing stopped). At a third timestep, the writing of the first buffer unit data into the consumer is complete. In response, the consumer sends a first write done token to the produce. In response, the write credit counter is incremented by one and therefore reactivated (write credit = 1) (Examiner notes: Where slice request execution is monitored by the credits available);
instructions for causing the computer to execute a read request with the local coordination (Col. 15, lines 7-11 and lines 12-22, The ready-to-read counter is configured to decrement when the producer begins writing a buffer data unit into the consumer. The size of the buffer data unit is s (e.g., 16 bytes, 64 bytes, 512 bytes … The ready-to-read credit counter is configured to increment when the producer receives from the consumer a read ready token. The read ready token indicates to the producer that the consumer has freed a buffer data unit and is ready to receive an additional buffer data unit. The producer stops writing data into the consumer when the ready-to-read credit counter has zero read credits. The producer resumes writing data into the consumer when the producer receives read ready token from the consumer) in the first batch of the plurality of batches in the each slice of the plurality of slices after monitoring the write request is completed (Col. 16, lines 31-40, At the fifth timestep, the producer begins writing a third buffer unit data into the consumer. In response, the ready-to-read credit counter is further decremented by one (read credit = 0) and the write credit counter is decremented by one (write credit = 1). The ready-to-read credit counter expires and therefore the producer stops writing data into the consumer (writing stopped). At a sixth timestep, the writing of the third buffer unit consumer is complete. In response, the consumer sends a third write done token to the producer); and
instructions for causing the computer to monitor the read request in the first batch of the plurality of batches in the each slice of the plurality of slices (Col. 16, lines 43-51, At a seventh timestep, a downstream consumer reads a buffer unit data from the consumer, i.e., the consumer is an upstream producer from the perspective of the downstream consumer and write the buffer unit data into the downstream consumer. This frees up space in the consumer equaling to a buffer unit data and therefore the consumer is ready to receive an additional buffer unit data from the producer. In response, the consumer sends a read ready token to the producer).
Chen reasonably teaches executing requests with local coordination in a slice (Col. 15). However, Chen does not explicitly teach a plurality of batches in each slice of the plurality of slices. Further, Chen reasonably teaches monitoring of requests through counters (Col. 15), however, does not explicitly teach local event tags set in the batches of the plurality of batches.
Chandra teaches coordination in a first batch of a plurality of batches (Col. 1, lines 46-50, According to the present invention, the foregoing and other objects are attained by providing corresponding sets of test instructions (Examiner notes: a batch of instructions) for a number of nodes in a computer network. The test instructions are partitioned into test modules; Col. 1, lines 64-67, The test modules have an ordered sequence. That is, there is a first module, a second module, a third module, and so on (Examiner notes: a workload partitioned into a plurality of ordered batches), in each set of modules on each one of the nodes) in each slice of a plurality of slice a plurality of batches in each slice of a plurality of slices (Col. 1, lines 51-56, When a node completes processing of one of its test modules, the node stores a test result corresponding to the test module. Meanwhile, another node processing it set of the test modules, stores test results for each one of the modules when it completes. This the two nodes process the test modules asynchronously (Examiner notes: the slice apparatus of Chen); Col. 2, lines 1-2, The nodes process the corresponding test modules in the same ordered sequence (Examiner notes: each node receives the same copy of the first batch)
setting a local event tag in the first batch of the plurality of batches in each of the plurality of slices (Col. 4, lines 54-67, If the comparison at 160 of node Nx and node Ny test module processing status indicates that node Nx has processed at least as many test modules as node Ny, then, at 178, node Nx sends the result of the highest order test module processed by node Nx to node Ny. Then, at 180 node Nx waits for node Ny to process additional test modules until node Ny catches up with node Nx. Once node Ny has processed a test module of the same order completed by node Nx, the at 185, node Nx gets the result of that test module from node Ny. Then, at 190, the results of the highest order test modules processed by both node Nx and node Ny are compared) which is substantially similar to claim 11 and therefore rejected with similar rationale.
Examiner notes: It would be obvious for one of ordinary skill in the art to recognize that the method of claim 11 is being substantially recited again as limitations for the non-transitory computer-readable medium of claim 27.
However, Chen and Chandra do not explicitly teach coordinate distributed information processing locally using an asynchronous local output sync signal.
Temam teaches using an asynchronous local output sync signal (Col. 10-Col. 11, The sync flag controls manage concurrency between executions of different instruction types within compute tile 200. Each compute operation associated with each instruction type will be executed in strict order of issuance (i.e., First-In, First-Out).) to indicate completion of the execution of a workload batch (Col. 21, every write by a tile to a ring output will trigger an increment of the corresponding sync flag count. Controller 702 may examine the payload data to determine the number of data chunks or segments that comprise the payload (Examiner notes: workload batch). Controller 702 then monitors execution by the tile to ensure the expected number of data segments are forwarded and/or consumed by the tile before another tile executes in master mode) in each slice independently (Col. 4, Each computing unit of the hardware computing system is self-contained and can independently execute computations required) which is substantially similar to claim 11 and therefore rejected with similar rationale.
Examiner notes: It would be obvious for one of ordinary skill in the art to recognize that the method of claim 11 is being substantially recited again for the non-transitory computer-readable medium of claim 27.
With regard to claim 28, Chen teaches the non-transitory computer-readable medium of claim 27, further comprising instructions for causing the computer to execute the write request asynchronously with respect to other slices of the plurality of slices (Col. 13, lines 20-23, The anti-congestion logic 232 is configured to configure each of the producers with a write credit counter that is initialized with one or more write credits; Col. 14, lines 35-40, Additionally, the producers, the compute nodes, and the consumers operate asynchronously and therefore use the anti-congestion flow control described herein to handle backpressure and avoid processing bottlenecks and buffer overflows between the producers and the consumers) and to execute the read request asynchronously (Col. 12, lines 61-67, The anti-congestion logic 232 is configured to configure each of the producer with a ready-to-read credit counter, such that the ready-to-read credit counter of a particular producer is initialized with as many read credits as a buffer depth of a corresponding consumer that reads data from the particular producer; Col. 14, lines 35-40, Additionally, the producers, the compute nodes, and the consumers operate asynchronously and therefore use the anti-congestion flow control described herein to handle backpressure and avoid processing bottlenecks and buffer overflows between the producers and the consumers)
However, Chen does not explicitly teach that the asynchronous execution is with respect to other slices of the plurality of slices
Chandra teaches with respect to other slices of the plurality of slices (Col. 3, lines 51-56, However, by partitioning signatures analysis programs into test modules producing a signature for each of the modules, as provided in the present invention, and coordinating the processing of the test modules concurrently but in an asynchronous mode among a number of nodes) which is substantially similar to claim 12 and therefore reject with similar rationale.
Examiner notes: It would be obvious for one of ordinary skill in the art to recognize that the method of claim 12 is being substantially recited again as limitations for the non-transitory computer-readable medium of claim 28.
With regard to claim 29, Chen teaches the non-transitory computer-readable medium of claim 28, further comprising instructions for causing the computer to monitor the write request until one or more of the following occur: a) all previous write requests in a path are completed; b) the path has a verified receipt of the local event tag; c) the write request in the first batch of the plurality of batches is completed (Col. 20, lines 34-45, At action 1062, the method includes (i) configuring each of the producers with a write credit counter, (ii) initializing the write credit counter with one or more write credits, (iii) decrementing the write credit counter when the particular producer begins writing the buffer data unit into the corresponding consumer along the data connection, and (iv) incrementing the write credit counter when the particular producer receives from the corresponding consumer a write done token along the control connection. The write done token indicates to the particular producer that the writing of the buffer data unit into the corresponding consumer has completed (Examiner notes: Condition B, receipt of local event completion tag and Condition C completion of first data batch).
With regard to claim 30, Chen teaches the non-transitory computer-readable medium of claim 28, further comprising instructions for causing the computer to monitor the read request until one or more of the following occur: a) all previous read requests in a path are completed; b) the path has a verified receipt of the local event tag; c) the read request in the first batch of the plurality of batches is completed (Col. 20, lines 19-32, At action 1052, the method further includes (i) configuring each of the producers with a read-to-read credit counter, (ii) initializing the ready-to-read credit counter of a particular producer with as many read credits as a buffer depth of a corresponding consumer that reads data from the particular producer, (iii) decrementing the ready-to-read credit counter when the particular producer begins writing a buffer data unit into the corresponding consumer along a data connection, and (iv) incrementing the ready-to-read credit counter when the particular producer receives from the corresponding consumer a read ready token along a control connection. The read ready token indicates to the particular producer that the corresponding consumer has freed a buffer data unit and is ready to receive an additional buffer data unit (Examiner notes: Condition B, receipt of local event completion tag and Condition C, awaiting saturation of read buffer after processing first data batch).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 2020/0089499 A1
teaches
Abstract, A processing system comprising multiple tiles and interconnect between the tiles. The interconnect is used to communicate between a group of some or all of the tiles according to a bulk synchronous parallel scheme, whereby each tile in the group performs an on-tile compute phase followed by an inter-tile exchange phase with the exchange phase being held back until all tiles in the group have completed the compute phase. Each tile in the group has a local exit state upon completion of the compute phase. The instruction set comprises a synchronization instruction for execution by each tile upon completion of its compute phase to signal a sync request to logic in the interconnect.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IVAN A CASTANEDA whose telephone number is (571)272-0465. The examiner can normally be reached Monday-Friday 9:30AM-5:30PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee Li can be reached at (571) 272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/I.A.C./Examiner, Art Unit 2195
/Aimee Li/Supervisory Patent Examiner, Art Unit 2195