DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1, 9, 13-15, 31-36, 38-50, 52-57, and 59-60 are pending in this office action and presented for examination. Claims 9, 38, and 40 are newly amended, and claim 51 is newly cancelled, by the response received January 5, 2026.
Drawings
The replacement drawings received January 5, 2026, are objected to because:
All drawings must be made by a process which will give them satisfactory reproduction characteristics. Every line, number, and letter must be durable, clean, black (except for color drawings), sufficiently dense and dark, and uniformly thick and well-defined. The weight of all lines and letters must be heavy enough to permit adequate reproduction. This requirement applies to all lines however fine, to shading, and to lines representing cut surfaces in sectional views. However, the replacement drawings received January 5, 2026, do not meet this requirement — see, for example, the array of white dots that causes the lines, text, and numbers to appear fuzzy and blurry. For example, see the interior of the “FIG. 3” label; the arrowheads coupled to the bottom of element 124 in FIG. 3; and the black speckles around instances of reference character 310 in FIG. 3. (For example, compare, in the file wrapper, the replacement drawings received January 5, 2026, with the replacement drawings received July 26, 2024.) At least some of these issues may be caused by dithering being applied when a conversion from greyscale to black has taken place; if so, Examiner recommends ensuring that any drawings to be filed do not contain any grey elements.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 9, 13-15, and 31-35 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt (US 20090300327 A1) in view of Lumsdaine et al. (Lumsdaine) (US 20070198621 A1) in view of Curry et al. (Curry) (US 20100315428 A1) in view of Matick et al. (Matick) (US 20050114606 A1) in view of Mohamed (US 6366998 B1).
Consider claim 1, Omtzigt discloses a computing device ([0027], line 2, data flow computer system) comprising: a memory for storing data and a domain flow program ([0027], lines 4-5, memory 110 that contains the data and the program to execute; [0027], lines 13-14, programming information; [0022], lines 1-2, an execution engine executes single assignment programs with affine dependencies; see the remarks dated 1/20/2012 in the parent application 12/467485, beginning on page 10, for a greater explanation of how Omtzigt teaches domain flow); a controller for requesting the data and the domain flow program from the memory ([0027], lines 6-10, execution starts by the controller 120 requesting a program from the memory 110. The controller 120 presents a read request via bus 121 to the memory controller 130 which translates the read request to a memory request and returns the data to the controller 120. This data contains the program instructions to execute a single assignment program); and a processor fabric for processing the data and the domain flow program ([0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams) via a plurality of processing elements ([0029], line 3, processing elements (PE) 310).
However, Omtzigt does not entail that, for sparse matrices, index structures are used to minimize memory bandwidth. Omtzigt also does not disclose a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory. Omtzigt also does not disclose the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.
On the other hand, Lumsdaine discloses that, for sparse matrices, index structures are used to minimize memory bandwidth (for example, see [0028], lines 3-10, the term "sparse matrix" refers to a matrix where the vast majority (usually at least 99% for large matrices) of the element values are zero. The number of nonzero elements in a given matrix is denoted by NNZ. Matrices with this property may be stored much more efficiently in memory if only the nonzero data values are stored, as well as some index data to identify where those data values fit in the matrix).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Lumsdaine with Omtzigt in order to increase memory efficiency (Lumsdaine, [0028], lines 3-10).
However, the combination thus far does not entail a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory. The combination thus far also does not disclose the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.
On the other hand, Curry discloses a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from memory ([0015], lines 8-10, functions can be chained together to process images without intermediate trips to external memory; [0017], lines 1-2, SIMD processors execute functions known as kernels).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Curry with the combination of Omtzigt and Lumsdaine in order to increase system performance.
However, the combination thus far does not entail a cache for recalling a previous kernel. The combination thus far also does not disclose the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.
On the other hand, Matick discloses a cache for recalling a previous kernel (for example, [0046], lines 9-10, kernels in the cache).
Matick’s teaching increases speed (Matick, [0002], lines 1-2, caches are used to speed up accesses to recently accessed information).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Matick with the combination of Omtzigt, Lumsdaine, and Curry in order to increase speed.
However, the combination thus far does not disclose the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.
On the other hand, Mohamed discloses a density such that parallel computations are expressed in 100 bytes or fewer (FIG. 6, which shows a vector MAC instruction being expressed in 2 bytes, and an overall instruction packet being expressed in 32 bytes).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Mohamed with the combination of Omtzigt, Lumsdaine, Curry, and Matick in order to save space relative to implementations which necessitate a greater number of bytes to express parallel computations.
Consider claim 9, the overall combination entails the computing device of claim 1 (see above) wherein the controller is configured for presenting a read request to a memory controller which translates the read request to a memory request and returns the data to the controller (Omtzigt, [0027], lines 7-10, the controller 120 presents a read request via bus 121 to the memory controller 130 which translates the read request to a memory request and returns the data to the controller 120).
Consider claim 13, the overall combination entails the computing device of claim 1 (see above) further comprising a crossbar configured for routing the data to rows or columns in the processor fabric (Omtzigt, [0027], lines 27-29, the crossbar 150 routes the data streams to the appropriate rows or columns in the processor fabric 160).
Consider claim 14, the overall combination entails the computing device of claim 13 (see above) wherein the processor fabric is configured for producing output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams).
Consider claim 15, the overall combination entails the computing device of claim 14 wherein the output data is written to the memory by traversing through the crossbar and whereupon the output data is presented to a memory controller which writes the output data into the memory (Omtzigt, [0027], lines 31-36, these output data streams are written back to memory 110 by traversing the crossbar 150 to the streamers 140 that associate memory addresses to the data streams, and then present them to the memory controller 130, which will write the data streams into memory 110).
Consider claim 31, the overall combination entails the computing device of claim 1 (see above) wherein the processor fabric is programmable (Omtzigt, [0032], lines 24-26, after the controller 120 is done programming the processor fabric 160, execution is able to commence).
Consider claim 32, the overall combination entails the computing device of claim 1 (see above) wherein the processor fabric receives the data, executes instructions on the data, and produces output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams).
Consider claim 33, the overall combination entails the computing device of claim 1 (see above) wherein the data is processed under control of the domain flow program which is a single assignment program (Omtzigt, [0033], lines 10-11, execution of a single assignment program on the fabric).
Consider claim 34, the overall combination entails the computing device of claim 33 (see above) wherein the plurality of processing elements are configured as a network of processing elements (Omtzigt, [0029], lines 16-17, processing element routing network 320) configured for executing the single assignment program (Omtzigt, [0033], lines 10-11, execution of a single assignment program on the fabric).
Consider claim 35, the overall combination entails the computing device of claim 1 (see above) wherein the plurality of processing elements are configured as a matrix of processing elements (Omtzigt, [0029], line 3, processor array).
Claim(s) 36 and 38-46 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt (US 20090300327 A1) in view of Lumsdaine et al. (Lumsdaine) (US 20070198621 A1) in view of Curry et al. (Curry) (US 20100315428 A1) in view of Matick et al. (Matick) (US 20050114606 A1) in view of Glasco (US 20030182514 A1).
Consider claim 36, Omtzigt discloses a method comprising: storing, in a memory of a device, data and a domain flow program ([0027], lines 4-5, memory 110 that contains the data and the program to execute; [0027], line 2, data flow computer system; [0027], lines 13-14, programming information; [0022], lines 1-2, an execution engine executes single assignment programs with affine dependencies; see the remarks dated 1/20/2012 in the parent application 12/467485, beginning on page 10, for a greater explanation of how Omtzigt teaches domain flow); requesting, with a controller, the data and the domain flow program from the memory ([0027], lines 6-10, execution starts by the controller 120 requesting a program from the memory 110. The controller 120 presents a read request via bus 121 to the memory controller 130 which translates the read request to a memory request and returns the data to the controller 120. This data contains the program instructions to execute a single assignment program); and processing, with a processor fabric, the data and the domain flow program ([0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams) via a plurality of elements ([0029], line 3, processing elements (PE) 310).
However, Omtzigt does not entail that for sparse matrices, index structures are used to minimize memory bandwidth. Omtzigt also does not disclose chaining a plurality of parallel kernels to avoid serializing intermediate data to and from the memory, wherein a cache is used for recalling a previous kernel of the plurality of parallel kernels. Omtzigt also does not disclose a plurality of processor fabrics configured to communicate with each other.
On the other hand, Lumsdaine discloses that for sparse matrices, index structures are used to minimize memory bandwidth (for example, see [0028], lines 3-10, the term "sparse matrix" refers to a matrix where the vast majority (usually at least 99% for large matrices) of the element values are zero. The number of nonzero elements in a given matrix is denoted by NNZ. Matrices with this property may be stored much more efficiently in memory if only the nonzero data values are stored, as well as some index data to identify where those data values fit in the matrix).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Lumsdaine with Omtzigt in order to increase memory efficiency (Lumsdaine, [0028], lines 3-10).
However, the combination thus far does not entail chaining a plurality of parallel kernels to avoid serializing intermediate data to and from the memory, wherein a cache is used for recalling a previous kernel of the plurality of parallel kernels. The combination thus far also does not disclose a plurality of processor fabrics configured to communicate with each other.
On the other hand, Curry discloses a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from memory ([0015], lines 8-10, functions can be chained together to process images without intermediate trips to external memory; [0017], lines 1-2, SIMD processors execute functions known as kernels).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Curry with the combination of Omtzigt and Lumsdaine in order to increase system performance.
However, the combination thus far does not entail that a cache is used for recalling a previous kernel. The combination thus far also does not disclose a plurality of processor fabrics configured to communicate with each other.
On the other hand, Matick discloses that a cache is used for recalling a previous kernel (for example, [0046], lines 9-10, kernels in the cache).
Matick’s teaching increases speed (Matick, [0002], lines 1-2, caches are used to speed up accesses to recently accessed information).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Matick with the combination of Omtzigt, Lumsdaine, and Curry in order to increase speed.
However, the combination thus far does not disclose a plurality of processor fabrics configured to communicate with each other.
On the other hand, Glasco discloses a plurality of processor fabrics configured to communicate with each other ([0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture).
Glasco’s teaching reduces the number of links used and further modularizes a multiprocessor system (Glasco, [0033], lines 15-18, in order to reduce the number of links used and to further modularize a multiprocessor system using a point-to-point architecture, multiple clusters are used). In addition, a plurality of processor fabrics increases system performance relative to one of that processor fabric.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Glasco with the combination of Omtzigt, Lumsdaine, Curry, and Matick in order to reduce the number of links used and to further modularize a multiprocessor system and to increase system performance.
Consider claim 38, the overall combination entails the method of claim 36 (see above) wherein the controller is configured for presenting a read request to a memory controller which translates the read request to a memory request and returns the data to the controller (Omtzigt, [0027], lines 7-10, the controller 120 presents a read request via bus 121 to the memory controller 130 which translates the read request to a memory request and returns the data to the controller 120).
Consider claim 39, the overall combination entails the method of claim 36 (see above) further comprising routing, with a crossbar, the data to rows or columns in the plurality of processor fabrics (Omtzigt, [0027], lines 27-29, the crossbar 150 routes the data streams to the appropriate rows or columns in the processor fabric 160; Glasco, [0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture).
Consider claim 40, the overall combination entails the method of claim 39 (see above) wherein the plurality of processor fabrics are configured for producing output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams; Glasco, [0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture).
Consider claim 41, the overall combination entails the method of claim 40 wherein the output data is written to the memory by traversing through the crossbar and whereupon the output data is presented to a memory controller which writes the output data into the memory (Omtzigt, [0027], lines 31-36, these output data streams are written back to memory 110 by traversing the crossbar 150 to the streamers 140 that associate memory addresses to the data streams, and then present them to the memory controller 130, which will write the data streams into memory 110).
Consider claim 42, the overall combination entails the method of claim 36 (see above) wherein the plurality of processor fabrics are programmable (Omtzigt, [0032], lines 24-26, after the controller 120 is done programming the processor fabric 160, execution is able to commence; Glasco, [0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture).
Consider claim 43, the overall combination entails the method of claim 36 (see above) wherein the plurality of processor fabrics receive the data, execute instructions on the data, and produce output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams; Glasco, [0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture).
Consider claim 44, the overall combination entails the method of claim 36 (see above) wherein the data is processed under control of the domain flow program which is a single assignment program (Omtzigt, [0033], lines 10-11, execution of a single assignment program on the fabric).
Consider claim 45, the overall combination entails the method of claim 36 (see above) wherein the plurality of processing elements are configured as a matrix of processing elements (Omtzigt, [0029], line 3, processor array).
Consider claim 46, the overall combination entails the method of claim 36 (see above) wherein the plurality of elements are configured as a network of processing elements (Omtzigt, [0029], lines 16-17, processing element routing network 320) configured for executing the single assignment program (Omtzigt, [0033], lines 10-11, execution of a single assignment program on the fabric).
Claim(s) 47 and 52-55 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt (US 20090300327 A1) in view of Lumsdaine et al. (Lumsdaine) (US 20070198621 A1) in view of Curry et al. (Curry) (US 20100315428 A1) in view of Matick et al. (Matick) (US 20050114606 A1) in view of Martin et al. (Martin) (US 20020156995 A1).
Consider claim 47, Omtzigt discloses a processor fabric for processing ([0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams) a domain flow program ([0027], lines 4-5, memory 110 that contains the data and the program to execute) comprising: a first set of processing elements ([0029], line 3, processing elements (PE) 310); and a second set of processing elements which communicate with the first set of processing elements ([0029], line 3, processing elements (PE) 310; [0029], lines 16-17, processing element routing network 320), wherein the first set of processing elements and the second set of processing elements are configured for processing data and the domain flow program ([0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams).
However, Omtzigt does not entail that for sparse matrices, index structures are used to minimize memory bandwidth. Omtzigt also does not disclose a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained. Omtzigt also does not disclose the first set of processing elements is implemented as an asynchronous execution pipeline.
On the other hand, Lumsdaine discloses that for sparse matrices, index structures are used to minimize memory bandwidth (for example, see [0028], lines 3-10, the term "sparse matrix" refers to a matrix where the vast majority (usually at least 99% for large matrices) of the element values are zero. The number of nonzero elements in a given matrix is denoted by NNZ. Matrices with this property may be stored much more efficiently in memory if only the nonzero data values are stored, as well as some index data to identify where those data values fit in the matrix).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Lumsdaine with Omtzigt in order to increase memory efficiency (Lumsdaine, [0028], lines 3-10).
However, the combination thus far does not entail a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained. The combination thus far also does not disclose the first set of processing elements is implemented as an asynchronous execution pipeline.
On the other hand, Curry discloses a plurality of parallel kernels, wherein the plurality of parallel kernels are chained ([0015], lines 8-10, functions can be chained together to process images without intermediate trips to external memory; [0017], lines 1-2, SIMD processors execute functions known as kernels).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Curry with the combination of Omtzigt and Lumsdaine in order to increase system performance.
However, the combination thus far does not entail a cache for recalling a previous kernel. The combination thus far also does not disclose the first set of processing elements is implemented as an asynchronous execution pipeline.
On the other hand, Matick discloses a cache for recalling a previous kernel (for example, [0046], lines 9-10, kernels in the cache).
Matick’s teaching increases speed (Matick, [0002], lines 1-2, caches are used to speed up accesses to recently accessed information).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Matick with the combination of Omtzigt, Lumsdaine, and Curry in order to increase speed.
However, the combination thus far does not disclose the first set of processing elements is implemented as an asynchronous execution pipeline.
On the other hand, Martin discloses a first set of processing elements is implemented as an asynchronous execution pipeline ([0047], line 2, asynchronous execution pipeline).
Martin’s teaching increases speed (Martin, [0047], lines 13-16).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Martin with the combination of Omtzigt, Lumsdaine, Curry, and Matick in order to increase speed.
Consider claim 52, the overall combination entails the processor fabric of claim 47 (see above) wherein the first set of processing elements and the second set of processing elements are further configured for receiving the data and producing output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams).
Consider claim 53, the overall combination entails the processor fabric of claim 52 wherein the output data is written to a memory by traversing through a crossbar to a memory controller which writes the output data into the memory (Omtzigt, [0027], lines 31-36, these output data streams are written back to memory 110 by traversing the crossbar 150 to the streamers 140 that associate memory addresses to the data streams, and then present them to the memory controller 130, which will write the data streams into memory 110).
Consider claim 54, the overall combination entails the processor fabric of claim 47 (see above) wherein the first set of processing elements and the second set of processing elements recognize a spatial tag of a computational event and take action under control of the domain flow program (Omtzigt, [0029], lines 21-24, the PEs 310 recognize the spatial tag called the signature of a computational event and take action under control of the single assignment program installed in their program store by controller 120).
Consider claim 55, the overall combination entails the processor fabric of claim 47 (see above) wherein the domain flow program evolves as multi-dimensional data match up with the first set of processing elements and the second set of processing elements and produce new multi-dimensional data (Omtzigt, [0029], lines 28-32, the overall computation represented by the single assignment program evolves as the multi-dimensional data streams match up within the processing elements 310 and producing potentially new multi-dimensional data streams).
Claim(s) 48 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Jensen et al. (Jensen) (US 20070094482 A1).
Consider claim 48, the combination thus far entails the processor fabric of claim 47 (see above), but does not entail that the first set of processing elements and the second set of processing elements process an instruction set architecture configured for a specific class of algorithms.
On the other hand, Jensen discloses processing elements processing an instruction set architecture configured for a specific class of algorithms ([0007], lines 14-19, accordingly, a system designer can specify a specific ISA for encoding a specific set of program functions (e.g., signal processing algorithms) and select other ISAs for encoding other types of program functions (e.g., operating system functions, I/O functions, general purpose functions)).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Jensen with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin in order to provide support for a specific class of algorithms, such as signal processing algorithms, thereby increasing device capability.
Claim(s) 49 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen as applied to claim 48 above, and further in view of Zakiya (US 20020032551 A1).
Consider claim 49, the combination thus far entails the processor fabric of claim 48 (see above), but does not entail the specific class of algorithms comprise hashing algorithms.
On the other hand, Zakiya discloses hashing algorithms ([0005], lines 1-2, The current most widely used hash algorithms are MD5 and the Secure Hash Algorithm (SHA-1)).
Zakiya’s teaching is useful for computing a unique condensed representation of a message or a data file, which is useful for security (Zakiya, [0003]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Zakiya with the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen in view of its usefulness for computing a unique condensed representation of a message or a data file, which is useful for security.
Claim(s) 50 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen as applied to claim 48 above, and further in view of Benkelman (US 6694064 A1).
Consider claim 50, the combination thus far entails the processor fabric of claim 48 (see above), but does not entail the specific class of algorithms are utilized for interpolations and resampling.
On the other hand, Benkelman discloses optimizations for interpolations and resampling (col. 5, lines 30-32, two common resampling algorithms are referred to as "bilinear interpolation" and "nearest neighbor.").
Benkelman’s teaching is fundamental in the rectification or transformation process (Benkelman, col. 5, lines 23-24).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Benkelman with the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen in view of its usefulness in the rectification or transformation process.
Claim(s) 56 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Eggers et al. (Eggers) (US 20060179429 A1).
Consider claim 56, the combination thus far entails the processor fabric of claim 47 (see above), wherein each processing element of the second set of processing elements includes a set of a content addressable memory (FIG. 6, Tag CAM 640), one or more functional units ([0036], line 28, functional units), and a router configured to generate affine routing vectors (Fig. 4, packet router 425; [0032], line 37, routing vector; [0032], lines 20-21, routing vector. This information defines some affine recurrence equation). However, the combination thus far does not disclose an instruction scheduling and dispatch queue.
On the other hand, Eggers discloses an instruction scheduling and dispatch queue ([0172], lines 1-2, the PE selects an instruction from the scheduling queue).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Eggers with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin, as this modification merely entails combining prior art elements (the prior art elements of the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin, and Eggers’s teaching of an instruction scheduling and dispatch queue) according to known methods (Examiner submits that use of an instruction scheduling and dispatch queue was very well-known to one of ordinary skill in the art before the effective filing date of the claimed invention) to yield predictable results (the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin, further modified to incorporate an instruction scheduling and dispatch queue), which is an example of a rationale that may support a conclusion of obviousness as per MPEP 2143. Alternatively, Examiner submits that an instruction scheduling and dispatch queue facilitates efficient instruction scheduling and dispatching relative to not using a queue.
Claim(s) 57 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Jensen et al. (Jensen) (US 20070094482 A1) and Bhanot et al. (Bhanot) (US 20040073590 A1).
Consider claim 57, the combination thus far entails the processor fabric of claim 47 (see above). However, the combination thus far does not entail that a plurality of instruction sets drive a plurality of global operators.
On the other hand, Jensen discloses a plurality of instruction sets ([0007], lines 14-19, accordingly, a system designer can specify a specific ISA for encoding a specific set of program functions (e.g., signal processing algorithms) and select other ISAs for encoding other types of program functions (e.g., operating system functions, I/O functions, general purpose functions))
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Jensen with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin in order to increase system capability by providing support for a wide range of program functions.
However, the combination thus far does not entail that the plurality of instruction sets drive a plurality of global operators.
On the other hand, Bhanot discloses global operators ([0003], line 4, global operation).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Bhanot with the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen, as this modification merely entails combining prior art elements (the prior art elements of the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen, and Bhanot’s teaching of a global operator) according to known methods (Examiner submits that the general concept of a global operator was well-known to one of ordinary skill in the art before the effective filing date of the claimed invention, and Bhanot also discloses execution of a global operation in the environment of a network of computing nodes) to yield predictable results (the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen, further modified to support global operations), which is an example of a rationale that may support a conclusion of obviousness as per MPEP 2143. Alternatively, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Bhanot with the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen in order to improve efficiency (Bhanot, [0002], lines 7-8).
Claim(s) 59 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Glasco (US 20030182514 A1).
Consider claim 59, the combination thus far entails the processor fabric of claim 47 (see above), but does not entail the processor fabric is implemented on a chip with one or more additional processor fabrics, wherein the processor fabric is configured to communicate with the one or more additional processor fabrics.
On the other hand, Glasco discloses a processor fabric is implemented on a chip with one or more additional processor fabrics, wherein the processor fabric is configured to communicate with the one or more additional processor fabrics ([0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture).
Glasco’s teaching reduces the number of links used and further modularizes a multiprocessor system (Glasco, [0033], lines 15-18, in order to reduce the number of links used and to further modularize a multiprocessor system using a point-to-point architecture, multiple clusters are used). In addition, a plurality of processor fabrics increases system performance relative to one of that processor fabric.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Glasco with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin in order to reduce the number of links used and to further modularize a multiprocessor system and to increase system performance.
Claim(s) 60 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Chan et al. (Chan) (US 20060282650 A1).
Consider claim 60, the combination thus far entails the processor fabric of claim 47 (see above), but does not entail each processing element of the second set of processing elements comprises a Single Instruction Multiple Data (SIMD) unit configured to process a plurality of operations per instruction.
On the other hand, Chan discloses processing elements each comprise a Single Instruction Multiple Data (SIMD) unit configured to process a plurality of operations per instruction (FIG. 1B, MFU 222; [0023], lines 10-17, the media functional units 222 are single-instruction-multiple-data (SIMD) media functional units. Each media functional unit 222 is capable of processing parallel 16-bit components, in addition to 32-bit operands. Various parallel 16-bit operations supply the single-instruction-multiple-data capability for processor 100 including add, multiply-add, shift, compare, and the like).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Chan with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin in order to increase performance via parallelism.
Response to Arguments
Applicant on page 8 argues: “Within the Office Action, Figures 3-5 have been objected to. By the above amendments, the drawings have been clarified. Applicants have ensured that the drawings are free from arrays of white dots. No new matter has been added. The objection should be withdrawn.”
However, the replacement drawings received January 5, 2026, do not appear to overcome the aforementioned objection to the drawings. Examiner generally notes that dithering showing up only in the drawings in the file wrapper does not preclude the root cause from being the pre-filed drawings. Examiner recommends ensuring that any drawings to be filed do not contain any grey elements, and comparing the hexadecimal or RGB code of the elements of the replacement drawings dated January 5, 2026, with the replacement drawings dated July 26, 2024 (which have satisfactory reproduction characteristics). (Examiner notes that there may be other reasons for the unsatisfactory reproduction characteristics, but dithering appears to be the most common. Examiner also notes that the drawings dated January 5, 2026, appears to have less satisfactory reproduction characteristics than the previous drawings dated September 15, 2025; for example, the drawings dated January 5, 2026 appear to have more black speckles; see, for example, the area at the bottom of element 124 in FIG. 3, where arrows converge.)
Applicant on page 8 argues: “Within the Office Action, Claims 40 and 41 have been objected to. By the above amendments, Claim 40 has been clarified. No new matter has been added by this clarifying amendment. The objection should be withdrawn.”
In view of the aforementioned amendment, the previously presented objections are withdrawn.
Applicant on page 8 argues: “Within the Office Action, Claims 9, 38 and 51 have been rejected under 35 U.S.C. § 112, second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention. Applicants respectfully disagree. However, Claims 9 and 38 have been clarified by the above amendments. Additionally, Claim 51 has been canceled by the above amendments. Therefore, the rejection should be withdrawn.”
In view of the aforementioned amendments and claim cancellation, the previously presented indefinite rejections are withdrawn.
Applicant on page 9 argues: “Applicants note that U.S. Patent Application Publication No. 2009/0300327 by Omtzigt is not prior art to the presently claimed invention as the present application claims priority to the '327 application in the preliminary amendment and the application data sheet. The U.S. Patent Application Publication No. 2009/0300327 by Omtzigt is Application No. 12/467,485, which is listed as a Parent Application on the Domestic Priority Section of the Filing Receipt for the Present Application. Therefore, the rejection should be withdrawn.”
However, as noted in the preliminary amendment and the application data sheet, parent application 14/185,841 (which lies between the aforementioned 12/467,485 application and the instant application in the priority chain) is a continuation-in-part of the aforementioned 12/467,485 application. As the claims of the instant application do not appear to be supported by the specification and claims of the aforementioned application 12/467,485, the claims of the instant application have an effective filing date equal to the filing date of parent application 14/185,841 (February 20, 2014). Therefore, Omtzigt (US 2009/0300327 A1) (application 12/467,485), published on December 3, 2009, appears to be prior art to the presently claimed invention.
Applicant on page 9 argues: “Additionally, Applicants respectfully disagree that Mohamed teaches the limitation: wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer. Although Figure 6 of Mohamed mentions an instruction packet having 32-bit instructions, Mohamed does not teach wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.”
Examiner notes that the Office Action relied upon Mohamed to render obvious the aforementioned limitation (in conjunction with other prior art references), rather than teach the entirety of the limitation by itself.
Applicant on page 9 argues: “Additionally, Mohamed, like Lumsdaine and Curry, is a stored program machine which is fundamentally different from the presently claimed invention, a domain flow execution engine. It is not proper to combine the teachings related to a stored program machine with a domain flow execution engine.”
Examiner generally submits that it would not necessarily be the case that any given teaching disclosed in the context/environment of a stored program machine would be improper to combine with a domain flow execution engine. Examiner further submits that the concept and desirability of code density and code size (e.g., that smaller code size is more desirable over larger code size, as, for example, less memory space is needed) is applicable regardless of the particular type of architecture (whether a stored program machine, a domain flow execution engine, or something else) the code is run on.
Applicant on page 9 argues: “As described previously, there is a significant difference between a domain flow execution engine (e.g., the presently claimed invention) and a stored program machine (e.g., used in Lumsdaine and Curry). Therefore, the teachings of Lumsdaine and Curry are not compatible with the teachings of Omtzigt. More specifically, in a stored program machine, the kernels are simple programs that can execute and request data from memory. Furthermore, the kernel in a stored program machine manages and schedules a linear sequence of instructions. In a domain flow engine, kernels are a collective state that programs the 'personality' of the compute fabric, so that each processing element has sufficient state to be autonomous when data arrives. Moreover, a domain flow execution engine kernel orchestrates a network of independent, concurrent processes connected by data flows. This is understood by one skilled in the art, and thus, it is also understood that although the same term "kernel" is used, there are clear differences between a kernel in a domain flow execution engine and a kernel for a stored program machine. Therefore, a kernel in a stored program machine cannot perform the same tasks as a domain flow execution engine kernel. In other words, the tasks a kernel performs in a stored-program machine are fundamentally different from those of a domain-flow execution engine. While a stored-program kernel is a low-level, privileged resource manager for a single computer, a domain-flow engine orchestrates high-level tasks and data across systems. As also described previously, in domain flow, it is a priori worked out how to configure the distributed data path as part of a kernel, whereas in a stored program machine, each instruction can do anything it wants, request far fetched data through cache hierarchy or network, jump to new subroutines, take an interrupt, etc. The whole control mechanism is different, and therefore, Lumsdaine and Curry have nothing to do with the domain flow execution engine machine of the presently claimed invention. Since they have nothing to do with the domain flow execution engine, their teachings are inherently incompatible with the teachings of Omtzigt, and thus the combination is improper.”
Examiner appreciates the explanation contrasting a domain flow execution engine and a
stored program machine. However, Examiner submits that the claims as currently recited do not
appear to preclude Lumsdaine and Curry from being relied upon in the manner of the above prior
art rejections, even if Lumsdaine and Curry are not directed to the domain flow execution engine
machine of the instant invention.
Examiner submits that even if Lumsdaine and Curry are not directed to the domain flow execution engine machine of the instant invention, such does not preclude specific teachings from Lumsdaine and Curry from being compatible with the teachings of Omtzigt. For example, Examiner submits that index structures being used for sparse matrices to minimize bandwidth, as taught by Lumsdaine, are not inherently incompatible with a domain flow engine. For example, Examiner submits that parallelism and chaining, as taught by Curry, are not inherently incompatible with kernels of a domain flow engine. Examiner generally submits that teachings that are not disclosed as being incorporated in a particular environment are not necessarily inherently incompatible with that environment.
Applicant on page 10 argues: “However, Lumsdaine does not teach, for example, a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory. Additionally, Lumsdaine does not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.”
However, Examiner relied upon further references to render obvious the aforementioned subject matter.
Applicant on page 11 argues: “However, Curry does not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.”
However, Examiner submits that Curry teaches the aforementioned subject matter, as cited. Examiner submits that the kernels can be reasonably considered to be “parallel” kernels, given that they are executed by SIMD (i.e., single-instruction-multiple-data) processors, which performs multiple computations in parallel. (Also note that Nordquist, cited as pertinent in a previous office action, similarly discloses parallel kernels, in disclosing data level parallelism in that data is processed in parallel computation units.) To any extent to which the kernels of the instant invention are parallel in a different manner, Examiner submits that the broadest reasonable interpretation of “parallel kernels” does not preclude Curry from being relied upon in the manner of the above prior art rejections.
Applicant on page 11 argues: “Additionally, Curry does not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.”
However, Examiner relied upon a further reference to render obvious the aforementioned subject matter.
Applicant on page 11 argues: “However, Matick does not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory”.
However, Examiner relied upon Curry to teach the aforementioned subject matter, as noted above.
Applicant on page 11 argues: “Additionally, Matick does not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.”
However, Examiner relied upon a further reference to render obvious the aforementioned subject matter.
Applicant on page 12 argues: “However, as described above, Mohamed does not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.”
Examiner notes that the Office Action relied upon Mohamed to render obvious the aforementioned limitation (in conjunction with other prior art references), rather than teach the entirety of the limitation by itself.
Applicant on page 12 argues: “As described above, Omtzigt, Lumsdaine, Curry, Matick, Mohamed and their combination do not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.” Applicant on page 12 further argues: “As also described above, Omtzigt, Lumsdaine, Curry, Matick, Mohamed and their combination do not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.” Applicant on page 13 argues: “As described above, Omtzigt, Lumsdaine, Curry, Matick and their combination do not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.” Applicant across pages 13-14 argues: “As also described above, Omtzigt, Lumsdaine, Curry, Matick, Glasco and their combination do not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.”
However, Examiner submits that Curry teaches the aforementioned subject matter, as noted above.
Applicant on page 12 argues: “Additionally, Omtzigt, Lumsdaine, Curry, Matick, Mohamed and their combination do not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.” Applicant on page 12 further argues: “Omtzigt, Lumsdaine, Curry, Matick, Mohamed and their combination do not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.”
However, Examiner submits that Mohamed renders obvious the aforementioned limitation (in conjunction with other prior art references), as addressed above.
Applicant on page 12 argues: “As described above, Omtzigt is not prior art to the present application.” Applicant on page 13 argues: “As described above, Applicants note that U.S. Patent Application Publication No. 2009/0300327 by Omtzigt is not prior art to the presently claimed invention as the present application claims priority to the '327 application in the preliminary amendment and the application data sheet. Therefore, the rejection should be withdrawn.” Applicant on page 13 argues: “As described above, Omtzigt is not prior art to the present application.” Applicant on page 14 argues: “As described above, Applicants note that U.S. Patent Publication No. 2009/0300327 by Omtzigt is not prior art to the presently claimed invention as the present application claims priority to the '327 application in the preliminary amendment and the application data sheet. Therefore, the rejection should be withdrawn.” Applicant on page 15 argues: “As described above, Omtzigt is not prior art to the present application.”
However, as addressed above, Omtzigt appears to be prior art to the claimed invention.
Applicant on page 13 argues: “Again, Glasco is a stored program machine, not a domain flow execution engine (e.g., the presently claimed invention). Therefore, Glasco is not applicable to the presently claimed invention.”
As addressed above, Examiner generally submits that it would not necessarily be the case that any given teaching disclosed in the context/environment of a stored program machine would be improper to combine with a domain flow execution engine.
Applicant on page 14 argues: “Again, Martin is a stored program machine, not a domain flow execution engine (e.g., the presently claimed invention). Therefore, Martin is not applicable to the presently claimed invention.” Applicant on page 15 argues: “As also described above, Omtzigt, Lumsdaine, Curry, Matick, Martin and their combination do not teach, for example, a first set of processing elements implemented as an asynchronous execution pipeline.”
As addressed above, Examiner generally submits that it would not necessarily be the case that any given teaching disclosed in the context/environment of a stored program machine would be improper to combine with a domain flow execution engine.
Applicant on page 15 argues: “Applicants also note that this rejection involves seven (7) references. Applicants understand that there is no specific limit to the number of references able to be used for an obviousness rejection, but it should be worth reconsidering an obviousness rejection when seven (7) references are needed.”
Examiner appreciates Applicant’s understanding that there is no specific limit to the number of references able to be used for an obviousness rejection and also understands Applicant’s general concern. However, in the instant case, Examiner submits that the claimed subject matter of claim 49 is wide-ranging in scope (encompassing distinct concepts that may be, but do not necessarily have to be, used together, including, but not limited to, hashing algorithms, a domain flow execution engine per se, sparse matrices, parallel kernels, chained kernels, and caching), such that use of seven references is reasonably commensurate with this scope. For example, Examiner submits that the inventive concept related to sparse matrix optimization is distinct enough from the inventive concept related to chaining kernels that the reliance of respective references to teach the aforementioned inventive concepts in a same claim, rather than relying on just a single reference, would not be a tip of an iceberg that reflects a deeper issue with the rejection.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEITH E VICARY whose telephone number is (571)270-1314. The examiner can normally be reached Monday to Friday, 9:00 AM to 5:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached at (571)270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KEITH E VICARY/Primary Examiner, Art Unit 2183