Prosecution Insights
Last updated: October 04, 2026
Application No. 18/211,447

EXECUTION ENGINE FOR EXECUTING SINGLE ASSIGNMENT PROGRAMS WITH AFFINE DEPENDENCIES

Non-Final OA §103§112
Filed
Jun 19, 2023
Priority
May 27, 2008 — provisional 61/130,114 +6 more
Examiner
VICARY, KEITH E
Art Unit
2183
Tech Center
2100 — Computer Architecture & Software
Assignee
Stillwater Supercomputing Inc.
OA Round
9 (Non-Final)
58%
Grant Probability
Moderate
9-10
OA Rounds
7m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 58% of resolved cases
58%
Career Allowance Rate
403 granted / 698 resolved
+2.7% vs TC avg
Strong +40% interview lift
Without
With
+40.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 10m
Avg Prosecution
36 currently pending
Career history
746
Total Applications
across all art units

Statute-Specific Performance

§101
10.1%
-29.9% vs TC avg
§103
34.6%
-5.4% vs TC avg
§102
11.4%
-28.6% vs TC avg
§112
37.2%
-2.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 698 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on April 28, 2026, has been entered. Claims 1, 9, 13-15, 31-36, 38-50, 52-57, and 59-60 are pending in this office action and presented for examination. Claim 36 is newly amended by the RCE received April 28, 2026. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 36 and 38-46 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 36 recites the limitation “a cache is used for recalling a previous kernel of the plurality of parallel kernels, wherein the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory” in lines 8-10. However, the original disclosure does not appear to provide support for this limitation. For example, the original disclosure (e.g., page 20, line 14; page 18, line 12) does not appear to provide support for a same cache both a) being used for recalling a previous kernel of the plurality of parallel kernels, and b) being further configured for accumulating page coherent data for writeback to dynamic random access memory. Claim 36 recites the limitation “the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory” in lines 9-10. However, the original disclosure does not appear to provide support for this limitation. For example, the original disclosure (e.g., page 18, lines 12-13) does not appear to provide support for a cache itself being “configured for” accumulating page coherent data for writeback to dynamic random access memory. Claims 38-46 are rejected for failing to alleviate the rejections of claim 36 above. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 36 and 38-46 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 36 recites the limitation “the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory” in lines 9-10. However, the metes and bounds of this limitation are indefinite, as the cache was not previously recited to be “configured” for other functionality. It is indefinite as to whether the “further” language is implicitly requiring the cache to be configured with other functionality. Claims 38-46 are rejected for failing to alleviate the rejection of claim 36 above. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 9, 13-15, and 31-35 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt (US 20090300327 A1) in view of Lumsdaine et al. (Lumsdaine) (US 20070198621 A1) in view of Curry et al. (Curry) (US 20100315428 A1) in view of Matick et al. (Matick) (US 20050114606 A1) in view of Mohamed (US 6366998 B1). Consider claim 1, Omtzigt discloses a computing device ([0027], line 2, data flow computer system) comprising: a memory for storing data and a domain flow program ([0027], lines 4-5, memory 110 that contains the data and the program to execute; [0027], lines 13-14, programming information; [0022], lines 1-2, an execution engine executes single assignment programs with affine dependencies; see the remarks dated 1/20/2012 in the parent application 12/467485, beginning on page 10, for a greater explanation of how Omtzigt teaches domain flow); a controller for requesting the data and the domain flow program from the memory ([0027], lines 6-10, execution starts by the controller 120 requesting a program from the memory 110. The controller 120 presents a read request via bus 121 to the memory controller 130 which translates the read request to a memory request and returns the data to the controller 120. This data contains the program instructions to execute a single assignment program); and a processor fabric for processing the data and the domain flow program ([0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams) via a plurality of processing elements ([0029], line 3, processing elements (PE) 310). However, Omtzigt does not entail that, for sparse matrices, index structures are used to minimize memory bandwidth. Omtzigt also does not disclose a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory. Omtzigt also does not disclose the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer. On the other hand, Lumsdaine discloses that, for sparse matrices, index structures are used to minimize memory bandwidth (for example, see [0028], lines 3-10, the term "sparse matrix" refers to a matrix where the vast majority (usually at least 99% for large matrices) of the element values are zero. The number of nonzero elements in a given matrix is denoted by NNZ. Matrices with this property may be stored much more efficiently in memory if only the nonzero data values are stored, as well as some index data to identify where those data values fit in the matrix). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Lumsdaine with Omtzigt in order to increase memory efficiency (Lumsdaine, [0028], lines 3-10). However, the combination thus far does not entail a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory. The combination thus far also does not disclose the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer. On the other hand, Curry discloses a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from memory ([0015], lines 8-10, functions can be chained together to process images without intermediate trips to external memory; [0017], lines 1-2, SIMD processors execute functions known as kernels). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Curry with the combination of Omtzigt and Lumsdaine in order to increase system performance. However, the combination thus far does not entail a cache for recalling a previous kernel. The combination thus far also does not disclose the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer. On the other hand, Matick discloses a cache for recalling a previous kernel (for example, [0046], lines 9-10, kernels in the cache). Matick’s teaching increases speed (Matick, [0002], lines 1-2, caches are used to speed up accesses to recently accessed information). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Matick with the combination of Omtzigt, Lumsdaine, and Curry in order to increase speed. However, the combination thus far does not disclose the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer. On the other hand, Mohamed discloses a density such that parallel computations are expressed in 100 bytes or fewer (FIG. 6, which shows a vector MAC instruction being expressed in 2 bytes, and an overall instruction packet being expressed in 32 bytes). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Mohamed with the combination of Omtzigt, Lumsdaine, Curry, and Matick in order to save space relative to implementations which necessitate a greater number of bytes to express parallel computations. Consider claim 9, the overall combination entails the computing device of claim 1 (see above) wherein the controller is configured for presenting a read request to a memory controller which translates the read request to a memory request and returns the data to the controller (Omtzigt, [0027], lines 7-10, the controller 120 presents a read request via bus 121 to the memory controller 130 which translates the read request to a memory request and returns the data to the controller 120). Consider claim 13, the overall combination entails the computing device of claim 1 (see above) further comprising a crossbar configured for routing the data to rows or columns in the processor fabric (Omtzigt, [0027], lines 27-29, the crossbar 150 routes the data streams to the appropriate rows or columns in the processor fabric 160). Consider claim 14, the overall combination entails the computing device of claim 13 (see above) wherein the processor fabric is configured for producing output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams). Consider claim 15, the overall combination entails the computing device of claim 14 wherein the output data is written to the memory by traversing through the crossbar and whereupon the output data is presented to a memory controller which writes the output data into the memory (Omtzigt, [0027], lines 31-36, these output data streams are written back to memory 110 by traversing the crossbar 150 to the streamers 140 that associate memory addresses to the data streams, and then present them to the memory controller 130, which will write the data streams into memory 110). Consider claim 31, the overall combination entails the computing device of claim 1 (see above) wherein the processor fabric is programmable (Omtzigt, [0032], lines 24-26, after the controller 120 is done programming the processor fabric 160, execution is able to commence). Consider claim 32, the overall combination entails the computing device of claim 1 (see above) wherein the processor fabric receives the data, executes instructions on the data, and produces output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams). Consider claim 33, the overall combination entails the computing device of claim 1 (see above) wherein the data is processed under control of the domain flow program which is a single assignment program (Omtzigt, [0033], lines 10-11, execution of a single assignment program on the fabric). Consider claim 34, the overall combination entails the computing device of claim 33 (see above) wherein the plurality of processing elements are configured as a network of processing elements (Omtzigt, [0029], lines 16-17, processing element routing network 320) configured for executing the single assignment program (Omtzigt, [0033], lines 10-11, execution of a single assignment program on the fabric). Consider claim 35, the overall combination entails the computing device of claim 1 (see above) wherein the plurality of processing elements are configured as a matrix of processing elements (Omtzigt, [0029], line 3, processor array). Claim(s) 36 and 38-46 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt (US 20090300327 A1) in view of Lumsdaine et al. (Lumsdaine) (US 20070198621 A1) in view of Curry et al. (Curry) (US 20100315428 A1) in view of Matick et al. (Matick) (US 20050114606 A1) in view of Glasco (US 20030182514 A1) in view of Ledbetter, JR. et al. (Ledbetter) (US 5119485) in view of Weisser et al. (Weisser) (US 5530941). Consider claim 36, Omtzigt discloses a method comprising: storing, in a memory of a device, data and a domain flow program ([0027], lines 4-5, memory 110 that contains the data and the program to execute; [0027], line 2, data flow computer system; [0027], lines 13-14, programming information; [0022], lines 1-2, an execution engine executes single assignment programs with affine dependencies; see the remarks dated 1/20/2012 in the parent application 12/467485, beginning on page 10, for a greater explanation of how Omtzigt teaches domain flow); requesting, with a controller, the data and the domain flow program from the memory ([0027], lines 6-10, execution starts by the controller 120 requesting a program from the memory 110. The controller 120 presents a read request via bus 121 to the memory controller 130 which translates the read request to a memory request and returns the data to the controller 120. This data contains the program instructions to execute a single assignment program); and processing, with a processor fabric, the data and the domain flow program ([0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams) via a plurality of elements ([0029], line 3, processing elements (PE) 310). However, Omtzigt does not entail that for sparse matrices, index structures are used to minimize memory bandwidth. Omtzigt also does not disclose chaining a plurality of parallel kernels to avoid serializing intermediate data to and from the memory, wherein a cache is used for recalling a previous kernel of the plurality of parallel kernels. Omtzigt also does not disclose a plurality of processor fabrics configured to communicate with each other. Omtzigt also does not disclose the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory. On the other hand, Lumsdaine discloses that for sparse matrices, index structures are used to minimize memory bandwidth (for example, see [0028], lines 3-10, the term "sparse matrix" refers to a matrix where the vast majority (usually at least 99% for large matrices) of the element values are zero. The number of nonzero elements in a given matrix is denoted by NNZ. Matrices with this property may be stored much more efficiently in memory if only the nonzero data values are stored, as well as some index data to identify where those data values fit in the matrix). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Lumsdaine with Omtzigt in order to increase memory efficiency (Lumsdaine, [0028], lines 3-10). However, the combination thus far does not entail chaining a plurality of parallel kernels to avoid serializing intermediate data to and from the memory, wherein a cache is used for recalling a previous kernel of the plurality of parallel kernels. The combination thus far also does not disclose a plurality of processor fabrics configured to communicate with each other. The combination thus far also does not disclose the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory. On the other hand, Curry discloses a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from memory ([0015], lines 8-10, functions can be chained together to process images without intermediate trips to external memory; [0017], lines 1-2, SIMD processors execute functions known as kernels). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Curry with the combination of Omtzigt and Lumsdaine in order to increase system performance. However, the combination thus far does not entail that a cache is used for recalling a previous kernel. The combination thus far also does not disclose a plurality of processor fabrics configured to communicate with each other. The combination thus far also does not disclose the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory. On the other hand, Matick discloses that a cache is used for recalling a previous kernel (for example, [0046], lines 9-10, kernels in the cache). Matick’s teaching increases speed (Matick, [0002], lines 1-2, caches are used to speed up accesses to recently accessed information). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Matick with the combination of Omtzigt, Lumsdaine, and Curry in order to increase speed. However, the combination thus far does not disclose a plurality of processor fabrics configured to communicate with each other. The combination thus far also does not disclose the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory. On the other hand, Glasco discloses a plurality of processor fabrics configured to communicate with each other ([0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture). Glasco’s teaching reduces the number of links used and further modularizes a multiprocessor system (Glasco, [0033], lines 15-18, in order to reduce the number of links used and to further modularize a multiprocessor system using a point-to-point architecture, multiple clusters are used). In addition, a plurality of processor fabrics increases system performance relative to one of that processor fabric. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Glasco with the combination of Omtzigt, Lumsdaine, Curry, and Matick in order to reduce the number of links used and to further modularize a multiprocessor system and to increase system performance. However, the combination thus far also does not disclose the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory. On the other hand, Ledbetter discloses a cache is configured for accumulating page coherent data for writeback to main memory (col. 2, lines 55-59, the PUSH instruction forces the write-back cache to search all cache entries for `dirty` data which the pending page-out operation may access, and to copy these entries back into main memory). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Ledbetter with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Glasco in order to efficiently implement coherency between a cache and a main memory. However, the combination thus far does not entail that the main memory is a dynamic random access memory in particular. On the other hand, Weisser discloses a main memory is a dynamic random access memory (col. 1, lines 59-60, main memory consisting of DRAMs). DRAM is denser and less expensive than SRAM, and can therefore be significantly larger (Weisser, col. 1, lines 59-62). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Weisser with the combination of Omtzigt, Lumsdaine, Curry, Matick, Glasco, and Ledbetter to obtain increased main memory size at low cost. Alternatively, this modification merely entails combining prior art elements (the prior art elements of the combination of Omtzigt, Lumsdaine, Curry, Matick, Glasco, and Ledbetter, and Weisser’s teaching of DRAM) according to known methods (Examiner submits that DRAM was well known, as reflected by Weisser) to yield predictable results (the combination of Omtzigt, Lumsdaine, Curry, Matick, Glasco, and Ledbetter, wherein the main memory is DRAM), which is an example of a rationale that may support a conclusion of obviousness as per MPEP 2143. Consider claim 38, the overall combination entails the method of claim 36 (see above) wherein the controller is configured for presenting a read request to a memory controller which translates the read request to a memory request and returns the data to the controller (Omtzigt, [0027], lines 7-10, the controller 120 presents a read request via bus 121 to the memory controller 130 which translates the read request to a memory request and returns the data to the controller 120). Consider claim 39, the overall combination entails the method of claim 36 (see above) further comprising routing, with a crossbar, the data to rows or columns in the plurality of processor fabrics (Omtzigt, [0027], lines 27-29, the crossbar 150 routes the data streams to the appropriate rows or columns in the processor fabric 160; Glasco, [0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture). Consider claim 40, the overall combination entails the method of claim 39 (see above) wherein the plurality of processor fabrics are configured for producing output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams; Glasco, [0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture). Consider claim 41, the overall combination entails the method of claim 40 wherein the output data is written to the memory by traversing through the crossbar and whereupon the output data is presented to a memory controller which writes the output data into the memory (Omtzigt, [0027], lines 31-36, these output data streams are written back to memory 110 by traversing the crossbar 150 to the streamers 140 that associate memory addresses to the data streams, and then present them to the memory controller 130, which will write the data streams into memory 110). Consider claim 42, the overall combination entails the method of claim 36 (see above) wherein the plurality of processor fabrics are programmable (Omtzigt, [0032], lines 24-26, after the controller 120 is done programming the processor fabric 160, execution is able to commence; Glasco, [0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture). Consider claim 43, the overall combination entails the method of claim 36 (see above) wherein the plurality of processor fabrics receive the data, execute instructions on the data, and produce output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams; Glasco, [0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture). Consider claim 44, the overall combination entails the method of claim 36 (see above) wherein the data is processed under control of the domain flow program which is a single assignment program (Omtzigt, [0033], lines 10-11, execution of a single assignment program on the fabric). Consider claim 45, the overall combination entails the method of claim 36 (see above) wherein the plurality of processing elements are configured as a matrix of processing elements (Omtzigt, [0029], line 3, processor array). Consider claim 46, the overall combination entails the method of claim 36 (see above) wherein the plurality of elements are configured as a network of processing elements (Omtzigt, [0029], lines 16-17, processing element routing network 320) configured for executing the single assignment program (Omtzigt, [0033], lines 10-11, execution of a single assignment program on the fabric). Claim(s) 47 and 52-55 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt (US 20090300327 A1) in view of Lumsdaine et al. (Lumsdaine) (US 20070198621 A1) in view of Curry et al. (Curry) (US 20100315428 A1) in view of Matick et al. (Matick) (US 20050114606 A1) in view of Martin et al. (Martin) (US 20020156995 A1). Consider claim 47, Omtzigt discloses a processor fabric for processing ([0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams) a domain flow program ([0027], lines 4-5, memory 110 that contains the data and the program to execute) comprising: a first set of processing elements ([0029], line 3, processing elements (PE) 310); and a second set of processing elements which communicate with the first set of processing elements ([0029], line 3, processing elements (PE) 310; [0029], lines 16-17, processing element routing network 320), wherein the first set of processing elements and the second set of processing elements are configured for processing data and the domain flow program ([0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams). However, Omtzigt does not entail that for sparse matrices, index structures are used to minimize memory bandwidth. Omtzigt also does not disclose a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained. Omtzigt also does not disclose the first set of processing elements is implemented as an asynchronous execution pipeline. On the other hand, Lumsdaine discloses that for sparse matrices, index structures are used to minimize memory bandwidth (for example, see [0028], lines 3-10, the term "sparse matrix" refers to a matrix where the vast majority (usually at least 99% for large matrices) of the element values are zero. The number of nonzero elements in a given matrix is denoted by NNZ. Matrices with this property may be stored much more efficiently in memory if only the nonzero data values are stored, as well as some index data to identify where those data values fit in the matrix). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Lumsdaine with Omtzigt in order to increase memory efficiency (Lumsdaine, [0028], lines 3-10). However, the combination thus far does not entail a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained. The combination thus far also does not disclose the first set of processing elements is implemented as an asynchronous execution pipeline. On the other hand, Curry discloses a plurality of parallel kernels, wherein the plurality of parallel kernels are chained ([0015], lines 8-10, functions can be chained together to process images without intermediate trips to external memory; [0017], lines 1-2, SIMD processors execute functions known as kernels). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Curry with the combination of Omtzigt and Lumsdaine in order to increase system performance. However, the combination thus far does not entail a cache for recalling a previous kernel. The combination thus far also does not disclose the first set of processing elements is implemented as an asynchronous execution pipeline. On the other hand, Matick discloses a cache for recalling a previous kernel (for example, [0046], lines 9-10, kernels in the cache). Matick’s teaching increases speed (Matick, [0002], lines 1-2, caches are used to speed up accesses to recently accessed information). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Matick with the combination of Omtzigt, Lumsdaine, and Curry in order to increase speed. However, the combination thus far does not disclose the first set of processing elements is implemented as an asynchronous execution pipeline. On the other hand, Martin discloses a first set of processing elements is implemented as an asynchronous execution pipeline ([0047], line 2, asynchronous execution pipeline). Martin’s teaching increases speed (Martin, [0047], lines 13-16). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Martin with the combination of Omtzigt, Lumsdaine, Curry, and Matick in order to increase speed. Consider claim 52, the overall combination entails the processor fabric of claim 47 (see above) wherein the first set of processing elements and the second set of processing elements are further configured for receiving the data and producing output data (Omtzigt, [0027], lines 29-31, the processor fabric 160 receives the incoming data streams, executes instructions on these streams and produces output data streams). Consider claim 53, the overall combination entails the processor fabric of claim 52 wherein the output data is written to a memory by traversing through a crossbar to a memory controller which writes the output data into the memory (Omtzigt, [0027], lines 31-36, these output data streams are written back to memory 110 by traversing the crossbar 150 to the streamers 140 that associate memory addresses to the data streams, and then present them to the memory controller 130, which will write the data streams into memory 110). Consider claim 54, the overall combination entails the processor fabric of claim 47 (see above) wherein the first set of processing elements and the second set of processing elements recognize a spatial tag of a computational event and take action under control of the domain flow program (Omtzigt, [0029], lines 21-24, the PEs 310 recognize the spatial tag called the signature of a computational event and take action under control of the single assignment program installed in their program store by controller 120). Consider claim 55, the overall combination entails the processor fabric of claim 47 (see above) wherein the domain flow program evolves as multi-dimensional data match up with the first set of processing elements and the second set of processing elements and produce new multi-dimensional data (Omtzigt, [0029], lines 28-32, the overall computation represented by the single assignment program evolves as the multi-dimensional data streams match up within the processing elements 310 and producing potentially new multi-dimensional data streams). Claim(s) 48 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Jensen et al. (Jensen) (US 20070094482 A1). Consider claim 48, the combination thus far entails the processor fabric of claim 47 (see above), but does not entail that the first set of processing elements and the second set of processing elements process an instruction set architecture configured for a specific class of algorithms. On the other hand, Jensen discloses processing elements processing an instruction set architecture configured for a specific class of algorithms ([0007], lines 14-19, accordingly, a system designer can specify a specific ISA for encoding a specific set of program functions (e.g., signal processing algorithms) and select other ISAs for encoding other types of program functions (e.g., operating system functions, I/O functions, general purpose functions)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Jensen with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin in order to provide support for a specific class of algorithms, such as signal processing algorithms, thereby increasing device capability. Claim(s) 49 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen as applied to claim 48 above, and further in view of Zakiya (US 20020032551 A1). Consider claim 49, the combination thus far entails the processor fabric of claim 48 (see above), but does not entail the specific class of algorithms comprise hashing algorithms. On the other hand, Zakiya discloses hashing algorithms ([0005], lines 1-2, The current most widely used hash algorithms are MD5 and the Secure Hash Algorithm (SHA-1)). Zakiya’s teaching is useful for computing a unique condensed representation of a message or a data file, which is useful for security (Zakiya, [0003]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Zakiya with the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen in view of its usefulness for computing a unique condensed representation of a message or a data file, which is useful for security. Claim(s) 50 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen as applied to claim 48 above, and further in view of Benkelman (US 6694064 A1). Consider claim 50, the combination thus far entails the processor fabric of claim 48 (see above), but does not entail the specific class of algorithms are utilized for interpolations and resampling. On the other hand, Benkelman discloses optimizations for interpolations and resampling (col. 5, lines 30-32, two common resampling algorithms are referred to as "bilinear interpolation" and "nearest neighbor."). Benkelman’s teaching is fundamental in the rectification or transformation process (Benkelman, col. 5, lines 23-24). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Benkelman with the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen in view of its usefulness in the rectification or transformation process. Claim(s) 56 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Eggers et al. (Eggers) (US 20060179429 A1). Consider claim 56, the combination thus far entails the processor fabric of claim 47 (see above), wherein each processing element of the second set of processing elements includes a set of a content addressable memory (FIG. 6, Tag CAM 640), one or more functional units ([0036], line 28, functional units), and a router configured to generate affine routing vectors (Fig. 4, packet router 425; [0032], line 37, routing vector; [0032], lines 20-21, routing vector. This information defines some affine recurrence equation). However, the combination thus far does not disclose an instruction scheduling and dispatch queue. On the other hand, Eggers discloses an instruction scheduling and dispatch queue ([0172], lines 1-2, the PE selects an instruction from the scheduling queue). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Eggers with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin, as this modification merely entails combining prior art elements (the prior art elements of the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin, and Eggers’s teaching of an instruction scheduling and dispatch queue) according to known methods (Examiner submits that use of an instruction scheduling and dispatch queue was very well-known to one of ordinary skill in the art before the effective filing date of the claimed invention) to yield predictable results (the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin, further modified to incorporate an instruction scheduling and dispatch queue), which is an example of a rationale that may support a conclusion of obviousness as per MPEP 2143. Alternatively, Examiner submits that an instruction scheduling and dispatch queue facilitates efficient instruction scheduling and dispatching relative to not using a queue. Claim(s) 57 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Jensen et al. (Jensen) (US 20070094482 A1) and Bhanot et al. (Bhanot) (US 20040073590 A1). Consider claim 57, the combination thus far entails the processor fabric of claim 47 (see above). However, the combination thus far does not entail that a plurality of instruction sets drive a plurality of global operators. On the other hand, Jensen discloses a plurality of instruction sets ([0007], lines 14-19, accordingly, a system designer can specify a specific ISA for encoding a specific set of program functions (e.g., signal processing algorithms) and select other ISAs for encoding other types of program functions (e.g., operating system functions, I/O functions, general purpose functions)) It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Jensen with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin in order to increase system capability by providing support for a wide range of program functions. However, the combination thus far does not entail that the plurality of instruction sets drive a plurality of global operators. On the other hand, Bhanot discloses global operators ([0003], line 4, global operation). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Bhanot with the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen, as this modification merely entails combining prior art elements (the prior art elements of the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen, and Bhanot’s teaching of a global operator) according to known methods (Examiner submits that the general concept of a global operator was well-known to one of ordinary skill in the art before the effective filing date of the claimed invention, and Bhanot also discloses execution of a global operation in the environment of a network of computing nodes) to yield predictable results (the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen, further modified to support global operations), which is an example of a rationale that may support a conclusion of obviousness as per MPEP 2143. Alternatively, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Bhanot with the combination of Omtzigt, Lumsdaine, Curry, Matick, Martin, and Jensen in order to improve efficiency (Bhanot, [0002], lines 7-8). Claim(s) 59 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Glasco (US 20030182514 A1). Consider claim 59, the combination thus far entails the processor fabric of claim 47 (see above), but does not entail the processor fabric is implemented on a chip with one or more additional processor fabrics, wherein the processor fabric is configured to communicate with the one or more additional processor fabrics. On the other hand, Glasco discloses a processor fabric is implemented on a chip with one or more additional processor fabrics, wherein the processor fabric is configured to communicate with the one or more additional processor fabrics ([0034], lines 1-3, the multiple processor clusters are interconnected using a point-to-point architecture). Glasco’s teaching reduces the number of links used and further modularizes a multiprocessor system (Glasco, [0033], lines 15-18, in order to reduce the number of links used and to further modularize a multiprocessor system using a point-to-point architecture, multiple clusters are used). In addition, a plurality of processor fabrics increases system performance relative to one of that processor fabric. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Glasco with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin in order to reduce the number of links used and to further modularize a multiprocessor system and to increase system performance. Claim(s) 60 is/are rejected under 35 U.S.C. 103 as being unpatentable over Omtzigt, Lumsdaine, Curry, Matick, and Martin as applied to claim 47 above, and further in view of Chan et al. (Chan) (US 20060282650 A1). Consider claim 60, the combination thus far entails the processor fabric of claim 47 (see above), but does not entail each processing element of the second set of processing elements comprises a Single Instruction Multiple Data (SIMD) unit configured to process a plurality of operations per instruction. On the other hand, Chan discloses processing elements each comprise a Single Instruction Multiple Data (SIMD) unit configured to process a plurality of operations per instruction (FIG. 1B, MFU 222; [0023], lines 10-17, the media functional units 222 are single-instruction-multiple-data (SIMD) media functional units. Each media functional unit 222 is capable of processing parallel 16-bit components, in addition to 32-bit operands. Various parallel 16-bit operations supply the single-instruction-multiple-data capability for processor 100 including add, multiply-add, shift, compare, and the like). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Chan with the combination of Omtzigt, Lumsdaine, Curry, Matick, and Martin in order to increase performance via parallelism. Response to Arguments Applicant on page 8 argues: “Within the previous Office Action, Figures 3-5 have been objected to. By the above amendments, the drawings have been clarified. Applicants have ensured that the drawings are free from arrays of white dots. No new matter has been added. The objection should be withdrawn.” In view of the aforementioned amendments, the previously presented objections to the drawings are withdrawn. Applicant on page 8 argues: ‘Within the Advisory Action, it has been asserted that Curry teaches the limitation: wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory. Within the Advisory Action, it has been argued that the kernels in Curry can be reasonably considered to be "parallel" kernels, since they are executed by SIMD processors which perform multiple computations in parallel. Applicants respectfully disagree. Although a SIMD performs multiple computations, it does not implement multiple independent kernel instances. Therefore, Curry does not teach the presently claimed limitation.’ However, the claims do not recite multiple independent kernel instances. Applicant across pages 8-9 argues: “Within the Advisory Action, it has been argued that Mohamed in conjunction with other prior art references teach the limitation: wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer. Applicants respectfully disagree and request further clarification as to how the references are combined to teach the presently claimed invention. Specifically, it is not clear how the instruction packet having 32-bit instructions of Mohamed is combined with other prior art to teach the limitation: the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer. However, as noted in the rejection, Mohamed discloses a density such that parallel computations are expressed in 100 bytes or fewer (FIG. 6, which shows a vector MAC instruction being expressed in 2 bytes, and an overall instruction packet being expressed in 32 bytes), and it is this teaching which is combined with the combination of Omtzigt, Lumsdaine, Curry, and Matick (which entails a domain flow program in particular) to teach the overall relevant limitation. Applicant across pages 8-9 argues: "As described previously, Applicants assert that several of the cited references are nonanalogous art. MPEP 2141.01(a) specifies that: I. TO RELY ON A REFERENCE UNDER 35 U.S.C. 103, IT MUST BE ANALOGOUS ART TO THE CLAIMED INVENTION In order for a reference to be proper for use in an obviousness rejection under 35 U.S.C. 103 , the reference must be analogous art to the claimed invention. In re Bigio, 381 F.3d 1320, 1325, 72 USPQ2d 1209, 1212 (Fed. Cir. 2004). A reference is analogous art to the claimed invention if: (1) the reference is from the same field of endeavor as the claimed invention (even if it addresses a different problem); or (2) the reference is reasonably pertinent to the problem faced by the inventor (even if it is not in the same field of endeavor as the claimed invention). As described previously, Mohamed, Lumsdaine and Curry are directed toward stored program machines, not domain flow execution engines. Therefore, these are nonanalogous art and should not be used." However, Examiner submits that the aforementioned cited references are from the same field of endeavor as the claimed invention (computing). Examiner further submits that the aforementioned cited references are reasonably pertinent to the problem faced by the inventor, as the advantages of the cited teachings of the aforementioned cited references are applicable regardless of whether the advantages are applied to stored program machines or domain flow execution engines. For example, Examiner submits that the concept and desirability of code density and code size (e.g., that smaller code size is more desirable over larger code size, as, for example, less memory space is needed) is applicable regardless of the particular type of architecture (whether a stored program machine, a domain flow execution engine, or something else) the code is run on. For example, Examiner submits that the concept and desirability of minimizing memory bandwidth for sparse matrices is applicable regardless of the particular type of architecture (whether a stored program machine, a domain flow execution engine, or something else) in which the sparse matrices are handled. Applicant on page 9 argues: "This further bolsters Applicants point that an excessive number of references are utilized to reject the presently claimed invention. Not only are seven (7) references used to reject some of the claims, but several of the references are nonanalogous art; thus, further suggesting that the rejections are improper, and the presently claimed invention is allowable." Examiner submits that the reference are analogous art; see above. In addition, Examiner submits that the claimed subject matter of the claims is wide-ranging in scope (encompassing distinct concepts that may be, but do not necessarily have to be, used together, including, but not limited to, hashing algorithms, a domain flow execution engine per se, sparse matrices, parallel kernels, chained kernels, and caching), such that use of seven references is reasonably commensurate with this scope. For example, Examiner submits that the inventive concept related to sparse matrix optimization is distinct enough from the inventive concept related to chaining kernels that the reliance of respective references to teach the aforementioned inventive concepts in a same claim, rather than relying on just a single reference, would not be a tip of an iceberg that reflects a deeper issue with the rejection. Applicant on page 9 argues: "Even if the references are considered analogous art, they are not functionally compatible with the other references. Therefore, one skilled in the art would not combine them." However, Examiner submits that the particular teachings of the secondary references that are relied upon are functionally compatible with the primary reference. For example, Examiner submits that the concept and desirability of code density and code size is not functionally incompatible with a domain flow execution engine. For example, Examiner submits that the concept and desirability of minimizing memory bandwidth for sparse matrices is not functionally incompatible with a domain flow execution engine. Applicant across pages 9-10 argues: “As described previously, there is a significant difference between a domain flow execution engine (e.g., the presently claimed invention) and a stored program machine (e.g., used in Lumsdaine and Curry). Therefore, the teachings of Lumsdaine and Curry are not compatible with the teachings of Omtzigt. More specifically, in a stored program machine, the kernels are simple programs that can execute and request data from memory. Furthermore, the kernel in a stored program machine manages and schedules a linear sequence of instructions. In a domain flow engine, kernels are a collective state that programs the 'personality' of the compute fabric, so that each processing element has sufficient state to be autonomous when data arrives. Moreover, a domain flow execution engine kernel orchestrates a network of independent, concurrent processes connected by data flows. This is understood by one skilled in the art, and thus, it is also understood that although the same term "kernel" is used, there are clear differences between a kernel in a domain flow execution engine and a kernel for a stored program machine. Therefore, a kernel in a stored program machine cannot perform the same tasks as a domain flow execution engine kernel. In other words, the tasks a kernel performs in a stored-program machine are fundamentally different from those of a domain-flow execution engine. While a stored-program kernel is a low-level, privileged resource manager for a single computer, a domain-flow engine orchestrates high-level tasks and data across systems. As also described previously, in domain flow, it is a priori worked out how to configure the distributed data path as part of a kernel, whereas in a stored program machine, each instruction can do anything it wants, request far fetched data through cache hierarchy or network, jump to new subroutines, take an interrupt, etc. The whole control mechanism is different, and therefore, Lumsdaine and Curry have nothing to do with the domain flow execution engine machine of the presently claimed invention. Since they have nothing to do with the domain flow execution engine, their teachings are inherently incompatible with the teachings of Omtzigt, and thus the combination is improper.” Examiner appreciates the explanation contrasting a domain flow execution engine and a stored program machine. However, Examiner submits that the claims as currently recited do not appear to preclude Lumsdaine and Curry from being relied upon in the manner of the above prior art rejections, even if Lumsdaine and Curry are not directed to the domain flow execution engine machine of the instant invention. Examiner submits that even if Lumsdaine and Curry are not directed to the domain flow execution engine machine of the instant invention, such does not preclude specific teachings from Lumsdaine and Curry from being compatible with the teachings of Omtzigt. For example, Examiner submits that index structures being used for sparse matrices to minimize bandwidth, as taught by Lumsdaine, are not inherently incompatible with a domain flow engine. For example, Examiner submits that parallelism and chaining, as taught by Curry, are not inherently incompatible with kernels of a domain flow engine. Examiner generally submits that teachings that are not disclosed as being incorporated in a particular environment are not necessarily inherently incompatible with that environment. Applicant on page 10 argues: “However, Lumsdaine does not teach, for example, a cache for recalling a previous kernel of a plurality of parallel kernels, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory. Additionally, Lumsdaine does not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.” However, Examiner relied upon further references to render obvious the aforementioned subject matter. Applicant on page 11 argues: “However, Curry does not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.” However, Examiner submits that Curry teaches the aforementioned subject matter, as cited. Examiner submits that the kernels can be reasonably considered to be “parallel” kernels, given that they are executed by SIMD (i.e., single-instruction-multiple-data) processors, which performs multiple computations in parallel. (Also note that Nordquist, cited as pertinent in a previous office action, similarly discloses parallel kernels, in disclosing data level parallelism in that data is processed in parallel computation units.) To any extent to which the kernels of the instant invention are parallel in a different manner, Examiner submits that the broadest reasonable interpretation of “parallel kernels” does not preclude Curry from being relied upon in the manner of the above prior art rejections. Applicant on page 11 argues: “Additionally, Curry does not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.” However, Examiner relied upon a further reference to render obvious the aforementioned subject matter. Applicant on page 11 argues: “However, Matick does not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory”. However, Examiner relied upon Curry to teach the aforementioned subject matter, as noted above. Applicant on page 11 argues: “Additionally, Matick does not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.” However, Examiner relied upon a further reference to render obvious the aforementioned subject matter. Applicant on page 12 argues: “However, as described above, Mohamed does not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.” Examiner notes that the Office Action relied upon Mohamed to render obvious the aforementioned limitation (in conjunction with other prior art references), rather than teach the entirety of the limitation by itself. Applicant on page 12 argues: "Applicants respectfully disagree that Mohamed teaches the limitation: wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer. Although Figure 6 of Mohamed mentions an instruction packet having 32-bit instructions, Mohamed does not teach wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer." Examiner notes that the Office Action relied upon Mohamed to render obvious the aforementioned limitation (in conjunction with other prior art references), rather than teach the entirety of the limitation by itself. Applicant on page 12 argues: "Additionally, Mohamed, like Lumsdaine and Curry, is a stored program machine which is fundamentally different from the presently claimed invention, a domain flow execution engine. It is not proper to combine the teachings related to a stored program machine with a domain flow execution engine." However, Examiner submits that it is not fundamentally improper to apply a teaching from a reference directed to a stored program machine to a reference directed to a domain flow execution engine. For example, Examiner submits that the concept and desirability of code density and code size is not functionally incompatible with a domain flow execution engine. Applicant on page 12 argues: “As described above, Omtzigt, Lumsdaine, Curry, Matick, Mohamed and their combination do not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.” Applicant on page 13 further argues: “As also described above, Omtzigt, Lumsdaine, Curry, Matick, Mohamed and their combination do not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.” Applicant on page 13 argues: “As described above, Omtzigt, Lumsdaine, Curry, Matick and their combination do not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.” Applicant on page 14 argues: “As described above, Omtzigt, Lumsdaine, Curry, Matick, Glasco and their combination do not teach, for example, wherein the plurality of parallel kernels are chained to avoid serializing intermediate data to and from the memory.” However, Examiner submits that Curry teaches the aforementioned subject matter, as noted above. Applicant on page 12 argues: “Additionally, Omtzigt, Lumsdaine, Curry, Matick, Mohamed and their combination do not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.” Applicant on page 13 further argues: “Omtzigt, Lumsdaine, Curry, Matick, Mohamed and their combination do not teach, for example, wherein the domain flow program comprises a density such that parallel computations are expressed in 100 bytes or fewer.” However, Examiner submits that Mohamed renders obvious the aforementioned limitation (in conjunction with other prior art references), as addressed above. Applicant on page 13 argues: "Additionally, Claim 36 has been amended to recite: wherein the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory." In view of the aforementioned amendment, Examiner is relying upon Ledbetter, JR. et al. and Weisser et al. Applicant on page 14 argues: “Again, Glasco is a stored program machine, not a domain flow execution engine (e.g., the presently claimed invention). Therefore, Glasco is not applicable to the presently claimed invention.” As addressed above, Examiner generally submits that it would not necessarily be the case that any given teaching disclosed in the context/environment of a stored program machine would be improper to combine with a domain flow execution engine. Applicant on page 14 argues: "Additionally, Glasco does not teach, for example, wherein the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory." In view of the aforementioned amendment, Examiner is relying upon Ledbetter, JR. et al. and Weisser et al. Applicant on page 14 argues: "Omtzigt, Lumsdaine, Curry, Matick, Glasco and their combination do not teach, for example, wherein the cache is further configured for accumulating page coherent data for writeback to dynamic random access memory." In view of the aforementioned amendment, Examiner is relying upon Ledbetter, JR. et al. and Weisser et al. Applicant on page 15 argues: “Again, Martin is a stored program machine, not a domain flow execution engine (e.g., the presently claimed invention). Therefore, Martin is not applicable to the presently claimed invention.” Applicant on page 15 argues: “As described above, Omtzigt, Lumsdaine, Curry, Matick, Martin and their combination do not teach, for example, a first set of processing elements implemented as an asynchronous execution pipeline.” As addressed above, Examiner generally submits that it would not necessarily be the case that any given teaching disclosed in the context/environment of a stored program machine would be improper to combine with a domain flow execution engine. Applicant on page 16 argues: “Applicants note that this rejection involves seven (7) references. Applicants understand that there is no specific limit to the number of references able to be used for an obviousness rejection, but it should be worth reconsidering an obviousness rejection when seven (7) references are needed.” Applicant on page 16 argues: "Applicants note that this rejection involves seven (7) references. Applicants understand that there is no specific limit to the number of references able to be used for an obviousness rejection, but it should be worth reconsidering an obviousness rejection when seven (7) references are needed." Applicant on page 17 further argues: "Applicants note that this rejection involves seven (7) references. Applicants understand that there is no specific limit to the number of references able to be used for an obviousness rejection, but it should be worth reconsidering an obviousness rejection when seven (7) references are needed." Examiner appreciates Applicant’s understanding that there is no specific limit to the number of references able to be used for an obviousness rejection and also understands Applicant’s general concern. However, in the instant case, Examiner submits that the claimed subject matter of claim 49 is wide-ranging in scope (encompassing distinct concepts that may be, but do not necessarily have to be, used together, including, but not limited to, hashing algorithms, a domain flow execution engine per se, sparse matrices, parallel kernels, chained kernels, and caching), such that use of seven references is reasonably commensurate with this scope. For example, Examiner submits that the inventive concept related to sparse matrix optimization is distinct enough from the inventive concept related to chaining kernels that the reliance of respective references to teach the aforementioned inventive concepts in a same claim, rather than relying on just a single reference, would not be a tip of an iceberg that reflects a deeper issue with the rejection. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEITH E VICARY whose telephone number is (571)270-1314. The examiner can normally be reached Monday to Friday, 9:00 AM to 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached at (571)270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KEITH E VICARY/ Primary Examiner, Art Unit 2183
Read full office action

Prosecution Timeline

Show 19 earlier events
Oct 02, 2025
Response after Non-Final Action
Nov 05, 2025
Non-Final Rejection mailed — §103, §112
Jan 05, 2026
Response Filed
Jan 30, 2026
Final Rejection mailed — §103, §112
Mar 24, 2026
Response after Non-Final Action
Apr 28, 2026
Request for Continued Examination
May 01, 2026
Response after Non-Final Action
Aug 12, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748596
ARRAY PROCESSOR HAVING AN INSTRUCTION SEQUENCER INCLUDING A PROGRAM STATE CONTROLLER AND LOOP CONTROLLERS
2y 1m to grant Granted Sep 29, 2026
Patent 12743280
PARALLEL INSTRUCTION DEMARCATOR
2y 4m to grant Granted Sep 22, 2026
Patent 12699569
SELF-PROVISIONING AND FLEXIBLE HARDWARE ACCELERATOR ARCHITECTURE
2y 2m to grant Granted Aug 04, 2026
Patent 12688397
PARTITIONABLE DIGITAL HARDWARE SYSTEM FOR IMPLEMENTING RECURRENT NEURAL NETWORKS
5y 2m to grant Granted Jul 21, 2026
Patent 12663994
Supporting Multiple Vector Lengths with Configurable Vector Register File
2y 11m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

9-10
Expected OA Rounds
58%
Grant Probability
98%
With Interview (+40.3%)
3y 10m (~7m remaining)
Median Time to Grant
High
PTA Risk
Based on 698 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month