DETAILED ACTION
This office action addresses Applicant’s response filed on 31 March 2026. Claims 1-20 are pending.
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-3, 16, and 18 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Wilder (US 2010/0274990).
Regarding claim 1, Wilder discloses a method for accelerating an execution of computational loops on an integrated circuit (¶44), the method comprising: programming a finite state machine (FSM) based on a loop iteration parameter comprising a number of computation cycles of a computational loop to be executed by a computational circuit (¶18);
at runtime, executing the FSM based on a start signal, wherein executing the FSM includes: (i) generating, by the FSM, a plurality of control signals, each control signal corresponding to a respective computation cycle of the computational loop based on the loop iteration parameter; and (ii) controlling, by the FSM, an operation of the computational circuit executing the computational loop based on a transmission of the plurality of control signals to the computational circuit (¶¶18-20, 97);
wherein, in response to execution of the FMS based on the start signal, instruction fetch and decode operations corresponding to the computational loop are disabled such that the computational loop is executed without fetching or decoding instruction sequences corresponding to the computational loops during execution of the FSM (¶¶16, 23).
Regarding claim 2, Wilder discloses that the FSM is controllably connected to a plurality of processing cores, each of the plurality of processing cores having at least one computational circuit (¶44).
Regarding claim 3, Wilder discloses that at runtime, the FSM is executed without performing fetches of computational loop instructions (¶16).
Regarding claim 16, Wilder discloses that at runtime, the FSM generates the plurality of controls signals causing an execution of an N-way multiply accumulate with computation weights and computation input data, wherein: N relates to a number of distinct multiply accumulate circuits concurrently executing a distinct computational loop, and N is greater than one (¶¶14, 65).
Regarding claim 18, Wilder discloses a method comprising: programming a finite state machine (FSM) based on one or more FSM initialization parameters, wherein the one or more FSM initialization parameters include a loop iteration parameter comprising a number of multiply-accumulate computation cycles of a convolutional loop (¶¶14, 17);
at runtime, implementing the FSM to enable one or more computations by: (i) generating, by the FSM, a plurality of convolutional loop control signals based on the loop iteration parameter; and (ii) controlling, by the FSM, an execution of a plurality of multiply-accumulate computation cycles of a multiply accumulator circuit (MAC) performing the convolutional loop based on transmitting the plurality of convolutional loop control signals until the number of multiply-accumulate computation cycles of the convolutional loop are completed (¶¶14, 17, 97, 98);
wherein, in response to implementation of the FMS based on the start signal, instruction fetch and decode operations corresponding to the convolutional loop are disabled such that the convolutional loop is executed without fetching or decoding instruction sequences corresponding to the multiply-accumulate computation cycles of the convolutional loop during implementation of the FSM (¶¶16, 23).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 4-7, 9, and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wilder in view of Ware (US 2022/0283779).
Regarding claim 4, Wilder does not appear to explicitly disclose programming the FSM based on a data movement parameter comprising at least one data movement instruction that, when executed, moves input data from a register file of a first processing core to data input ports of neighboring processing cores. Ware discloses these limitations (¶¶5, 7). It would have been obvious to persons having ordinary skill in the art before the effective filing date of the application to combine the teachings of Wilder and Ware, because doing so would have involved merely the routine combination of known elements according to known techniques, or the substitution of an element for a known equivalent, to produce merely the predictable results of correctly computing convolutions using a MAC pipeline. KSR Int’l Co. v. Teleflex Inc., 82 U.S.P.Q.2d 1385, 1395. Wilder discloses controlling a plurality of MAC units to perform repeated MAC computations. Ware teaches performing convolutions with repeated MAC computations, where the input data is moved between MAC units. The teachings of Ware are directly applicable to Wilder in the same way, so that Wilder would similarly move input data between MAC units to perform convolution operations using the MAC units.
Regarding claim 5, Wilder discloses that at runtime, the FSM is executed without performing fetches of data movement instructions (¶16).
Regarding claim 6, Wilder does not appear to explicitly disclose that the register file is associated with one or more data output ports of the first processing core and data input ports of the first processing core, wherein the data input ports of the first processing core are directly connected to data output ports of the neighboring processing cores; and executing the at least one data movement instruction causes the input data to rotate an angle from the data input ports of the first processing core to the one or more data output ports of the first processing core. Ware discloses these limitations (¶¶5, 7, 21). Motivation to combine remains consistent with claim 4.
Regarding claim 7, Wilder discloses that at runtime, executing the FSM causes an execution of the computational loop based on the loop iteration parameter (¶¶18, 97), but does not appear to explicitly disclose subsequently, executing the FSM causes an execution of one or more computational loops based on the loop iteration parameter and the data movement parameter. Ware discloses executing the FSM causes an execution of one or more computational loops based on the loop iteration parameter and the data movement parameter (¶¶4, 7, 12). Motivation to combine remains consistent with claim 4.
Regarding claim 9, Wilder does not appear to explicitly disclose that programming the FSM includes identifying, by the FSM, a distinct data movement control signal for each of the number of computation cycles of the computational loop based on the loop iteration parameter and a data movement parameter. Ware discloses these limitations (¶21). Motivation to combine remains consistent with claim 4.
Regarding claim 10, Wilder does not appear to explicitly disclose that controlling the operation of the computational circuit executing the computational loop includes transmitting, by the FSM, the distinct data movement control signal for each of the number of computation cycles of the computational loop until the number of computation cycles of the computational loop are completed. Ware discloses these limitations (¶21). Motivation to combine remains consistent with claim 4.
Claim(s) 8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wilder in view of Ware and Li (US 2018/0267809).
Regarding claim 8, Wilder discloses that at runtime, the FSM generates: a first set of control signals of the plurality of control signals for executing the computational loop based on the loop iteration parameter; and, a second set of control signals of the plurality of control signals for executing (a) the number of computation cycles of the computational loop (¶18). Wilder does not appear to explicitly disclose the second set of control signals of the plurality of control signals for executing (a) the number of computation cycles of the computational loop and (b) the at least one data movement instruction based on the loop iteration parameter and the data movement parameter; Ware discloses these limitations (¶¶4, 7, 12, 21). Motivation to combine remains consistent with claim 4.
Wilder does not appear to explicitly disclose that the generation of the second set of control signals is in response to completing the computational loop based on the loop iteration parameter; Li discloses these limitations (¶138). It would have been obvious to persons having ordinary skill in the art before the effective filing date of the application to combine the teachings of Wilder, Ware, and Li, because doing so would have involved merely the routine combination of known elements according to known techniques to produce merely the predictable results of executing multiple instructions. KSR Int’l Co. v. Teleflex Inc., 82 U.S.P.Q.2d 1385, 1395. Wilder discloses controlling execution of a repeated computation instruction. Ware teaches data movement instructions to control data flow between processing elements when executing computations. Li teaches that after completing execution of a repeated computation instruction, the processing elements execute subsequent repeated computation instructions. The teachings of Li are directly applicable to Wilder, so that Wilder would similarly execute subsequent repeated computation instructions after execution of a repeated computation is complete, in order to execute multiple repeated computation instructions.
Claim(s) 11-13 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wilder in view of Hou (US 2021/0240483).
Regarding claim 11, Wilder does not appear to explicitly disclose that programming the FSM includes encoding a starting memory address parameter to a start memory address register file accessible to one or more computational circuits controllable by the FSM. Hou discloses these limitations (Fig. 1; ¶22). It would have been obvious to persons having ordinary skill in the art before the effective filing date of the application to combine the teachings of Wilder and Hou, because doing so would have involved merely the routine combination of known elements according to known techniques to produce merely the predictable results of retrieving operands from memory. KSR Int’l Co. v. Teleflex Inc., 82 U.S.P.Q.2d 1385, 1395. Wilder discloses computing matrix operations by controlling processing cores comprising MACs. Hou teaches encoding starting memory address and filter size parameters so that the correct matrix operands for the computations are retrieved by the MACs. The teachings of Hou are directly applicable to Wilder in the same way, so that Wilder would similarly encode starting memory address and filter sizes to retrieve the correct operands from memory for MAC operation.
Regarding claim 12, Wilder does not appear to explicitly disclose that the starting memory address parameter comprises a register file pointer that points to a head of input data at a location within an n-dimensional memory stored within at least one processing core controllable by the FSM. Hou discloses these limitations (Fig. 1; ¶31). Motivation to combine remains consistent with claim 11.
Regarding claim 13, Wilder does not appear to explicitly disclose that programming the FSM includes encoding a convolution filter size parameter to a convolution register file of at least one processing core controllable by the FSM. Hou discloses these limitations (Fig. 1; ¶22). Motivation to combine remains consistent with claim 11.
Regarding claim 19, Wilder discloses that programming the FSM includes programming one or more iteration parameters at one or more iteration register files accessible to the FSM (¶95), but does not appear to explicitly disclose (i) programming a starting memory address parameter at a start memory address register file accessible to the MAC; (ii) programming a convolution filter size parameter at a convolution register file accessible to the MAC. Hou discloses these limitations (Fig. 1; ¶22). Motivation to combine remains consistent with claim 11.
Claim(s) 14 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wilder in view of Hou and Henry (US 2018/0276035).
Regarding claim 14, Wilder does not appear to explicitly disclose that the convolution filter size parameter comprises a value that maps to one of a plurality of distinct convolutional filter sizes for a given convolutional computation by a multiply accumulator circuit of the at least one processing core. Hou discloses these limitations (¶22). If Hou is found to be unclear regarding these limitations, Henry discloses the same (¶222). It would have been obvious to persons having ordinary skill in the art before the effective filing date of the application to combine the teachings of Wilder, Hou, and Henry, because doing so would have involved merely the routine combination of known elements according to known techniques to produce merely the predictable results of retrieving filter values from memory according to conventional filter sizes. KSR Int’l Co. v. Teleflex Inc., 82 U.S.P.Q.2d 1385, 1395. Wilder teaches controlling convolution operations. As discussed above with regard to claim 11, Hou teaches encoding filter size to correctly retrieve filter values from memory for convolution operations. Persons having ordinary skill in the art, reading Hou, would understand that defining the size of the filter in terms of row and column dimensions necessarily maps to distinct convolution filter sizes, as taught by Henry. The teachings of Hou and Henry are directly applicable to Wilder, so that Wilder would similarly define distinct filter dimensions to correctly retrieve filter values from memory.
Regarding claim 17, Wilder discloses that if a convolution filter size parameter of the FSM includes a value that maps to one of a plurality of distinct convolutional filter sizes that is greater than a 1x1 convolutional filter size, the FSM broadcasts input data pointed to by a starting memory address parameter to a collection of processing cores in neighboring proximity to the FSM (¶¶15, 78, 86). If Wilder is found to be unclear regarding the filter size parameter and starting memory address parameter, Hou discloses the same (¶22). If Wilder is found to be unclear regarding the plurality of distinct convolutional filter sizes, Henry discloses the same (¶222). Motivation to combine remains consistent with claim 14.
Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wilder in view of Li.
Regarding claim 15, Wilder does not appear to explicitly disclose that programming the FSM includes encoding the loop iteration parameter to a combination of distinct iteration register files of at least one processing core controllable by the FSM. Li discloses these limitations (claim 5; ¶¶8, 128). It would have been obvious to persons having ordinary skill in the art before the effective filing date of the application to combine the teachings of Wilder and Li, because doing so would have involved merely the routine combination of known elements according to known techniques to produce merely the predictable results of independently configuring and controlling processing elements. KSR Int’l Co. v. Teleflex Inc., 82 U.S.P.Q.2d 1385, 1395. Wilder discloses performing computations using a plurality of processing elements. Li teaches that each processing element receives configuration data for performing the computations. The teachings of Li are directly applicable to Wilder the same way, so that Wilder’s processing elements would similarly receive configuration data to allow independent control of the processing elements.
Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wilder in view of Ware and Hou.
Regarding claim 20, Wilder discloses a method for implementing finite state machine (FSM)-controlled convolutional computations on an integrated circuit (¶¶18, 44), the method comprising: configuring an FSM based on one or more FSM programming instructions, wherein the FSM controls: (a) computations of multiply accumulator circuits (MACs) of a plurality of distinct processing cores (¶44);
wherein configuring the FSM includes: (3) encoding an iteration value to at least one iteration register file associated with the FSM, wherein the iteration value identifies a number of cycles of a convolutional loop performed by at least one of the MACs (¶¶18, 95); and executing a Boolean switch based on the configuring of the FSM that starts an operation of the FSM for generating control signals to the MACs for automatically executing one or more distinct convolutional loops (¶¶18, 97);
wherein, in response to execution of the FMS based on execution of the Boolean switch, instruction fetch and decode operations corresponding to the computational loop are disabled such that the one or more distinct computational loop are executed without fetching or decoding instruction sequences corresponding to the one or more distinct computational loops during execution of the FSM (¶¶16, 23).
Wilder does not appear to explicitly disclose data movement operations of data ports of the plurality of distinct processing cores. Ware discloses a method for implementing finite state machine (FSM)-controlled convolutional computations on an integrated circuit (¶¶12, 18), the method comprising: configuring an FSM based on one or more FSM programming instructions, wherein the FSM controls: (b) data movement operations of data ports of the plurality of distinct processing cores (¶7). It would have been obvious to persons having ordinary skill in the art before the effective filing date of the application to combine the teachings of Wilder and Ware, because doing so would have involved merely the routine combination of known elements according to known techniques, or the substitution of an element for a known equivalent, to produce merely the predictable results of correctly computing convolutions using a MAC pipeline. KSR Int’l Co. v. Teleflex Inc., 82 U.S.P.Q.2d 1385, 1395. Wilder discloses controlling a plurality of MAC units to perform repeated MAC computations. Ware teaches performing convolutions with repeated MAC computations, where the input data is moved between MAC units. The teachings of Ware are directly applicable to Wilder in the same way, so that Wilder would similarly move input data between MAC units to perform convolution operations using the MAC units.
Wilder does not appear to explicitly disclose (1) encoding a starting memory address value to an address register file accessible to the MACs of the plurality of distinct processing cores, and (2) encoding a convolutional filter size to a convolutional register file associated with the FSM. Hou discloses these limitations (Fig. 1; ¶22). It would have been obvious to persons having ordinary skill in the art before the effective filing date of the application to combine the teachings of Wilder, Ware, and Hou, because doing so would have involved merely the routine combination of known elements according to known techniques to produce merely the predictable results of retrieving operands from memory. KSR Int’l Co. v. Teleflex Inc., 82 U.S.P.Q.2d 1385, 1395. Wilder discloses computing matrix operations by controlling processing cores comprising MACs. Hou teaches encoding starting memory address and filter size parameters so that the correct matrix operands for the computations are retrieved by the MACs. The teachings of Hou are directly applicable to Wilder in the same way, so that Wilder would similarly encode starting memory address and filter sizes to retrieve the correct operands from memory for MAC operation.
Response to Arguments
Applicant's arguments filed 31 March 2026 have been fully considered but they are not persuasive.
Applicant asserts that the claims have been amended to overcome Wilder; specifically, Applicant asserts that Wilder fails to teach per-cycle control signals where “each control signal corresponding to a respective computation cycle of the computational loop” and a ‘fetchless’ mode “disabling instruction fetch, disabling instruction decode, or executing a computational loop without fetching instruction sequences”. Remarks 2. The examiner disagrees. Regarding per-cycle control signals, Wilder explicitly states at ¶18: “In one particular embodiment, one of the control signals provided to the SIMD data processing circuitry identifies the number of iterations M required, and the state machine generates internal control signals which are altered dependent on the iteration being performed, and are used to select the input data elements and the single coefficient data element for each iteration” (emphasis added). Regarding fetchless mode, Wilder’s ¶¶16 and 23 explicitly disclose Applicant’s fetchless mode:
a single instruction can be used to cause the SIMD data processing circuitry to perform a plurality of iterations of a multiply-accumulate process determined by a scalar value provided as an input operand of that instruction, in order to directly produce a plurality of multiply-accumulate results. Since all of the data elements required for all of the specified iterations can be derived directly from the first and second vectors provided as input operands of the instruction, a significant reduction in energy consumption can be realised when compared with the known prior art techniques which require the execution of a program loop multiple times, with accesses to memory during each time through the loop. In particular, the invention provides a single instruction that can execute without further register or instruction reads in order to generate a plurality of multiply-accumulate results, saving significant energy consumption when compared with known prior art techniques.
…
To alleviate unnecessary power consumption resulting from such issues, in one embodiment the state machine determines the number of iterations M from the scalar value, and asserts a stall signal to one or more components of the data processing apparatus whilst at least one of the plurality of iterations are being performed. In one particular example, the stall signal is used to suspend instruction fetching whilst the stall signal is asserted.
(emphasis added). Notably, Wilder explicitly contemplates the advantage of reducing power consumption from exactly Applicant’s technique of eliminating further instruction fetching to execute looped instructions, by programming the FSM to control loop execution based on one instruction. Thus, contrary to Applicant’s assertions, Wilder clearly discloses the new limitations.
Applicant also asserts that the examiner’s inherency position is improper, and that the examiner “does not identify any disclosure in Wilder that requires disabling instruction fetch or decode operations, nor does the Office Action provide any technical reasoning demonstrating that such disabling necessarily occurs in Wilder. Instead, the rejection appears to assume that the presence of FSM-based control inherently results in the claimed functionality.” Remarks 2. The examiner disagrees. The examiner has not relied on inherency to reject the claims under §§102 or 103 using Wilder. As discussed above, Wilder explicitly discloses disabling instruction fetch and decode operations; the rejection nowhere “assume[s] that the presence of FSM-based control inherently results in the claimed functionality”, as asserted by Applicant.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ARIC LIN whose telephone number is (571)270-3090. The examiner can normally be reached M-F 07:30-17:00 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jack Chiang can be reached at 571-272-7483. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
19 August 2026
/ARIC LIN/ Examiner, Art Unit 2851
/JACK CHIANG/ Supervisory Patent Examiner, Art Unit 2851