Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office Action has been withdrawn pursuant to 37 CFR 1.114.
Detailed Action
This Non-Final Office Action is responsive to Applicants’ RCE submission dated 6/5/26. Claims 1-10, 12, 14-15, and 18 are presently pending, of which claims 1, 9, and 15 are independent.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office Action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-10, 12, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Patent Application Publication No. 2022/0292334 (“Pandey”) in view of Non-Patent Literature “DMA: Its Importance and Applications Explained” (“Solomon”).
Regarding claim 1, PANDEY teaches A method comprising:
storing a first activation input tensor1 to one or more first storage devices coupled to a microcontroller unit (MCU) by a memory bus (as discussed in relation to FIGs. 4A-4B, [0048] teaching “In some instances a model 108 and/or an input data into model 108 may be too large to fit into high-speed cache 136 (or any other internal memory) of edge computing device 130. A conventional approach to performing MLM operations in such instances is to load network parameters (e.g., weights and biases) and/or input data from memory 134 to cache 136 prior to performing computations where such parameters and/or input data are used. Parameters and/or input data are then overwritten in the next iteration until all operations are complete. Since the same data may be used in multiple computations, it not unusual to load the same parameters and/or input data into cache 136 multiple times. As depicted schematically in FIG. 4A, in some implementations of the present disclosure, a MLM may be factorized into two or more partitions of such a size that the network parameters and/or input data of each partition can fit into cache 136.”, and where the input data / “input data grid” as taught (e.g., element 402 in the FIGs.) is akin to a stored activation input tensor as recited as stored in memory (i.e., “one or more first storage devices”), and where a memory as taught per [0048] is shown as element 134 in FIG. 1A and would be understood by one of ordinary skill in the art to be situated in a manner as coupled to a bus for the depicted edge computing device relative to controller elements for that same device, such that the bus is a conduit between the memory and the controller elements and the cache for example (see [0078] for a teaching of a “bus interconnect” serving as a conduit between memory and processor elements), and where the controller elements shown per FIG. 1A and as discussed in the reference generally are equivalent to the recited MCU);
partitioning the stored first activation input tensor into a plurality of tensor segments ([0048]-[0049] discussing the partitioning of the input grid such as the inputs can be loaded into cache in successive iterations to perform the computations necessary to evaluate a machine learning model, e.g. [0049] “… For example, network parameters may be able to fit in cache memory but the input data may be too large to be loaded at once. In such instances, the input data may be factorized into smaller portions that can be loaded into cache memory. Partition A may include operations that use input data A 410 to compute output data A 411 (e.g., a first portion of output data grid 406) and partition(s) B (C, etc.) may include operations that use input data B 420 to compute output data B 421 (output data C 431, etc.). After input data A 410 has been loaded to cache 136 and output data A 411 has been computed, input data B 420 (and, similarly, input data into subsequent partitions) may be loaded into cache 136 and output data B 421 may be computed. In some implementations, network parameters of neuron 404 (and other neurons that are not shown explicitly) may similarly be partitioned into portions and loaded into cache 136 together with the inputs of the corresponding partitions.”); and
sequentially loading individual tensor segments of the stored first activation input tensor from the one or more first storage devices to one or more second storage devices via a sequence of direct memory access (DMA) transactions (in reference to the scheme introduced and discussed in [0048]-[0049], as noted just prior, [0051] further clarifies that “ Input buffer 452 may store all N input values {Ij} that are loaded from a system memory 460 during cycle 1 of direct memory access (DMA) operations.”) … the one or more second storage devices being integrated in the MCU (where the cache, loaded by way of the DMA operations, per [0048]-[0049] and [0051], is understood to be part of a kernel/chip per [0044] but especially [0077] (“For example, the first memory device may be cache (e.g., high-speed memory on the processor chip).”)) with processing circuitry to apply one or more activation functions associated with the first activation input tensor ([0048]-[0049]’s taught memory management scheme facilitates the “performing [of] MLM operations”, with [0049] getting into how the factorized information including the inputs and weights are managed to compute output data for the model, which the Examiner understands to include/rely upon activations (i.e., the evaluation of an activation function as recited) resulting from computations featuring those same inputs and weights (see [0028] for an explicit teaching of the activation function being applied: “For example, operations performed on a set of input data (e.g., partitioned among multiple neurons) by various neurons may represent one layer, operations performed on the output of that layer may represent another layer, and so on. A neuron may represent any set of computations that takes two or more input numbers and produces an output number (e.g., via weight multiplication, bias addition, application of an activation function, etc.).”)).
Regarding the DMA transactions addressed just above, Applicants’ claim further clarifies that they are executed independently of processing circuitry of the MCU. Pandey does not teach this, and rather, the Examiner relies upon SOLOMON to teach what Pandey otherwise lacks, see e.g., Solomon’s 1st paragraph discussing the use of “separate chip” DMA controllers which can move blocks of data without using/burdening the primary CPU element, which is further clarified in the following paragraph that starts at the bottom of the 1st page and continues onto the 2nd page.
Like Pandey, Solomon relates to the use of direct memory access in a larger architecture to implement a more efficient memory management scheme that improves computational throughput/efficiency. Hence, the references are similarly directed and therefore analogous. It would have been obvious to one of ordinary skill in the art to implement Pandey’s DMA feature with one or more DMA controllers, per Solomon, and with a reasonable expectation of success, so as to improve the computation performance of not just the machine learning task at hand but also the edge device’s processing capability.
Regarding claim 2, Pandey in view of Solomon teach the method of claim 1, as discussed above. The aforementioned references teach the additional limitations wherein sequentially loading individual tensor segments of the stored first activation input tensor further comprises:
loading a first tensor segment of the individual tensor segments to a first portion of the one or more second storage devices local to processing circuitry to apply a first activation function to the first tensor segment of the individual tensor segments and completing of loading of a second tensor segment to a second portion of the one or more second storage devices local to the processing circuitry to apply the first activation function subsequent to commencement of application of the first activation function to the first tensor segment (Pandey: [0048]-[0049] discussing the numerous partitions A, B, C, etc., which are subject to evaluation with the corresponding weights by some activation function for the machine learning model that is subjected to the taught memory management scheme, where each partition is subject to an activation function (per [0028], for example, as noted just above in relation to claim 1) and because there are many such iterations of this, then it reads on the successive/iterative loading of inputs and weights as discussed in these cited-to portions of the reference in a manner that reads on these further limitations).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 3, Pandey in view of Solomon teach the method of claim 1, as discussed above. The aforementioned references teach the additional limitations further comprising:
storing weights in at least one of the one or more first storage devices and partitioning the stored weights according to the one or more activation functions and sequentially loading individual stored weights to the one or more second storage devices local to the processing circuitry to apply the one or more activation functions (Pandey: [0048]-[0049] clarify that weights are subject to the same successive/iterative loading and computation process as discussed above in relation to claim 1, and [0028] clarifying the involvement of an activation function in each iteration).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 4, Pandey in view of Solomon teach the method of claim 3, as discussed above. The aforementioned references teach the additional limitations wherein at least one of the one or more activation functions comprises a dot product (Pandey’s [0028]: “For example, operations performed on a set of input data (e.g., partitioned among multiple neurons) by various neurons may represent one layer, operations performed on the output of that layer may represent another layer, and so on. A neuron may represent any set of computations that takes two or more input numbers and produces an output number (e.g., via weight multiplication, bias addition, application of an activation function, etc.).”).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 5, Pandey in view of Solomon teach the method of claim 1, as discussed above. The aforementioned references teach the additional limitations further comprising:
storing a second activation input tensor to at least one of the one or more first storage devices and partitioning the stored second activation input tensor into a plurality of second tensor segments and sequentially loading individual second tensor segments of the stored second activation input tensor to the one or more second storage devices to processing circuitry to apply at least one of the one or more activation functions associated with the second activation input tensor (Pandey: a parameter of a particular machine learning model is an activation function, per [0078], and it follows that the taught framework as applied to different models would then involve different activation functions, even if the same underlying circuitry and general optimization schemes stay the same; hence, the mappings as applied to claim 1 are the same for a different model that is subject to the same modified framework of Pandey in view of Solomon, even if the particular inputs, weights, activation functions, and so forth differ as a function of the particular model as parameterized).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 6, Pandey in view of Solomon teach the method of claim 5, as discussed above. The aforementioned references teach the additional limitations wherein at least one of the one or more activation functions comprises an operation to additively combine an associated tensor segment of the plurality of tensor segments and an associated tensor segment of the plurality of second tensor segments ([0056]: “During subsequent stages, additional portions of the input values and the corresponding portions of weights are used to compute additional portions of the output values. For example, during the first cycle of stage k (cycle (k−1)M+1), the k-th portion of the input values {Ij}(k) is loaded to input buffer 452 and the k-th portion of the weights {W1j}(k) is loaded into weight buffer 454. The computation logic 458 then computes the portion O1 (k) of the output value O1 by adding O1 (k) to the accumulator that stores the previously computed sum O1 (1)+O1 (2)+ . . . +O1 (k−1). During subsequent cycles of stage k, further portions of weights {Wij}(k) are loaded to weight buffer 454 and new portions Oi (k) of the output values Oi are computed. After completion of all n stages, M final values {Oi} are stored in the output buffer 456.”, where the taught accumulation/summing discussed here is equivalent to Applicants’ recitation of an additive combination).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 7, Pandey in view of Solomon teach the method of claim 1, as discussed above. The aforementioned references teach the additional limitations wherein the one or more first storage devices are external to the MCU (Pandey: FIG. 1A showing memory element 134 as being external to CPU 132).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 8, Pandey in view of Solomon teach the method of claim 1, as discussed above. The aforementioned references teach the additional limitations wherein sequentially loading individual tensor segments of the stored first activation input tensor to memories local to processing circuitry comprises executing the sequence of DMA transactions (Pandey: where the cache, loaded by way of the DMA operations, per [0048]-[0049] and [0051], is understood to be part of a kernel/chip per [0044] but especially [0077] (“For example, the first memory device may be cache (e.g., high-speed memory on the processor chip).”)).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 9, the claim includes the same or similar limitations as discussed above in relation to claim 1, and is therefore rejected under the same rationale.
Regarding claim 10, Pandey in view of Solomon teach the computing device of claim 9, as discussed above. The aforementioned references teach the additional limitations wherein:
the one or more first storage devices comprise a dynamic random access memory (DRAM) device, a flash memory device, or a combination thereof; and the one or more second storage devices comprise at least one static random access memory (SRAM) device (Pandey: [0098]: “The implementations of methods, hardware, software, firmware or code set forth above may be implemented via instructions or code stored on a machine-accessible, machine readable, computer accessible, or computer readable medium which are executable by a processing element. “Memory” includes any mechanism that provides (i.e., stores and/or transmits) information in a form readable by a machine, such as a computer or electronic system. For example, “memory” includes random-access memory (RAM), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; magnetic or optical storage medium; flash memory devices; electrical storage devices; optical storage devices; acoustical storage devices, and any type of tangible machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).”, from which the Examiner believes it possible to understand an edge computing device as might be used in relation to FIG. 1A’s framework for example to feature memory implementations one or both of SRAM or DRAM).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 12, Pandey in view of Solomon teach the computing device of claim 9, as discussed above. The aforementioned references teach the additional limitations further comprising:
circuitry to store weights in at least one of the one or more first storage devices and circuitry to partition the stored weights according to the one or more activation functions and circuitry to sequentially load individual stored weights to the one or more second storage devices (Pandey: FIG. 1A’s element 132 (CPU) which is understood to perform some of the steps discussed above in relation to claim 1 and Pandey’s [0048]-[0049] for example, as is a kernel/chip per [0044] but especially [0077] (“For example, the first memory device may be cache (e.g., high-speed memory on the processor chip)”) understood to perform the sequential loading step to what is equivalent to the second storage device as recited / cache).
The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 15, the claim includes the same or similar limitations as discussed above in relation to claim 1, and is therefore rejected under the same rationale.
Claims 14 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Pandey in view of Solomon and further in view of U.S. Patent Application Publication No. 2022/0012563 (“Castro Gonzalez”).
Regarding claim 14, Pandey in view of Solomon teach the computing device of claim 9, as discussed above in relation to claim 1. The aforementioned references do not teach the additional limitation wherein the memory bus comprises a Serial Peripheral Interface (SPI). Rather, the Examiner relies upon CASTRO GONZALEZ to teach what Pandey etc. otherwise lack, see e.g., Castro Gonzalez’s comparable framework for edge inference of neural network models ([0062]-[0065]) as facilitated by a framework that features SPI bus elements specifically ([0193]).
Both references are similarly directed with respect to providing frameworks and teachings to facilitate edge computing of machine-learning/neural network models. Hence, they are analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate well-known bus elements as contemplated by Castro Gonzalez into Pandey’s framework, with a reasonable expectation of success, on the basis that various types of bus elements as broadly contemplated by the primary reference can satisfy its needs including a specific bus element as contemplated by the secondary reference.
Regarding claim 18, the claim includes the same or similar limitations as discussed above in relation to claim 14, and is therefore rejected under the same rationale.
Response to Arguments
Applicants’ arguments with respect to the pending claims, as received 6/5/26, have been carefully considered but are moot in view of the newly-formulated grounds of rejection. The Examiner notes that Pandey is still relied upon, but however using different teachings and mappings than before. If Applicants still believe the reference is not apt, or that its combination with other references are unsound, the Examiner recommends a telephonic interview to discuss any purported discrepancies. Given the present breadth in scope of Applicants’ claim, the reference is believed to be a fair and reasonable reading on it as presented in the obviousness rejection.
Conclusion
10. The prior art made of record and not relied upon is considered pertinent to Applicants’ disclosure:
Non-Patent Literature “Enhancing ARM-based embedded SoC performance in high-bandwidth human-interface applications” (“Wilbrink”)
US 2022/0067536 (“Datla”), which the Examiner believes is relevant to Applicants’ claimed invention:
[0012] As shown in FIG. 2, a method S100 is executed by a neural network at a processor system 100 including a shared memory unit 140, a set of primary memory units 120, and a set of processing units 150. The method S100 includes storing, in the shared memory unit 140: a first weight tensor at a first source address, the first weight tensor including a set of weight tensor partitions; and a first input tensor at a second source address, the first input tensor larger than the first weight tensor in Block S110. The method S100 also includes broadcasting the first input tensor from the second source address to a first relative destination address in the set of primary memory units 120 in Block S120. The method S100 additionally includes, for each processing unit 150 in the set of processing units 150: transferring a weight tensor partition in the set of weight tensor partitions from the first source address to a first destination address in the primary memory unit 120 of the processing unit 150 in Block S130; and, at the processing unit 150, generating an output tensor partition of a first output tensor based on the first input tensor and the weight tensor partition in Block S140. The method S100 further includes storing, in the shared memory unit 140: a second weight tensor at a third source address; a second input tensor at a fourth source address, the second input tensor: including a set of input tensor partitions; and smaller than the second weight tensor in Block S150. The method S100 further includes broadcasting the second weight tensor from the third source address to a second relative destination address in the set of primary memory units 120 in Block S160. The method S100 further includes, for each processing unit 150 in the set of processing units 150: transferring an input tensor partition in the set of input tensor partitions from the fourth source address to a second destination address in the primary memory unit 120 of the processing unit 150 in Block S170; and, at each processing unit 150 in the set of processing units 150, generating an output tensor partition of a second output tensor based on the second weight tensor and the input tensor partition in Block S180.
[0069] Generally, the processor system 100 can store a layer of an artificial neural network (e.g., such as a CNN), including an input tensor and a weight tensor, in response to a set of scheduled transfer operations between the main memory of the processor system 100 and the shared memory unit 140 of the processor. The processor system 100 can store an input tensor and a set of weight tensor partitions for an input-broadcast output-stationary dataflow or the processor system 100 can store a weight tensor and a set of input tensor partitions for a weight-broadcast output-stationary dataflow. More specifically, the processor system 100 can store, in the shared memory unit 140: a first weight tensor at a first source address, the first weight tensor comprising a set of weight tensor partitions; and a first input tensor at a second source address, the first input tensor larger than the first weight tensor in Block S110. Additionally, in Block S150, the processor system 100 can store, in the shared memory unit 140: a second weight tensor at a third source address; a second input tensor at a fourth source address, the second input tensor: including a set of input tensor partitions; and smaller than the second weight tensor. Thus, the processor system 100 can access inputs for a scheduled layer from the shared memory unit 140.
[0075] Generally, the processor system 100 can broadcast the first input tensor from a source address to a relative destination address in the set of primary memory units 120 in Block S120. More specifically, the processor system 100 can: issue a first read request for a first input tensor at a second source address, via a direct memory access core 110; in response to the first read request, load the first input tensor into a data buffer 112; via the direct memory access core 110, issue a first write request specifying the first relative destination address to the broadcast subsystem 130; and, via a set of broadcast buses 132 of the processor system 100, simultaneously transfer the first input tensor from the data buffer 112 to the first relative destination address in each primary memory unit 120 in the set of primary memory units 120. Thus, in Block S120, the processor system 100 broadcasts an input tensor to each of a set of primary memory units 120 in the processor system 100.
[0076] In order to compute an output for the input-broadcast layer of the artificial neural network, the processor system 100 also distributes a set of weight partitions, each partition including a subsection of the weight tensor for the input-broadcast layer. More specifically, the processor system 100 can, for each processing unit 150 in the set of processing units 150, transfer a weight tensor partition in the set of weight tensor partitions from the first source address to a first destination address in the primary memory unit 120 of the processing unit 150 in Block S130. Thus, the processor system 100 can broadcast an input tensor of a layer of an artificial neural network and serially unicast a set of weight partitions, thereby making available input data and weight partition data to each processor unit in the set of processor units.
[0077] Upon receiving both input data and weight partition data at each primary memory unit 120 in the set of target primary memory units 120, the processor system 100 can, via a set of processor units corresponding to the target set of primary memory units 120, calculate an output partition based on the input tensor and the weight partition for each primary memory unit 120 in the set of target primary memory units 120. More specifically, the processor system 100 can: at the processing unit 150, generate an output tensor partition of a first output tensor based on the first input tensor and the weight tensor partition in Block S140. Thus, by repeatedly executing Blocks S110, S120, S130, and S140 over successive layers of an artificial neural network, the processor system 100 can continually execute the scheduled parallel process.
[0080] In order to compute an output for the weight-broadcast layer of the artificial neural network, the processor system 100 also distributes a set of input partitions, each input partition including a subsection of the input tensor for the weight-broadcast layer. More specifically, the processor system 100 can, for each processing unit 150 in the set of processing units 150, transfer an input tensor partition in the set of input tensor partitions from the fourth source address to a second destination address in the primary memory unit 120 of the processing unit 150 in Block S170. Thus, the processor system 100 can broadcast a weight tensor of a layer of an artificial neural network and serially unicast a set of input tensor partitions, thereby making available input data and weight partition data to each processor unit in the set of processor units for the weight-broadcast layers.
[0081] Upon receiving both input partition data and weight data at each primary memory unit 120 in the set of target primary memory units 120 the processor system 100 can, via a set of processor units corresponding to the target set of primary memory units 120, calculate an output partition based on the input tensor partition and the weight tensor for each primary memory unit 120 in the set of target primary memory units 120. More specifically, the processor system 100 can: at the processing unit 150, generate an output tensor partition of a first output tensor based on the second weight tensor and the input tensor partition in Block S180. Thus, by repeatedly executing Blocks S150, S160, S170, and S180 over successive layers of an artificial neural network, the processor system 100 can continually execute the scheduled parallel process.
US 12235756 (“Aga”)
11. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHOURJO DASGUPTA whose telephone number is (571)272-7207. The examiner can normally be reached M-F 8am-5pm CST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 571 272 4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHOURJO DASGUPTA/Primary Examiner, Art Unit 2144
1 The Examiner construes the term “activation input tensor” in accordance with Applicants’ definition as provided in [0018] of the published specification, e.g., “an expression of one or more activation input values according to a particular structure, dimension and/or format.”