DETAILED ACTION
This office action is in response to amendments filed on 05/11/2026.
Claims 1-5, 7-8, 10-14, 16, and 18-20 have been amended. Claims 1-20 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Rejections Under 35 USC § 112(b):
In light of applicant’s amendments to the claims (pg. 2-6), the rejections under 35 USC § 112(b) have been withdrawn.
Prior Art Rejections:
Applicant's arguments regarding the prior art rejections have been fully considered, but they are not persuasive, or are moot because the new ground of rejection does not rely on the reference applied in the prior rejection of record for the teaching or matter specifically challenged in the argument.
Applicant argues (pg. 10-11) that the cited references Mills and Mital do not disclose or suggest the three input data broadcasting modes recited in the amended independent claim (default, multicast, simulcast). Applicant specifically argues that while Mital may disclose two broadcast modes in paragraph 0007 and another broadcast mode in paragraph 0033, two of these three broadcast modes operate on instructions rather than input data. Examiner respectfully disagrees with the assertion that one of the broadcast modes disclosed in Mital’s paragraph 0007 operates solely on channel instructions. In the first broadcast mode in 0007, which corresponds to the claimed multicast mode, “the memory manager slices the data set into data set chunks evenly spread across a cluster of components.” Spreading input data set chunks across the component clusters amounts to broadcasting input data. As applicant acknowledges, in the second broadcast mode in 0007, which corresponds to the claimed default mode, “the memory manager… broadcasts the data set to every cluster.” Thus, Mital discloses both a multicast mode of the input data and a default broadcast mode of the input data. Examiner additionally notes that, as can be seen in the rejection below, the Zbiciak reference has been brought in to teach the simulcast mode recited in the amended independent claim.
The prior art rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over
Mills, U.S. Patent Application Publication US-20190340501-A1 (published 11/07/2019), in view of
Mital et al. (hereinafter Mital), U.S. Patent Application Publication US-20230120227-A1 (filed 10/18/2022) and
Zbiciak et al. (hereinafter Zbiciak), U.S. Patent Application Publication US-20170249150-A1 (published 08/31/2017).
Regarding Claim 1,
Mills teaches A system-on-a-chip circuit, comprising: (0021: “FIG. 2 is a block diagram illustrating components in device 100… the device 100 may include, among other components, image sensor 202, system-on-a chip (SOC) component 204…”)
a neural processor circuit, comprising: (0027: “SOC component 204 may include, among other subcomponents… neural processor circuit 218…”)
a plurality of neural engines, at least one of the plurality of neural engines configured to perform computational operations on input data with variable dimensions; and (0039: “Neural processor circuit 218 is a configurable circuit that performs neural network operations on the input data based at least on kernel data 340. For this purpose, neural processor circuit 218 may include, among other components, neural task manager 310, a plurality of neural engines 314A through 314N…” 0058: “A work unit is a portion of the input data having a size that produces output values that fit into accumulator 414 of neural engine 314… work units can be shaped to one of 16×16, 32×8, 64×4, 128×2 or 256×1 dimension.” The neural processor circuit includes a plurality of neural engines which operate on work units (i.e. input data) of variable dimensions.)
a data processor circuit coupled to the plurality of neural engines, the data processor circuit configured to provide multiple modes of data broadcasting [comprising a multicast mode of the input data and a simulcast mode of the input data], wherein the input data is broadcasted from the data processor circuit to the plurality of neural engines [via one or more broadcast buses]; (0041: “…the neural task manager 310 sends rasterizer information to the components of the neural processor circuit 218 to enable each of the components to track, retrieve or process appropriate portions of the input data…” 0065: “…rasterizer 718 in data buffer 318 broadcasts in sequence work units for processing by the neural engines 314…” 0043: “Data buffer 318 may be operated in a broadcast mode where data input data of all input channels are fed to all neural engines 314 or in a unicast mode where data input data of a subset of input channels are fed to each neural engine 314.” The neural task manager, data buffer, and rasterizer (i.e. data processor circuit) broadcast work units (i.e. input data) to the neural engines, and can be operated in broadcast or unicast mode (i.e. provide multiple modes of data broadcasting).)
a central processor unit coupled to the neural processor circuit, the central processor unit configured to: (0027: “SOC component 204 may include, among other subcomponents…a central processor unit (CPU) 208…”)
Mills does not appear to explicitly disclose the remaining features of claim 1.
However, Mital teaches the multiple modes of data broadcasting comprising a multicast mode of the input data (0007: “The memory manager configured to when a data size of a data set from an AI-based processing model layer using the AI processor is larger than a weight size, the memory manager slices the data set into data set chunks evenly spread across a cluster of components, broadcasts channel instructions from the AI-based processing model layer to every cluster of components, and processes the data set chunk in the cluster of components according to the channel instructions of the AI-based processing model layer.” Slicing the input data set into data set chunks and broadcasting each data set chunk to a different component cluster is a multicast mode (see the explanation of multicast mode in specification paragraph 0088 of the instant application).)
a data processor circuit for broadcasting data via one or more broadcast buses (0030: “Within the AI processor 110, the scheduler is responsible for sending data to each of the multiple ALUs [arithmetic logic units] connected to it via the broadcast bus for parallel processing.” The scheduler (i.e. data processor circuit) broadcasts data via broadcast bus.)
the plurality of neural engines being assigned to multiple clusters processing different work units of the input data simultaneously by the multicast mode [or the simulcast mode], (0007: “The artificial intelligence processor can have multiple clusters of components including multiple arithmetic logic units each configured to have one or more computing engines to perform the computations for the AI system… the memory manager slices the data set into data set chunks evenly spread across a cluster of components, broadcasts channel instructions from the AI-based processing model layer to every cluster of components, and processes the data set chunk in the cluster of components according to the channel instructions of the AI-based processing model layer.” The neural computing engines each process different input data set chunks (i.e. work units of the input data) simultaneously. The neural computing engines assigned to a common input data set chunk are a cluster.)
the multicast mode [and the simulcast mode] of the input data being different from a default broadcast mode of the input data with the plurality of neural engines being assigned to a single cluster for the default broadcast mode of the input data; (0007: “In addition, memory manager configured to when the data size of the data set is smaller than a weight size of the AI-based processing model layer, the memory manager slices the AI-based processing model layer into channel chunks, assigns a channel chunk to a channel cluster, broadcasts the data set to every cluster, and processes the data set chunk according to channel instructions of the channel chunk.” Slicing the input data set into channel chunks and broadcasting the full input data set to all components is a default broadcast mode (see the explanation of default mode in specification paragraph 0088 of the instant application). The neural computing engines each process the full input data set (i.e. are all assigned to a single cluster).)
determine, based in part on a neural network description, a data broadcasting mode and an input data dimension configuration mode; (0029: “A compiler for the AI processor 110 uses a descriptor/instruction set with specific instructions crafted to efficiently handle various operations for neural networks… The descriptor/instruction set includes categories of descriptors/instructions including, for example… Data descriptors/instructions…” 0007: “The memory manager configured to when a data size of a data set from an AI-based processing model layer using the AI processor is larger than a weight size, the memory manager slices the data set into data set chunks evenly spread across a cluster of components, broadcasts channel instructions from the AI-based processing model layer to every cluster of components, and processes the data set chunk in the cluster of components according to the channel instructions of the AI-based processing model layer. In addition, memory manager configured to when the data size of the data set is smaller than a weight size of the AI-based processing model layer, the memory manager slices the AI-based processing model layer into channel chunks, assigns a channel chunk to a channel cluster, broadcasts the data set to every cluster, and processes the data set chunk according to channel instructions of the channel chunk.” Based on the size of the data (i.e. neural network description), the system determines whether to partition the input data by spatial dimension or by channel (i.e. determines an input data dimension configuration mode) and determines whether to broadcast the entire dataset or data chunks to each component cluster (i.e. determines a data broadcasting mode).)
generate one or more task descriptors, a task descriptor indicating the data broadcasting mode comprising the multicast mode [and the simulcast mode] of the input data and the input data dimension configuration mode, the multicast mode [and the simulcast mode] of the input data being different from the default broadcast mode of the input data; (See the previously cited portions of 0007. The AI processor determines instructions (i.e. task descriptors) for the component clusters which indicate whether the data is partitioned by spatial dimension or by channel (i.e. the input data dimension configuration mode) and whether the entire input dataset (default) or input data chunks (multicast) are broadcast (i.e. the data broadcasting mode comprising the multicast mode of the input data). As explained above, the multicast mode is different from the default mode.)
instruct, using the task descriptors, the data processor circuit to broadcast the input data to the plurality of neural engines according to the determined data broadcasting mode; and (See the previously cited portion of 0007 and 0030. The instructions (i.e. task descriptors) are broadcast to the component clusters where the scheduler (i.e. data processor circuit) sends input data (either the entire dataset or a data chunk, in accordance with the data broadcasting mode) to the plurality of components (i.e. neural engines).)
instruct, using the task descriptors, a neural engine to perform the computational operations according to the determined input data dimension configuration mode. (0027: “The two or more clusters of components connect to a broadcast bus for the memory manager 132 to broadcast a same instruction to the two or more clusters of components at a same time to evenly divide a computation across the two of more clusters of components so that each cluster of components performs a same computation but on a different portion of data…” Components (i.e. neural engines) receive instructions (i.e. task descriptors) to perform computations (i.e. computational operations) on partitions of input data (i.e. according to the input data dimension configuration mode).)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Mills and Mital. Mills teaches a neural processing circuit which partitions data into work units for processing by neural engines. Mital teaches dynamically determining how to partition and broadcast data to clusters for neural computation based on neural network descriptors. One of ordinary skill would have motivation to combine Mills and Mital in order to “efficiently do computations for AI systems, such as neural networks, as well as have a scalable architecture to adapt to most Artificial Intelligence (AI) networks, as well as optimize memory accesses and allocation” (Mital, 0025).
Mills and Mital do not appear to explicitly disclose a simulcast mode of data broadcasting.
However, Zbiciak teaches a simulcast mode of data broadcasting. (0005-0006: “This invention is a streaming engine employed in a digital signal processor… The streaming engine includes for each stream an address generator which produces addresses of data elements and a steam head register which stores data elements next to be supplied to functional units for use as operands. The two streams share two memory ports. A toggling preference of stream to port ensures fair allocation. The arbiters permit one stream to borrow the other's interface when the other interface is idle. Thus one stream may issue two memory requests, one from each memory port, if the other stream is idle. This spreads the bandwidth demand for each stream across both interfaces, ensuring neither interface becomes a bottleneck.” Examiner notes that, per specification paragraph 0092 of the instant application, simulcast mode is interpreted as a variant of multicast mode in which two work units are fetched and broadcast simultaneously when bandwidth permits. As shown above, Zbiciak teaches issuing memory requests for (i.e. fetching) and streaming (i.e. broadcasting) two data elements simultaneously when another stream is idle (i.e. when bandwidth is available).)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine Mills, Mital, and Zbiciak. Mills teaches a neural processing circuit which partitions data into work units for processing by neural engines. Mital teaches dynamically determining how to partition and broadcast data to clusters for neural computation based on neural network descriptors. Zbiciak teaches maximizing the efficiency of a digital data processor by fetching and streaming two data elements simultaneously when bandwidth is available. One of ordinary skill would have motivation to combine Mills, Mital, and Zbiciak because Zbiciak’s opportunistic dual data broadcasting “increases the available bandwidth to the functional units” (Zbiciak, 0078).
Regarding Claim 2, Mills, Mital, and Zbiciak teach The system-on-a-chip circuit of claim 1, as shown above.
Mills, Mital, and Zbiciak also teach wherein the simulcast mode is a function of the multicast mode. (Examiner notes that, as shown above in regard to claim 1, Mital teaches the multicast mode, and Zbiciak teaches modifying a broadcast mode to operate on two data elements simultaneously (i.e. the simulcast mode). Thus, in combination, Mital and Zbiciak teach the simulcast mode as a function of the multicast mode.)
Regarding Claim 3, Mills, Mital, and Zbiciak teach The system-on-a-chip circuit of claim 2, as shown above.
Mital also teaches wherein the multicast mode further comprises:
assigning each of the plurality of neural engines to a cluster; (0007: “The artificial intelligence processor can have multiple clusters of components including multiple arithmetic logic units each configured to have one or more computing engines to perform the computations for the AI system…”)
broadcasting a different portion of the input data associated with the computational operation to each cluster to generate an output value; (0007: “the memory manager slices the data set into data set chunks evenly spread across a cluster of components, broadcasts channel instructions from the AI-based processing model layer to every cluster of components, and processes the data set chunk in the cluster of components according to the channel instructions of the AI-based processing model layer.” 0026: “Note, at least one or more of the clusters of ALUs has an output that connects to its neighboring cluster.”)
and performing, at each cluster, the computational operations. (0027: “…each cluster of components performs a same computation but on a different portion of data…”)
Regarding Claim 4, Mills, Mital, and Zbiciak teach The system-on-a-chip circuit of claim 3, as shown above.
Mital also teaches wherein the simulcast mode further comprises:
broadcasting each block of the input data to different clusters; and (0007: “the memory manager slices the data set into data set chunks evenly spread across a cluster of components, broadcasts channel instructions from the AI-based processing model layer to every cluster of components, and processes the data set chunk in the cluster of components according to the channel instructions of the AI-based processing model layer.”)
performing, at each cluster, the computational operations. (0027: “…each cluster of components performs a same computation but on a different portion of data…”)
Mills teaches fetching consecutive blocks of the input data from a buffer; (0065: “…rasterizer 718 in data buffer 318 broadcasts in sequence work units for processing by the neural engines 314.” The rasterizer broadcasts (i.e. fetches) in sequence work units (i.e. consecutive blocks of input data) from the data buffer.)
Regarding Claim 5, Mills, Mital, and Zbiciak teach The system-on-a-chip circuit of claim 1, as shown above.
Mital also teaches wherein the input data dimension configuration mode comprises:
a first input data dimension configuration mode, which increases a first dimension of a block of the input data, and (0038: “When a data size of the data set is larger than a processing model layer for processing the data set, the memory manager 132 is configured to slice the data set into data set chunks.” Partitioning the dataset into chunks by spatial dimension is a first input data dimension configuration mode which results in a larger channel dimension in each chunk (i.e. increases a first dimension of a block of input data).)
a second input data dimension configuration mode, which increases a second dimension of a block of the input data. (0040: “When the data size of the data set is smaller than the processing model layer, the memory manager 132 is configured to slice the processing model layer into channel chunks.” Partitioning the dataset into chunks by channel is a second input data dimension configuration mode which results in larger spatial dimensions in each chunk (i.e. increases a second dimension of a block of input data).)
Regarding Claim 6, Mills, Mital, and Zbiciak teach The system-on-a-chip circuit of claim 3, as shown above.
Mital also teaches wherein a total number of clusters is determined based in part on the determined input data configuration mode. (0052: “The memory manager 132 can scale an amount of instances of the clusters to perform the computations for the AI system via a user configurable register transfer language parameter fed into the compiler at compile time (Block 606).” The amount of instances of clusters (i.e. total number of clusters) is scaled (i.e. determined) to perform the computations on the partitioned input data (i.e. based on the input data dimension configuration mode).)
Regarding Claim 7, Mills, Mital, and Zbiciak teach The system-on-a-chip circuit of claim 1, as shown above.
Mital also teaches wherein the neural network description includes the computational operations and a number of output channels. (0029: “For example, the compiler for the AI processor 110 uses a descriptor/instruction set with specific instructions crafted to efficiently handle various operations, addressing modes, data types, ability to address memory locations, etc., for neural networks. These neural networks can have sparse weights, manipulate one or more dimensional data, e.g., height, width, and channels and other dimensions such as images/frames per second… The descriptor/instruction set includes categories of descriptors/instructions including, for example… Data descriptors/instructions (used for both input and output)…” 0041: “In another example… The data is pretty low but the model because of the number of channels is pretty large, so what the compiler cooperating with the scheduler does is divide these 576 channels into the four clusters, so that each of the clusters is going to generate 144 channels.” The compiler uses the descriptor/instruction set (i.e. neural network description) which includes operations for neural networks (i.e. computational operations) and data descriptors such as number of output channels.)
Regarding Claim 8, Mills, Mital, and Zbiciak teach The system-on-a-chip circuit of claim 1, as shown above.
Mital also teaches wherein [a buffer] is configured to operate in the simulcast mode based in part on the input data dimension configuration mode and the computational operations. (Examiner notes that, per specification paragraph 0092 of the instant application, simulcast mode is a function of multicast mode, and thus the determination to operate in simulcast mode is based on the determination to operate in multicast mode. 0007: “The memory manager configured to when a data size of a data set from an AI-based processing model layer using the AI processor is larger than a weight size, the memory manager slices the data set into data set chunks evenly spread across a cluster of components, broadcasts channel instructions from the AI-based processing model layer to every cluster of components, and processes the data set chunk in the cluster of components according to the channel instructions of the AI-based processing model layer.” The system is configured to broadcast each input data set chunk to a different component cluster (i.e. operate in the multicast mode) when each cluster is performing the same channel instructions (i.e. based on the computational operations) on a different partition of data (i.e. based on the input data dimension configuration mode).)
Mills teaches broadcasting via a buffer (0065: “…rasterizer 718 in data buffer 318 broadcasts in sequence work units for processing by the neural engines 314.”)
Regarding Claim 9, Mills, Mital, and Zbiciak teach The system-on-a-chip circuit of claim 1, as shown above.
Mills also teaches wherein the input data is sized by a rasterizer, the rasterizer configured to divide the input data based on the input data dimension configuration mode determined [by a compiler]. (0066: “To perform their functions, each of rasterizers 714, 718, 720, 722 receives task information 710 indicating how the input data and/or kernel data are to be segmented and to be handled by each component of the neural processor circuit 218. The task information includes information about particulars of the current layer (e.g., dimensions of input and output data, dimension of an associated kernel, types of padding at the boundaries of input data).” Rasterizers perform segmentation (i.e. sizing and division) of input data based on information including dimensions of input data (i.e. the input data dimension configuration mode).)
Mital teaches the input data dimension configuration mode is determined by a compiler. (0039: “At a user selectable threshold, a size/amount of the data is compared to a size/amount of the weights will transition thru that threshold and change from moving data a single time and broadcasting weights over to moving weights one time and broadcasting data. At this point, the memory manager sub-module 132 of the compiler will switch the AI processor 110 from frame sub-layering across clusters over to channel sub-layering across clusters.” Switching from frame sub-layering to channel sub-layering is determining the input data dimension configuration mode, and this switch is performed by the compiler.)
Claims 10-17 are method claims containing substantially the same elements as system claims 1-7 and 9, respectively. Mills, Mital, and Zbiciak teach the elements of claims 1-7 and 9, as shown above.
Claims 18-20 are device claims containing substantially the same elements as system claims 1-3, respectively. Mills, Mital, and Zbiciak teach the elements of claims 1-3, as shown above.
Mills also teaches An electronic device, comprising: a system memory configured to store a neural network; and a system-on-a-chip circuit coupled to the memory, the system-on-a-chip circuit configured to: (0021: “FIG. 2 is a block diagram illustrating components in device 100, according to one embodiment. Device 100 may perform various operations including image processing. For this and other purposes, the device 100 may include, among other components, image sensor 202, system-on-a chip (SOC) component 204, system memory 230…” 0016: “Embodiments of the present disclosure relate to… performing neural network operations.” 0025: “System memory 230 is a component for storing instructions for execution by SOC component 204 and for storing data processed by SOC component 204.” The device performs neural network operations, and includes an SOC coupled to a system memory which stores instructions and data for SOC execution (i.e. the memory is configured to store a neural network).)
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BENJAMIN M ROHD whose telephone number is (571)272-6445. The examiner can normally be reached Mon-Thurs 8:00-6:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/B.M.R./Examiner, Art Unit 2147
/ERIC NILSSON/Primary Examiner, Art Unit 2151