Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 13 objected to because of the following informalities: Applicant claims status twice, “the synchronization status of the other processing circuitry status indicates to the first circuitry…” Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
Claims 8-14 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claims 8-14 claim a first circuitry and a second circuitry. The first and second circuitry are never described in the specification. The roles of the first circuitry, other circuitry and the second circuitry are changed in claims 8-14, in a way that is never described in the specification. This appears to be a drafting mistake.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
Claims 8-14 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The “other processing circuitry”, “other data”, “the one or more other sections” and other “other[s]” in claim 8-14, in a way, lack antecedent basis because they are referred to at “the other” throughout a claim where are there are multiple different components with the same functional description – e.g. the first processing circuitry, the second processing circuitry and the other processing circuitry. Further, this circuit setup is never described in the specification so it is not possible to understand what Applicant is attempting to claim. For purposes of examination, claims 8-14 will be interpreted similar to claims 1-7, because any other interpretation is not described and not defined.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-19 are rejected under 35 U.S.C. 103 as being unpatentable over US20200073702A1 to Han US20230315655A1 and to Choquette et al.
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over US20200073702A1 to Han, US20230315655A1 to Choquette et al and US20220391701A1 to Tanaka et al.
Han teaches claim 1. A system comprising:
a first processing core; and
a second processing core coupled to the first processing core and configurable to: (Han abs “The system can include: a task manager; and a plurality of cores…”)
prior to executing a current layer of a neural network, determine a synchronization status of the first processing core with respect to a previous layer of the neural network; (Han para 59 “task manager 310 can determine that convolution computation for a first iteration is finished.” Han para 23 “In a CNN model, such operations can repeat for a plurality of layers. During these operations, feature output map 106 of a previous layer serves as an input data of a next layer.” Han para 62 “after barrier instructions from all cores have been received, high level program 404 can be executed. … and provide an ‘collect&reduce’ instruction to the core to collect data (e.g., intermediate feature maps) from the other cores and to reduce them into one feature map, which can be used as an input for a next layer.”
execute the current layer of the neural network based on data from the previous layer computed by the first processing core and by the second processing core; and (Han para 62 “after barrier instructions from all cores have been received, high level program 404 can be executed. … and provide an ‘collect&reduce’ instruction to the core to collect data (e.g., intermediate feature maps) from the other cores and to reduce them into one feature map, which can be used as an input for a next layer…. Core 2024 can possess all intermediate feature maps for further processing.” The next layer is Applicant’s current layer.)
upon executing the current layer of the neural network, update the first processing core with a synchronization status of the second processing core with respect to the current layer of the neural network. (Han para 59 “task manager 310 can determine that convolution computation for a first iteration is finished.” It does this determination for each layer, including the next/current layer. See signal in fig. 4.)
Han doesn’t teach that the second core determines the flag or updates the first core.
However, Choquette teaches a second processing core coupled to the first processing core and configurable to:… determine a synchronization status… upon executing …, update the first processing core with a synchronization status… (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile. The consumer issues a remote arrive operation on the producer barrier to indicate the data in the tile was read, and then waits on the local consumer barrier until the tile becomes live again.” Consumer is the claimed second core, and producer is the claimed first core.)
Choquette, the claims and Han all distribute processing. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to provide inter-core communication, like in Choquette, because “the latency can be significantly reduced…” Choquette para 78.
Han teaches claim 2. The system of claim 1, wherein the data computed by the first processing core corresponds to a first section of an input tensor associated with the first processing core; and
wherein the data computed by the second processing core corresponds to a second section of the input tensor associated with the second processing core. (Han para 24 “input data 102 sized H×W×C can be divided into, for example, four portions 112, 114, 116, and 118, and distributed to four processing cores (e.g., cores 132, 134, 136, and 138). Therefore, each processing core is responsible for ¼ of input data 102.”)
Han teaches claim 3. The system of claim 2 wherein the first section of the input tensor comprises a non-overlapping section of the input tensor with respect to the second section of the input tensor. (Han para 24 “input data 102 sized H×W×C can be divided into, for example, four portions 112, 114, 116, and 118…” Fig. 1b shows the input is non-overlapping.)
Choquette teaches claim 4. The system of claim 1, wherein the first processing core is configurable to write the synchronization status of the first processing core to shared memory, and (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile….” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core.)
wherein the second processing core, to determine the synchronization status of the first processing core, is configurable to read the synchronization status of the first processing core from the shared memory. (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile ….” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core. Choquette para 91 “splitting the synchronization functionality across two barriers: one co-located with the consumer, and one co-located with the producer.”)
Choquette teaches claim 5. The system of claim 4 wherein the second processing core, to update the first processing core with the synchronization status of the second processing core, is configurable to write the synchronization status of the second processing core to the shared memory for consumption by the first processing core. (Choquette para 96 “The producer performs one or more remote combined store and arrive operations, each of which fills a tile with line data and performs a remote arrive operation on the corresponding barrier and calls wait() on the local producer barriers to block until the tile becomes dead (i.e. has no data) again.” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core. Choquette para 91 “splitting the synchronization functionality across two barriers: one co-located with the consumer, and one co-located with the producer.”))
Choquette teaches claim 6. The system of claim 5 wherein each respective instance of the synchronization status of the first processing core indicates that the first processing core has completed executing (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile.” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core. Choquette para 91 “splitting the synchronization functionality across two barriers: one co-located with the consumer, and one co-located with the producer.”))
Choquette doesn’t limit itself to layers of a neural network.
However, Han teaches the previous layer of the neural network. (Han para 23 “In a CNN model, such operations can repeat for a plurality of layers. During these operations, feature output map 106 of a previous layer serves as an input data of a next layer.”)
Choquette teaches claim 7. The system of claim 6 wherein each respective instance of the synchronization status of the second processing core indicates that the second processing core has completed executing ( Choquette para 91 “In this scheme, the “c-barrier” (a barrier co-located with the consumer) is waited on (e.g. using wait() instruction) by the consumer and arrived on (e.g. using arrive()) by the producer. Conversely the “p-barrier” (a barrier co-located with the producer) is arrived on by the consumer and waited on by the producer.” Choquette para 91 “splitting the synchronization functionality across two barriers: one co-located with the consumer, and one co-located with the producer.” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core.)
Choquette is not limited to neural networks.
However, Han teaches the current layer of the neural network. (Han para 23 “In a CNN model, such operations can repeat for a plurality of layers. During these operations, feature output map 106 of a previous layer serves as an input data of a next layer.”)
Han teaches claim 8. Processing circuitry comprising:
first circuitry configurable to: (Han abs “The system can include: a task manager; and a plurality of cores…”)
prior to executing a current layer of a neural network, determine a synchronization status of other processing circuitry with respect to a previous layer of the neural network; and(Han para 59 “task manager 310 can determine that convolution computation for a first iteration is finished.” Han para 23 “In a CNN model, such operations can repeat for a plurality of layers. During these operations, feature output map 106 of a previous layer serves as an input data of a next layer.” Han para 62 “after barrier instructions from all cores have been received, high level program 404 can be executed. … and provide an ‘collect&reduce’ instruction to the core to collect data (e.g., intermediate feature maps) from the other cores and to reduce them into one feature map, which can be used as an input for a next layer.”)
upon executing the current layer of the neural network, update the other processing circuitry with a synchronization status of the processing circuitry with respect to the current layer of the neural network; and (Han para 62 “after barrier instructions from all cores have been received, high level program 404 can be executed. … and provide an ‘collect&reduce’ instruction to the core to collect data (e.g., intermediate feature maps) from the other cores and to reduce them into one feature map, which can be used as an input for a next layer…. Core 2024 can possess all intermediate feature maps for further processing.” The next layer is Applicant’s current layer.)
second circuitry configurable to execute the current layer of the neural network based on data from the previous layer computed by the second circuitry and other data from the previous layer computed by the other processing circuitry. (Han para 62 “after barrier instructions from all cores have been received, high level program 404 can be executed.” Han para 63 “ task manager 310 can send to all cores “resume” instructions to resume the convolution computation for a second iteration.” The second iteration is executing the Han’t next layer which is Applicant’s current layer.)
Han doesn’t teach that the second core determines the flag or updates the first core.
However, Choquette teaches first circuitry configurable to:… determine a synchronization status… upon executing …, update the other processing circuitry with a synchronization status… (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile. The consumer issues a remote arrive operation on the producer barrier to indicate the data in the tile was read, and then waits on the local consumer barrier until the tile becomes live again.” Consumer is the claimed second core, and producer is the claimed first core.)
Choquette, the claims and Han all distribute processing. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to provide inter-core communication, like in Choquette, because “the latency can be significantly reduced…” Choquette para 78.
Han teaches claim 9. The processing circuitry of claim 8 wherein the data computed by the second circuitry, with respect to the previous layer of the neural network, corresponds to a section of an input tensor associated with the processing circuitry, and wherein the other data computed by the other processing circuitry corresponds to one or more other sections of the input tensor associated with the other processing circuitry. (Han para 24 “input data 102 sized H×W×C can be divided into, for example, four portions 112, 114, 116, and 118, and distributed to four processing cores (e.g., cores 132, 134, 136, and 138). Therefore, each processing core is responsible for ¼ of input data 102.”)
Han teaches claim 10. The processing circuitry of claim 9 wherein the section of the input tensor and the one or more other sections of the input tensor each comprise a non-overlapping section of the input tensor with respect to each other section of the input tensor. (Han para 24 “input data 102 sized H×W×C can be divided into, for example, four portions 112, 114, 116, and 118…” Fig. 1b shows the input is non-overlapping.)
Choquette teaches claim 11. The processing circuitry of claim 8 wherein the other processing circuitry writes the synchronization status of the other processing circuitry to shared memory, and (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile….” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core.) wherein the first circuitry, to determine the synchronization status of the other processing circuitry, reads the synchronization status of the other processing circuitry from the shared memory. (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile ….” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core. Choquette para 91 “splitting the synchronization functionality across two barriers: one co-located with the consumer, and one co-located with the producer.”)
Choquette teaches claim 12. The processing circuitry of claim 11 wherein the first circuitry, to update the other processing circuitry with the synchronization status of the processing circuitry, writes the synchronization status of the processing circuitry to the shared memory for consumption by the other processing circuitry. (Choquette para 96 “The producer performs one or more remote combined store and arrive operations, each of which fills a tile with line data and performs a remote arrive operation on the corresponding barrier and calls wait() on the local producer barriers to block until the tile becomes dead (i.e. has no data) again.” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core. Choquette para 91 “splitting the synchronization functionality across two barriers: one co-located with the consumer, and one co-located with the producer.”)
Choquette teaches claim 13. The processing circuitry of claim 8 wherein the synchronization status of the other processing circuitry status indicates to the first circuitry that the other processing circuitry has completed executing (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile.” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core. Choquette para 91 “splitting the synchronization functionality across two barriers: one co-located with the consumer, and one co-located with the producer.”))
Choquette doesn’t limit itself to layers of a neural network.
However, Han teaches the previous layer of the neural network. (Han para 23 “In a CNN model, such operations can repeat for a plurality of layers. During these operations, feature output map 106 of a previous layer serves as an input data of a next layer.”)
Choquette teaches claim 14. The processing circuitry of claim 8 wherein the synchronization status of the processing circuitry indicates to the other processing circuitry that the second circuitry has completed executing ( Choquette para 91 “In this scheme, the “c-barrier” (a barrier co-located with the consumer) is waited on (e.g. using wait() instruction) by the consumer and arrived on (e.g. using arrive()) by the producer. Conversely the “p-barrier” (a barrier co-located with the producer) is arrived on by the consumer and waited on by the producer.” Choquette para 91 “splitting the synchronization functionality across two barriers: one co-located with the consumer, and one co-located with the producer.” The tile is the shared memory. Consumer is the claimed second core, and producer is the claimed first core.)
Choquette is not limited to neural networks.
However, Han teaches the current layer of the neural network. (Han para 23 “In a CNN model, such operations can repeat for a plurality of layers. During these operations, feature output map 106 of a previous layer serves as an input data of a next layer.”)
Han teaches claim 15. A system, comprising;
multiple processing cores; and Han abs “The system can include: a task manager; and a plurality of cores…”)
wherein each of the multiple processing cores is configurable to:
prior to executing a current layer of a neural network:
read, from the (Han para 23 “In a CNN model, such operations can repeat for a plurality of layers. During these operations, feature output map 106 of a previous layer serves as an input data of a next layer.” Han para 62 “after barrier instructions from all cores have been received, high level program 404 can be executed. … and provide an ‘collect&reduce’ instruction to the core to collect data (e.g., intermediate feature maps) from the other cores and to reduce them into one feature map, which can be used as an input for a next layer.”)
determine, based on the synchronization status of each of the one or more other processing cores, that the one or more processing cores have completed processing with respect to the previous layer of the neural network; (Han para 59 “task manager 310 can determine that convolution computation for a first iteration is finished.”)
execute the current layer of the neural network based on data from the previous layer computed by the processing core and other data from the previous layer computed by the one or more other processing cores; and (Han para 62 “after barrier instructions from all cores have been received, high level program 404 can be executed. … and provide an ‘collect&reduce’ instruction to the core to collect data (e.g., intermediate feature maps) from the other cores and to reduce them into one feature map, which can be used as an input for a next layer…. Core 2024 can possess all intermediate feature maps for further processing.” The next layer is Applicant’s current layer.)
upon executing the current layer of the neural network, write a synchronization status of the processing core to the shared portion of the memory, wherein the synchronization status of the processing core indicates that the processing core has completed processing with respect to the current layer of the neural network. (Han para 59 “task manager 310 can determine that convolution computation for a first iteration is finished.” It does this determination for each layer, including the next/current layer. See signal in fig. 4.)
Han doesn’t teach that the second core determines the flag or updates the first core.
However, Choquette teaches each of the multiple processing cores configurable to:… determine a synchronization status… upon executing …, write a synchronization status of the processing core to the shared portion of the memory … (Choquette para 96 “Each consumer, which was waiting on a consumer barrier until the corresponding tile became live (i.e. have data to consume), would be released in response to the producer’s remote arrive operation on the consumer’s local barrier and proceeds to load (“LDS”) the data from the tile. The consumer issues a remote arrive operation on the producer barrier to indicate the data in the tile was read, and then waits on the local consumer barrier until the tile becomes live again.” Consumer is the claimed second core, and producer is the claimed first core. The tile is the shared memory)
Choquette, the claims and Han all distribute processing. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to provide inter-core communication, like in Choquette, because “the latency can be significantly reduced…” Choquette para 78.
Han teaches claim 16. The system of claim 15 wherein the memory includes non-shared portions corresponding to the multiple processing cores, and wherein the system further comprises a memory controller configurable to perform a direct memory access (DMA) transfer of the other data from one or more portions of the non-shared portions of the memory, corresponding to the one or more other processing cores, to a portion of the non-shared portions of the memory corresponding to the processing core. (Han DMA unit 208, para 32 “memory controller 206 can manage read/write data coming from outside chip communication system 202 (e.g., from DMA unit 208 or a DMA unit corresponding with another NPU) or from inside chip communication system 202 (e.g., from a local memory in core 2024 via a 2D mesh controlled by a task manager of global manager 2022).”)
Han teaches claim 17. The system of claim 16 wherein the data computed by the processing core corresponds to a section of an input tensor associated with the processing core. (Han para 24 “input data 102 sized H×W×C can be divided into, for example, four portions 112, 114, 116, and 118, and distributed to four processing cores (e.g., cores 132, 134, 136, and 138). Therefore, each processing core is responsible for ¼ of input data 102.”)
Han teaches claim 18. The system of claim 17 wherein the other data computed by the one or more other processing cores corresponds to one or more other sections of the input tensor associated with the one or more other processing cores. (Han para 24 “input data 102 sized H×W×C can be divided into, for example, four portions 112, 114, 116, and 118, and distributed to four processing cores (e.g., cores 132, 134, 136, and 138). Therefore, each processing core is responsible for ¼ of input data 102.”)
Han teaches claim 19. The system of claim 18 wherein the section of the input tensor and the one or more other sections of the input tensor each comprise a section of the input tensor that is non-overlapping with respect to each other section of the input tensor. (Han para 24 “input data 102 sized H×W×C can be divided into, for example, four portions 112, 114, 116, and 118…” Fig. 1b shows the input is non-overlapping.)
Han teaches claim 20. The system of claim 15 further comprising a general-purpose central processing unit configurable to distribute execution of the neural network to the multiple processing cores, and (Han Global manager 2022 fig. 2a) wherein each of the multiple processing cores comprises a (Han fig. 2a cores)
Han doesn’t teach a DSP.
However, Tanaka teaches multiple processing cores comprises a digital signal processor (DSP). (Tanaka para 40-41 “each of the computers for distributed processing 10-1 and 10-2 includes: a plurality of accelerators 13-1 to 13-4, …. the accelerators 13-1 to 13-4, for example, a GPU, a field-programmable gate array (FPGA), a digital signal processor (DSP),…”)
Han, the claims and Tanaka are all directed to distributed computing. It would have been obvious to a person having ordinary skill in the art, at the time of filing, to use an accelerator like a DSP because Han says that it’s cores are accelerators such as a GPU or NPU, (Han para 29) and a DSP is just another well-known type of accelerator. Tanaka para 41.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Austin Hicks whose telephone number is (571)270-3377. The examiner can normally be reached Monday - Thursday 8-4 PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AUSTIN HICKS/Primary Examiner, Art Unit 2142