Detailed Action
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is nonfinal and is in response to claims filed on 01/20/2023. Claims 1-20 are pending for examination.
Information Disclosure Statement
The Information Disclosure Statement (IDS) submitted on 01/14/2025 is in compliance with the provisions of 37 CFR 1.97, 1.98, and MPEP § 609. It has been placed in the application file, and the information referred to therein has been considered as to the merits.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Interpretation
Examiner notes that Claim 16 is directed to a method that recites conditional language. Claim 16 recites “performing the multiplication and reformatting operations on the corresponding input and weight data elements only if the difference is less than a difference threshold”.
The conditional nature of this claim language allows for an interpretation where any prior art meets the broadest reasonable interpretation of the claim without having the conditional language even occurring (and thus only the preceding limitations required by the prior art). For example, in claim 16, the “performing the multiplication and reformatting operations on the corresponding input and weight data elements” limitation is not required to be taught by the prior art. See MPEP 2111.04(II); see also Ex parte Schulhauser. Examiner notes that the broadest reasonable interpretation of the method of claim 16 require that the “performing the multiplication and reformatting operations” condition is not required to occur and thus the claim language ends. Examiner encourages claim amendments that specifically removes the conditional language of the claims and thus expressly has the claimed scenarios occur. Examiner respectfully reiterates that without changing the conditional nature of the claim, the method claims carry no patentable weight as noted above.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract ideas without significantly more.
With regards to claim 1, at step 1, the claim is directed to a machine, which is a statutory category of invention.
At Step 2A Prong 1, the examiner notes that the claim is directed to mental processes and/or mathematical concepts. The claim language has been reproduced below:
A circuit comprising: (mental process, evaluation)
a multiplier circuit configured to: (mental process, evaluation)
receive a signed mantissa of each data element of a plurality of input data elements and a plurality of weight data elements, and (mental process, evaluation)
generate a plurality of two's complement products by performing multiplication and reformatting operations on some or all of the signed mantissas of the plurality of input data elements and some or all of the signed mantissas of the plurality of weight data elements; (mathematical calculation)
a summing circuit configured to: (mental process, evaluation)
receive an exponent of each data element of the plurality of input data elements and the plurality of weight data elements, and (mental process, evaluation)
generate a plurality of sums by adding each exponent of the plurality of input data elements to each exponent of the plurality of weight data elements; (mathematical calculation)
a shifting circuit configured to (mental process, evaluation) shift each product of the plurality of products by an amount equal to a difference between a corresponding sum of the plurality of sums and a maximum sum; and (mathematical calculation)
an adder tree configured to (mental process, evaluation) generate a mantissa sum from the plurality of shifted products. (mathematical calculation)
Each of the non-bolded limitations are mental processes and/or mathematical calculations. The “A circuit comprising” limitation is an evaluation mental process that can be performed by choosing what the circuit comprises. The “a multiplier circuit configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier circuit is configured to do. The “receive a signed mantissa of each data element of a plurality” limitation is an evaluation mental process that can be performed by choosing what the circuit receives. The “generate a plurality of two's complement products” limitation is a mathematical calculation that can be performed by generating the plurality of two's complement products by hand using pen and paper. The “a summing circuit configured to” limitation is an evaluation mental process that can be performed by choosing what the summing circuit is configured to do. The “receive an exponent of each data element of the plurality of input data elements” limitation is an evaluation mental process that can be performed by choosing what the circuit receives. The “generate a plurality of sums” limitation is a mathematical calculation that can be performed by generating plurality of sums by hand using pen and paper. The “a shifting circuit configured to” limitation is an evaluation mental process that can be performed by choosing what the shifting circuit is configured to do. The “shift each product of the plurality of products by an amount equal” limitation is a mathematical calculation that can be performed by shifting each product by hand using pen and paper. The “an adder tree configured to” limitation is an evaluation mental process that can be performed by choosing what the adder tree is configured to do. The “generate a mantissa sum from the plurality of shifted products” limitation is a mathematical calculation that can be performed by generating the mantissa sum by hand using pen and paper.
At step 2A Prong 2, the additional elements are bolded above. The “receive” limitations, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “receive” in the context of the claim encompasses mere data gathering. The remaining additional elements amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f).
At Step 2B, the claim recites “receive a signed mantissa of each data element of a plurality of input data elements and a plurality of weight data elements”, “receive an exponent of each data element of the plurality of input data elements and the plurality of weight data elements”, and, per MPEP 2106.05(d) (Il), the courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity:
i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); and
iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93.
Regarding claim 13, It recites similar language as claim 1, and is rejected for at least the same reasons therein. Herein, claim 13 is directed towards the statutory category of a method, thus also satisfying step 1. Under steps 2A Prong 2 and 2B, the claim does not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claim 2, it is directed to mental processes and/or mathematical concepts. The “wherein the multiplier and summing circuits are configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier and summing circuits are configured to do. The “receive each data element of the plurality of input data elements and the plurality of weight data elements having a BF16 format” limitation is an evaluation mental process and mathematical relationship that can be performed by choosing the format of the inputs. The “the multiplier circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier circuit is configured to do. The “generate the plurality of products as 17-bit data elements” limitation is a mathematical calculation that can be performed by generating the plurality of products by hand using pen and paper and choosing the size of them. The “the summing circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the summing circuit is configured to do. The “generate the plurality of sums as nine-bit data elements” limitation is a mathematical calculation that can be performed by generating the plurality of sums by hand using pen and paper and choosing the size of them. The “the shifting circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the shifting circuit is configured to do. The “generate the plurality of shifted products as 21-bit data elements” limitation is a mathematical calculation that can be performed by generating the plurality of shifted products by hand using pen and paper and choosing the size of them. The “the adder tree is configured to” limitation is an evaluation mental process that can be performed by choosing what the adder tree is configured to do. The “generate the mantissa sum as a 25-bit data element” limitation is a mathematical calculation that can be performed by generating the mantissa sum by hand using pen and paper and choosing the size of them. Under step 2A Prong 2, The “receive” limitation, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “receive” in the context of the claim encompasses mere data gathering. The remaining additional elements amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). None of the additional elements regarding the generic computer components (i.e. the multiplier circuit, the summing circuit, the shifting circuit, the adder tree, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under Step 2B, the claim recites “wherein the multiplier and summing circuits are configured to receive each data element of the plurality of input data elements and the plurality of weight data elements having a BF16 format”, and, per MPEP 2106.05(d) (Il), the courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity:
i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); and
iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93.
With regards to claim 3, it is directed to mental processes and/or mathematical concepts. The “wherein the multiplier and summing circuits are configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier and summing circuits are configured to do. The “receive each data element of the plurality of input data elements and the plurality of weight data elements having a FP16 format” limitation is an evaluation mental process and mathematical relationship that can be performed by choosing the format of the inputs. The “the multiplier circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier circuit is configured to do. The “generate the plurality of products as 23-bit data elements” limitation is a mathematical calculation that can be performed by generating the plurality of products by hand using pen and paper and choosing the size of them. The “the summing circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the summing circuit is configured to do. The “generate the plurality of sums as six-bit data elements” limitation is a mathematical calculation that can be performed by generating the plurality of sums by hand using pen and paper and choosing the size of them. The “the shifting circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the shifting circuit is configured to do. The “generate the plurality of shifted products as 27-bit data elements” limitation is a mathematical calculation that can be performed by generating the plurality of shifted products by hand using pen and paper and choosing the size of them. The “the adder tree is configured to” limitation is an evaluation mental process that can be performed by choosing what the adder tree is configured to do. The “generate the mantissa sum as a 31-bit data element” limitation is a mathematical calculation that can be performed by generating the mantissa sum by hand using pen and paper and choosing the size of them. Under step 2A Prong 2, The “receive” limitation, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “receive” in the context of the claim encompasses mere data gathering. The remaining additional elements amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). None of the additional elements regarding the generic computer components (i.e. the multiplier circuit, the summing circuit, the shifting circuit, the adder tree, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under Step 2B, the claim recites “wherein the multiplier and summing circuits are configured to receive each data element of the plurality of input data elements and the plurality of weight data elements having a FP16 format”, and, per MPEP 2106.05(d) (Il), the courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity:
i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); and
iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93.
With regards to claim 4, it is directed to mental processes and/or mathematical concepts. The “wherein the multiplier circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier circuit is configured to do. The “perform the multiplication and reformatting operations by reformatting the signed mantissas of the some or all of the pluralities of input and weight data elements to two's complement” limitation is a mathematical calculation that can be performed by reformatting the signed mantissas by hand using pen and paper. The “multiplying the some or all of the reformatted mantissas” limitation is a mathematical calculation that can be performed by multiplying the some or all of the reformatted mantissas by hand using pen and paper. Under step 2A Prong 2, none of the additional elements regarding the generic computer components (i.e. multiplier circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under step 2B, the claims do not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claim 5, it is directed to mental processes and/or mathematical concepts. The “wherein the multiplier circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier circuit is configured to do. The “perform the multiplication and reformatting operations by generating a plurality of sign bits by performing an exclusive OR operation on sign bits of the signed mantissas” limitation is a mathematical calculation that can be performed by performing an exclusive OR operation on sign bits by hand using pen and paper. The “generating a corresponding plurality of mantissa products by multiplying mantissa bits” limitation is a mathematical calculation that can be performed by multiplying the mantissa bits by hand using pen and paper. The “reformatting the pluralities of sign bits and mantissa products to two's complement” limitation is a mathematical calculation that can be performed by reformatting the pluralities of sign bits and mantissa products to two's complement by hand using pen and paper. Under step 2A Prong 2, none of the additional elements regarding the generic computer components (i.e. multiplier circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under step 2B, the claims do not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claim 6, it is directed to mental processes and/or mathematical concepts. The “wherein the multiplier and summing circuits are configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier and summing circuits are configured to do. The “receive each of the plurality of input data elements and the plurality of weight data elements having a total of four data elements” limitation is an evaluation mental process and mathematical relationship that can be performed by choosing the number of inputs. The “the multiplier circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier circuit is configured to do. The “perform a total of sixteen or fewer multiplication operations” limitation is an evaluation mental process and mathematical calculation that can be performed by choosing the number of operations and then performing them by hand using pen and paper. The “the summing circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the summing circuit is configured to do. The “perform a total of sixteen summing operations on the plurality of input data elements and the plurality of weight data elements limitation is an evaluation mental process and mathematical calculation that can be performed by choosing the number of operations and then performing them by hand using pen and paper. Under step 2A Prong 2, The “receive” limitation, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “receive” in the context of the claim encompasses mere data gathering. The remaining additional elements amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). None of the additional elements regarding the generic computer components (i.e. the multiplier circuit, the summing circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under Step 2B, the claim recites “wherein the multiplier and summing circuits are configured to receive each of the plurality of input data elements and the plurality of weight data elements having a total of four data elements”, and, per MPEP 2106.05(d) (Il), the courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity:
i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); and
iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93.
With regards to claim 7, it is directed to mental processes and/or mathematical concepts. The “wherein the shifting circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the shifting circuit is configured to do. The “for each product of the plurality of products: right-shift the product by the amount” limitation is a mathematical calculation that can be performed by right shifting the product by hand using pen and paper. The “add a number of leading sign bits to the shifted product, the number being equal to the amount,” limitation is a mathematical calculation that can be performed by adding the number of leading sign bits by hand using pen and paper. The “add one or more trailing zero bits corresponding to the amount being less than a difference threshold” limitation is a mathematical calculation that can be performed by adding the trailing zero bits by hand using pen and paper. Under step 2A Prong 2, none of the additional elements regarding the generic computer components (i.e. the shifting circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under step 2B, the claims do not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claims 8 and 17, they are directed to mental processes and/or mathematical concepts. The “wherein the shifting circuit comprises” limitation is an evaluation mental process that can be performed by choosing what the shifting circuit comprises. The “a first stage configured to” limitation is an evaluation mental process that can be performed by choosing what the first stage is configured to do. “generate a plurality of intermediate data elements from the plurality of products based on the two least significant bits of the corresponding differences” limitation is a mathematical calculation that can be performed by generating the plurality of intermediate data elements by hand using pen and paper. The “a second stage configured to” limitation is an evaluation mental process that can be performed by choosing what the second stage is configured to do. “generate the plurality of shifted products from the plurality of intermediate data elements based on the other bits of the corresponding differences” limitation is a mathematical calculation that can be performed by generating the plurality shifted products by hand using pen and paper. Under step 2A Prong 2, none of the additional elements regarding the generic computer components (i.e. the shifting circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under step 2B, the claims do not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claim 9, it is directed to mental processes and/or mathematical concepts. The “further comprising” limitation is an evaluation mental process that can be performed by choosing what the circuit comprises. The “a difference circuit configured to” limitation is an evaluation mental process that can be performed by choosing what the difference circuit is configured to do. The “determine the maximum sum of the plurality of sums” limitation is a mathematical calculation that can be performed by determining the maximum sum by hand using pen and paper. The “calculate each difference by subtracting the corresponding sum of the plurality of sums from the maximum sum” limitation is a mathematical calculation that can be performed by calculating each difference by hand using pen and paper. The “output each difference to the shifting circuit” limitation is an evaluation mental process that can be performed by choosing where the differences are output. Under step 2A Prong 2, none of the additional elements regarding the generic computer components (i.e. the shifting circuit, the difference circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under step 2B, the claims do not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claim 10, it is directed to mental processes and/or mathematical concepts. The “wherein the shifting circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the shifting circuit is configured to do. The “for each difference: based on the difference being less than a difference threshold, generate the corresponding shifted product of the plurality of shifted products from the corresponding product of the plurality of products” limitation is a mathematical calculation that can be performed by generating the corresponding shifted product by hand using pen and paper. The “based on the difference being greater than or equal to the difference threshold, generate the corresponding shifted product” limitation is a mathematical calculation that can be performed by generating the corresponding shifted product by hand using pen and paper. Under step 2A Prong 2, none of the additional elements regarding the generic computer components (i.e. the shifting circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under step 2B, the claims do not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claims 11 and 16, they are directed to mental processes and/or mathematical concepts. The “wherein the multiplier circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the multiplier circuit is configured to do. The “receive each difference from the difference circuit” limitation is an evaluation mental process that can be performed by choosing where the data comes from. The “and for each difference, perform the multiplication and reformatting operations on the corresponding input and weight data elements only if the difference is less than a difference threshold” limitation is an evaluation mental process and mathematical calculation that can be performed by choosing when to perform the operation and then performing the operations by hand using pen and paper. Under step 2A Prong 2, none of the additional elements regarding the generic computer components (i.e. the difference circuit, the multiplier circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under step 2B, the claims do not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claims 12 and 18, they are directed to mental processes and/or mathematical concepts. The “wherein the circuit is configured to” limitation is an evaluation mental process that can be performed by choosing what the circuit is configured to do. The “convert the mantissa sum to a sign bit plus a plurality of mantissa bits” limitation is a mathematical calculation that can be performed by converting the mantissa sum by hand using pen and paper. Under step 2A Prong 2, none of the additional elements regarding the generic computer components (i.e. the circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under step 2B, the claims do not recite any additional elements that integrate the abstract idea into a practical application, nor do they amount to significantly more than the judicial exception.
With regards to claim 14, it is directed to mental processes and/or mathematical concepts. The “wherein the circuit comprises” limitation is an evaluation mental process that can be performed by choosing what the circuit comprises. The “the receiving the signed mantissa and exponent of each data element comprises” limitation is an evaluation mental process that can be performed by choosing what the receiving comprises. The “receiving each data element from the plurality of input data elements and the plurality of weight data elements stored in the memory array” limitation is an evaluation mental process that can be performed by choosing where the data comes from. Under step 2A Prong 2, The “receiving” limitations, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “receiving” in the context of the claim encompasses mere data gathering. The remaining additional elements amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). None of the additional elements regarding the generic computer components (i.e. the circuit, the memory array, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under Step 2B, the claim recites “receiving the signed mantissa and exponent of each data element comprises”, “receiving the signed mantissa and exponent of each data element comprises receiving each data element from the plurality of input data elements and the plurality of weight data elements stored in the memory array”, and, per MPEP 2106.05(d) (Il), the courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity:
i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); and
iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93.
With regards to claim 15, it is directed to mental processes and/or mathematical concepts. The “wherein the receiving the signed mantissa and exponent of each data element comprises” limitation is an evaluation mental process that can be performed by choosing what the receiving comprises. The “receiving each data element of the plurality of input data elements and the plurality of weight data elements having either a BF16 format or a FP16 format” is an evaluation mental process and mathematical relationship that can be performed by choosing the format of the inputs. Under step 2A Prong 2, The “receiving” limitations, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “receiving” in the context of the claim encompasses mere data gathering. The remaining additional elements amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). The claim does not recite any additional elements that integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under Step 2B, the claim recites “wherein the receiving the signed mantissa and exponent of each data element comprises”, “receiving each data element of the plurality of input data elements and the plurality of weight data elements having either a BF16 format or a FP16 format”, and, per MPEP 2106.05(d) (Il), the courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity:
i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); and
iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93.
With regards to claim 19, at step 1, the claim is directed to a machine, which is a statutory category of invention.
At Step 2A Prong 1, the examiner notes that the claim is directed to mental processes and/or mathematical concepts. The claim language has been reproduced below:
A compute-in-memory (CIM) circuit comprising: (mental process, evaluation)
a memory array configured to (mental process, evaluation) store a plurality of input data elements and a plurality of weight data elements; (mental process, evaluation; mathematical relationship)
a multiply-accumulate (MAC) unit configured to (mental process, evaluation) generate a sequence of partial sums based on the plurality of input data elements and the plurality of weight data elements; (mathematical calculation)
an adder configured to (mental process, evaluation) generate a sequence of accumulated sums by adding each partial sum of the sequence of partial sums to a stored accumulated sum; and (mathematical calculation)
a buffer configured to (mental process, evaluation)
store each accumulated sum of the sequence of accumulated sums as the stored accumulated sum, (mental process, evaluation)
output each stored accumulated sum to the adder, and (mental process, evaluation)
output a final stored accumulated sum from the CIM circuit. (mental process, evaluation)
Each of the non-bolded limitations are mental processes and/or mathematical calculations. The “A compute-in-memory (CIM) circuit comprising” limitation is an evaluation mental process that can be performed by choosing what the CIM comprises. The “a memory array configured to” limitation is an evaluation mental process that can be performed by choosing what the memory array is configured to do. The “store a plurality of input data elements and a plurality of weight data elements” limitation is an evaluation mental process and mathematical relationship that can performed by choosing what the array stores. The “a multiply-accumulate (MAC) unit configured to” limitation is an evaluation mental process that can be performed by choosing what the MAC is configured to do. The “generate a sequence of partial sums” limitation is a mathematical calculation that can be performed by generating the sequence of partial sums by hand using pen and paper. The “an adder configured to” limitation is an evaluation mental process that can be performed by choosing what the adder is configured to do. The “generate a sequence of accumulated sums” limitation is a mathematical calculation that can be performed by generating the sequence of accumulated sums by hand using pen and paper. The “a buffer configured to” limitation is an evaluation mental process that can be performed by choosing what the buffer is configured to do. The “store each accumulated sum of the sequence” limitation is an evaluation mental process that can be performed by choosing what the buffer stores. The “output each stored accumulated sum” limitation is an evaluation mental process that can be performed by choosing what the buffer outputs and where it outputs it. The “output a final stored accumulated sum” limitation is an evaluation mental process that can be performed by choosing what the buffer outputs and where it outputs it.
At step 2A Prong 2, the additional elements are bolded above. The “store” limitations, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “store” in the context of the claim encompasses mere data gathering. The “output” limitations, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “output” in the context of the claim encompasses mere data gathering. The remaining additional elements amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f).
At Step 2B, the claim recites “store a plurality of input data elements and a plurality of weight data elements”, “store each accumulated sum of the sequence of accumulated sums as the stored accumulated sum”, “output each stored accumulated sum to the adder”, “output a final stored accumulated sum from the CIM circuit”, and, per MPEP 2106.05(d) (Il), the courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity:
i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); and
iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93.
With regards to claim 20, it is directed to mental processes and/or mathematical concepts. The “wherein the buffer is configured to” limitation is an evaluation mental process that can be performed by choosing what the buffer is configured to do. The “output the final stored accumulated sum to a memory array of another CIM circuit” limitation is an evaluation mental process that can be performed by choosing what the buffer outputs and where it outputs it. Under step 2A Prong 2, The “output” limitation, as claimed under BRI, are additional elements that are insignificant extra-solution activity. The “output” in the context of the claim encompasses mere data gathering. The remaining additional elements amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). None of the additional elements regarding the generic computer components (i.e. the buffer, the memory array, the CIM circuit, etc.) are more than high level generic computer components that amount to no more than components comprising mere instructions to apply the exception and do not integrate the judicial exception into a practical application. See MPEP 2106.05(f). Under Step 2B, the claim recites “output the final stored accumulated sum to a memory array of another CIM circuit”, and, per MPEP 2106.05(d) (Il), the courts have recognized the following computer functions as well-understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity:
i. Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); TLI Communications LLC v. AV Auto. LLC, 823 F.3d 607, 610, 118 USPQ2d 1744, 1745 (Fed. Cir. 2016) (using a telephone for image transmission); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); and
iv. Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 4, 9-10, 12-15, and 18-19 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Song et al. (US 20220229633 A1) hereinafter Song.
With regards to claim 1, Song teaches A circuit comprising: a multiplier circuit configured to: receive a signed mantissa of each data element of a plurality of input data elements and a plurality of weight data elements, (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC)
and generate a plurality of two's complement products by performing multiplication and reformatting operations on some or all of the signed mantissas of the plurality of input data elements and some or all of the signed mantissas of the plurality of weight data elements; (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC; Song [0523]: The number of two's complement circuits 6231(1)-6231(8) and the number of multiplexers 6232(1)-6232(8) constituting the negative number processing circuit 6230 may be equal to or greater than the number of multipliers MUL0-MUL7 constituting the multiplication circuit)
a summing circuit configured to: receive an exponent of each data element of the plurality of input data elements and the plurality of weight data elements, and generate a plurality of sums by adding each exponent of the plurality of input data elements to each exponent of the plurality of weight data elements; (Song [0275]: The exponent processing circuit 1120 may include a first exponent adder 1121 and a second exponent adder 1122. The first exponent adder 1121 may receive exponent bits E1[7:0] of the first weight data W0_FLT and exponent bits E2[7:0] of the first vector data V0_FLT. The first exponent adder 1121 may add the exponent bits E1[7:0] of the first weight data W0_FLT and the exponent bits E2[7:0] of the first vector data V0_FLT, and output addition result data; Song Fig. 31: shows multiple multipliers that each include an exponent summing circuit)
a shifting circuit configured to shift each product of the plurality of products by an amount equal to a difference between a corresponding sum of the plurality of sums and a maximum sum; (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data)
and an adder tree configured to generate a mantissa sum from the plurality of shifted products (Song [0005]: The adder tree may be configured to add the pre-processed mantissa data to output mantissa data of multiplication addition data).
With regards to claim 4, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the multiplier circuit is configured to perform the multiplication and reformatting operations by reformatting the signed mantissas of the some or all of the pluralities of input and weight data elements to two's complement, (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC; Song [0523]: The number of two's complement circuits 6231(1)-6231(8) and the number of multiplexers 6232(1)-6232(8) constituting the negative number processing circuit 6230 may be equal to or greater than the number of multipliers MUL0-MUL7 constituting the multiplication circuit)
and multiplying the some or all of the reformatted mantissas of the plurality of input data elements with the some or all of the reformatted mantissas of the plurality of weight data elements (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC; Song [0523]: The number of two's complement circuits 6231(1)-6231(8) and the number of multiplexers 6232(1)-6232(8) constituting the negative number processing circuit 6230 may be equal to or greater than the number of multipliers MUL0-MUL7 constituting the multiplication circuit).
With regards to claim 9, Song teaches all of the limitations of claim 1 above. Song further teaches further comprising a difference circuit configured to: determine the maximum sum of the plurality of sums, (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data)
calculate each difference by subtracting the corresponding sum of the plurality of sums from the maximum sum, (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data)
and output each difference to the shifting circuit (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data).
With regards to claim 10, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the shifting circuit is configured to, for each difference: based on the difference being less than a difference threshold, generate the corresponding shifted product of the plurality of shifted products from the corresponding product of the plurality of products, (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data; Song [0400]: In an embodiment, when overflow and underflow do not occur, the overflow/underflow checker 4710 may output an overflow/underflow signal OUF[1:0] of ‘00’)
or based on the difference being greater than or equal to the difference threshold, generate the corresponding shifted product of the plurality of shifted products as a zero-value data element (Song [0425]: in response to an overflow/underfloor signal OUF[1:0] of ‘01’. The first 3:1 multiplexer 4733-1 may output the first mantissa minimum value MINM1 inputted through the third input terminal IN3 as the first data type FP16 10-bit mantissa bits FP16_MAN[22:1 3] in response to an overflow/underflow signal OUF[1:0] of ‘10’).
With regards to claim 12, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the circuit is configured to convert the mantissa sum to a sign bit plus a plurality of mantissa bits (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0274]: the first multiplier MUL0 may include a sign processing circuit 1110… The sign processing circuit 1110 may include an exclusive OR (hereinafter, referred to as “XOR”) gate 1111. The XOR gate 1111 may receive a sign bit S1[0] of the first weight data W0_FLT and a sign bit S2[0] of the first vector data V0_FLT; Song Fig. 33: shows the mantissa being appended to the sign).
Claim 13 is directed to a method that implements the same or similar features as the system of claim 1 and is therefore rejected for at least the same reasons therein.
With regards to claim 14, Song teaches all of the limitations of claim 13 above. Song further teaches wherein the circuit comprises a memory array, (Song [0145]: FIG. 2 is a block diagram illustrating a PIM system… The first memory bank (BANK0) 111 and the second memory bank (BANK1) 112 may represent a memory region for storing data… Each of the first and second memory banks 111 and 112 may include at least one cell array which includes memory unit cells located at cross points of a plurality of rows and a plurality of columns)
and the receiving the signed mantissa and exponent of each data element comprises receiving each data element from the plurality of input data elements and the plurality of weight data elements stored in the memory array (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC; Song [0145]: FIG. 2 is a block diagram illustrating a PIM system… The first memory bank (BANK0) 111 and the second memory bank (BANK1) 112 may represent a memory region for storing data… Each of the first and second memory banks 111 and 112 may include at least one cell array which includes memory unit cells located at cross points of a plurality of rows and a plurality of columns).
With regards to claim 15, Song teaches all of the limitations of claim 13 above. Song further teaches wherein the receiving the signed mantissa and exponent of each data element comprises receiving each data element of the plurality of input data elements and the plurality of weight data elements having either a BF16 format or a FP16 format (Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type... this is only an example, and the types of the first weight data W0_FLT and the first vector data V0_FLT may be types other than the 16-bit brain floating-point (BF16) type, such as 16-bit floating-point (FP16) type).
With regards to claim 18, Song teaches all of the limitations of claim 13 above. Song further teaches further comprising converting the mantissa sum to a sign bit plus a plurality of mantissa bits (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0274]: the first multiplier MUL0 may include a sign processing circuit 1110… The sign processing circuit 1110 may include an exclusive OR (hereinafter, referred to as “XOR”) gate 1111. The XOR gate 1111 may receive a sign bit S1[0] of the first weight data W0_FLT and a sign bit S2[0] of the first vector data V0_FLT; Song Fig. 33: shows the mantissa being appended to the sign).
With regards to claim 19, Song teaches A compute-in-memory (CIM) circuit comprising: a memory array configured to store a plurality of input data elements and a plurality of weight data elements; (Song [0145]: FIG. 2 is a block diagram illustrating a PIM system… The first memory bank (BANK0) 111 and the second memory bank (BANK1) 112 may represent a memory region for storing data… Each of the first and second memory banks 111 and 112 may include at least one cell array which includes memory unit cells located at cross points of a plurality of rows and a plurality of columns; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC)
a multiply-accumulate (MAC) unit configured to generate a sequence of partial sums based on the plurality of input data elements and the plurality of weight data elements; (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC)
an adder configured to generate a sequence of accumulated sums by adding each partial sum of the sequence of partial sums to a stored accumulated sum; (Song [0005]: The adder tree may be configured to add the pre-processed mantissa data to output mantissa data of multiplication addition data)
and a buffer configured to store each accumulated sum of the sequence of accumulated sums as the stored accumulated sum, (Song [0532]: FIG. 89 is a circuit diagram illustrating an example of a configuration of the accumulator… and the latch circuit 6450 of the accumulator; Song Fig. 89: shows that the latch circuit stores the sum)
output each stored accumulated sum to the adder, (Song [0532]: FIG. 89 is a circuit diagram illustrating an example of a configuration of the accumulator… and the latch circuit 6450 of the accumulator; Song Fig. 89: shows that the latch circuit stores the sum and then outputs it back to the adder)
and output a final stored accumulated sum from the CIM circuit (Song [0137]: After MAC operations, the MAC operator may output MAC result data. The MAC result data may be stored in the data storage region 11 or output from the PIM device; Song [0532]: FIG. 89 is a circuit diagram illustrating an example of a configuration of the accumulator… and the latch circuit 6450 of the accumulator; Song Fig. 89: shows that the latch circuit stores the sum and outputs the sum).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Song in view of Burgess et al. (US 20220129245 A1) hereinafter Burgess further in view of Ulrich et al. (US 20210255830 A1) hereinafter Ulrich.
With regards to claim 2, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the multiplier and summing circuits are configured to receive each data element of the plurality of input data elements and the plurality of weight data elements having a BF16 format, (Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type)
the multiplier circuit is configured to generate the plurality of products as [17-bit data elements,] (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit)
the summing circuit is configured to generate the plurality of sums as [nine-bit data elements,] (Song [0275]: The exponent processing circuit 1120 may include a first exponent adder 1121 and a second exponent adder 1122. The first exponent adder 1121 may receive exponent bits E1[7:0] of the first weight data W0_FLT and exponent bits E2[7:0] of the first vector data V0_FLT. The first exponent adder 1121 may add the exponent bits E1[7:0] of the first weight data W0_FLT and the exponent bits E2[7:0] of the first vector data V0_FLT, and output addition result data)
the shifting circuit is configured to generate the plurality of shifted products as 21-bit data elements, (Song [0642]: the shifted mantissa data M_SFT_LATCH[7:0] of the latch data transmitted from the second mantissa shifting circuit 6422D to generate and output accumulative mantissa data M_ACC[20:0]. In an example, one carry bit may be added during the accumulative addition operation in the second accumulative adder 6423D, and accordingly, the accumulative mantissa data M_ACC[20:0] may have a size of 21 bits)
and the adder tree is configured to generate the mantissa sum as a 25-bit data element (Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song [0005]: The adder tree may be configured to add the pre-processed mantissa data to output mantissa data of multiplication addition data; Song Fig. 32: shows the output being 25 bits long).
Song fails to teach 17-bit data elements.
However, Burgess teaches 17-bit data elements, (Burgess [0051]: The 16-bit bfloat16 product significand and the +sign bit of the SIZD field are converted into to a 17-bit 2's-complement number).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the 17-bit data elements as taught by Burgess. One of ordinary skill in the art would be motivated to make this combination because the various embodiments of the data item enable fast and simple hardware for computing dot products, often used in machine learning applications as taught by Burgess (Burgess [0025]).
Song in view of Burgess fails to teach nine-bit data elements.
However, Ulrich teaches nine-bit data elements, (Ulrich [0048]: a 9-bit adder is used to add exponents of fp16 and bfloat16 format numbers because the largest number of exponent bits for the two formats is 8 exponent bits for bfloat16 and the sum of two 8-bit numbers can be a 9-bit number).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song I view of Burgess with the nine-bit data elements as taught by Ulrich. One of ordinary skill in the art would be motivated to make this combination because practical and technological benefits of the disclosed device include improved efficiency and performance of multiplication operations, e.g., through more efficient use of integrated circuit chip area and reduced power consumption as taught by Ulrich (Ulrich [0011]).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Song in view of Ulrich further in view of Abdallah et al. (US 20040268094 A1) hereinafter Abdallah.
With regards to claim 3, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the multiplier and summing circuits are configured to receive each data element of the plurality of input data elements and the plurality of weight data elements having a FP16 format, (Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31… and the types of the first weight data W0_FLT and the first vector data V0_FLT may be types other than the 16-bit brain floating-point (BF16) type, such as 16-bit floating-point (FP16) type)
the multiplier circuit is configured to generate the plurality of products as [23-bit data elements,] (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit)
the summing circuit is configured to generate the plurality of sums as [six-bit data elements,] (Song [0275]: The exponent processing circuit 1120 may include a first exponent adder 1121 and a second exponent adder 1122. The first exponent adder 1121 may receive exponent bits E1[7:0] of the first weight data W0_FLT and exponent bits E2[7:0] of the first vector data V0_FLT. The first exponent adder 1121 may add the exponent bits E1[7:0] of the first weight data W0_FLT and the exponent bits E2[7:0] of the first vector data V0_FLT, and output addition result data)
the shifting circuit is configured to generate the plurality of shifted products as [27-bit data elements,] (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data)
and the adder tree is configured to generate the mantissa sum as a [31-bit data element] (Song [0005]: The adder tree may be configured to add the pre-processed mantissa data to output mantissa data of multiplication addition data).
Song fails to teach 23-bit data elements, six-bit data elements, and a 31-bit data element.
However, Ulrich teaches 23-bit data elements, (Ulrich [0043]: the output of the multiplication is placed into another format… An example of the third format representation is the single-precision floating-point format (also referred to herein as “fp32”), which includes a sign bit, 8 exponent bits, and 23 mantissa bits)
six-bit data elements, (Ulrich [0014]: an fp16 dot product unit may require a 6-bit adder (adding two 5-bit exponents can result in a 6-bit result))
a 31-bit data element (Ulrich [0043]: the output of the multiplication is placed into another format… An example of the third format representation is the single-precision floating-point format (also referred to herein as “fp32”), which includes a sign bit, 8 exponent bits, and 23 mantissa bits; (this shows that the output is 31 bits, which includes the exponent and mantissa)).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the 23-bit data elements, the six-bit data elements, and the 31-bit data element as taught by Ulrich. One of ordinary skill in the art would be motivated to make this combination because practical and technological benefits of the disclosed device include improved efficiency and performance of multiplication operations, e.g., through more efficient use of integrated circuit chip area and reduced power consumption as taught by Ulrich (Ulrich [0011]).
Song in view of Ulrich fails to teach 27-bit data elements.
However, Abdallah teaches 27-bit data elements, (Abdallah [0205]: The fraction field contains a binary fraction of 27 bits).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song in view of Ulrich with the 27-bit data elements as taught by Abdallah. One of ordinary skill in the art would be motivated to make this combination because a decreased number of instructions in the processing of graphics data, no requirement for duplicated floating point execution resources, and higher application processing efficiency as taught by Abdallah (Abdallah [0053]).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Song in view of Burgess.
With regards to claim 5, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the multiplier circuit is configured to perform the multiplication and reformatting operations by generating a plurality of sign bits by performing an exclusive OR operation on sign bits of the signed mantissas of the some or all of the pluralities of input and weight data elements, (Song [0274]: the first multiplier MUL0 may include a sign processing circuit 1110… The sign processing circuit 1110 may include an exclusive OR (hereinafter, referred to as “XOR”) gate 1111. The XOR gate 1111 may receive a sign bit S1[0] of the first weight data W0_FLT and a sign bit S2[0] of the first vector data V0_FLT)
generating a corresponding plurality of mantissa products by multiplying mantissa bits of the signed mantissas of the some or all of the plurality of input data elements with mantissa bits of the signed mantissas of the some or all of the plurality of weight data elements, (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC)
and reformatting the pluralities of sign bits and mantissa [products] to two's complement (Song [0523]: The number of two's complement circuits 6231(1)-6231(8) and the number of multiplexers 6232(1)-6232(8) constituting the negative number processing circuit 6230 may be equal to or greater than the number of multipliers MUL0-MUL7 constituting the multiplication circuit).
Song fails to teach reformatting the products.
However, Burgess teaches reformatting the products (Burgess [0051]: The 16-bit bfloat16 product significand and the +sign bit of the SIZD field are converted into to a 17-bit 2's-complement number).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with reformatting the products as taught by Burgess. One of ordinary skill in the art would be motivated to make this combination because the various embodiments of the data item enable fast and simple hardware for computing dot products, often used in machine learning applications as taught by Burgess (Burgess [0025]).
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Song in view of Czajkowski et al. (US 20150067010 A1) hereinafter Czajkowski further in view of DiBrino et al. (US 20230053261 A1) hereinafter DiBrino.
With regards to claim 6, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the multiplier and summing circuits are configured to receive each of the plurality of input data elements and the plurality of weight data elements having a total of [four data elements,] (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC)
the multiplier circuit is configured to perform a total of sixteen or fewer multiplication operations on the plurality of input data elements and the plurality of weight data elements, (Song [0255]: Specifically, the multiplying circuit 1100 may include a plurality of multipliers, for example, first to eighth multipliers MUL0-MUL7 arranged in parallel with each other. Here, the parallel arrangement may mean an arrangement structure in which data input/output and arithmetic operations are independently performed, and this may be applied in the same manner hereinafter. Each of the multipliers MUL0-MUL7 may receive weight data W0_FLT-W7_FLT and vector data V0_FLT-V7_FLT)
and the summing circuit is configured to perform [a total of sixteen summing operations] on the plurality of input data elements and the plurality of weight data elements (Song [0275]: The exponent processing circuit 1120 may include a first exponent adder 1121 and a second exponent adder 1122. The first exponent adder 1121 may receive exponent bits E1[7:0] of the first weight data W0_FLT and exponent bits E2[7:0] of the first vector data V0_FLT. The first exponent adder 1121 may add the exponent bits E1[7:0] of the first weight data W0_FLT and the exponent bits E2[7:0] of the first vector data V0_FLT, and output addition result data; Song Fig. 31: shows multiple multipliers that each include an exponent summing circuit).
Song fails to teach having a total of four data elements.
However, Czajkowski teaches having a total of four data elements, (Czajkowski [0032]: An illustrative diagram of the addition of these four floating-point numbers by an adder tree such as adder tree 400 is shown in FIG. 3).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the four inputs as taught by Czajkowski. One of ordinary skill in the art would be motivated to make this combination because it would allow the system to process multiple inputs in parallel, increasing efficiency.
Song in view of Czajkowski fails to teach a total of sixteen summing operations.
However, DiBrino teaches a total of sixteen summing operations (DiBrino [0080]: on the exponent side in stage 1, the exponents of the multiplier and multiplicand for the dot-product (DP) for each component go into a corresponding one of the (16 in this example) exponent adders).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song in view of Czajkowski with the sixteen summing operations as taught by DiBrino. One of ordinary skill in the art would be motivated to make this combination because the low-latency embodied can be used to achieve higher clock speeds as taught by DiBrino (Dibrino [0116]).
Claims 7, 11, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Song in view of Finch et al. (US 20220405051 A1) hereinafter Finch.
With regards to claim 7, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the shifting circuit is configured to, for each product of the plurality of products: right-shift the product by the amount, (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data)
[add a number of leading sign bits to the shifted product, the number] being equal to the amount, (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data).
Song fails to teach add a number of leading sign bits to the shifted product, the number [being equal to the amount,] and add one or more trailing zero bits corresponding to the amount being less than a difference threshold.
However, Finch teaches add a number of leading sign bits to the shifted product, the number [being equal to the amount,] (Finch [0025]: padding the normalized mantissa multiplication with leading 0s and trailing 0s; Finch [0026]: replacing the padded normalized mantissa multiplication with a twos complement of the padded normalized mantissa multiplication if the sign bit is 1)
and add one or more trailing zero bits corresponding to the amount being less than a difference threshold (Finch [0025]: padding the normalized mantissa multiplication with leading 0s and trailing 0s; Finch [0061]: If the threshold is not exceeded, the sum is normalized with MAX_EXP to form the floating point result, as previously described).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the leading sign bits, trailing zeroes, and threshold as taught by Finch. One of ordinary skill in the art would be motivated to make this combination because it is desired to provide a scalable high speed, low power multiply-accumulate (MAC) apparatus and method operative to form dot products from the addition of large numbers of floating point multiplicands as taught by Finch (Finch [0004]).
With regards to claim 11, Song teaches all of the limitations of claim 9 above. Song further teaches wherein the multiplier circuit is configured to: receive each difference from the difference circuit, (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data)
and for each difference, perform the multiplication and reformatting operations on the corresponding input and weight data elements [only if the difference is less than a difference threshold] (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit).
Song fails to teach only if the difference is less than a difference threshold.
However, Finch teaches only if the difference is less than a difference threshold (Finch [0061]: If the threshold is not exceeded, the sum is normalized with MAX_EXP to form the floating point result, as previously described).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the threshold as taught by Finch. One of ordinary skill in the art would be motivated to make this combination because it is desired to provide a scalable high speed, low power multiply-accumulate (MAC) apparatus and method operative to form dot products from the addition of large numbers of floating point multiplicands as taught by Finch (Finch [0004]).
With regards to claim 16, Song teaches all of the limitations of claim 13 above. Song further teaches wherein the generating the plurality of two's complement products comprises: for each difference between the corresponding sum of the plurality of sums and the maximum sum, performing the multiplication and reformatting operations on the corresponding input and weight data elements [only if the difference is less than a difference threshold] (Song [0006]: A multiplication-accumulation (MAC) according to an embodiment of the present disclosure may include a multiplication circuit; Song [0270]: FIG. 32 illustrates an embodiment of data formats of input data and output data of the first multiplier in the MAC operator of FIG. 31... In the present embodiment, it is premised that the input data, that is, the first weight data W0_FLT and the first vector data V0_FLT are in 16-bit brain floating-point (BF16 ) type; Song Fig. 5: shows that a plurality of weights and inputs are used in the MAC; Song [0523]: The number of two's complement circuits 6231(1)-6231(8) and the number of multiplexers 6232(1)-6232(8) constituting the negative number processing circuit 6230 may be equal to or greater than the number of multipliers MUL0-MUL7 constituting the multiplication circuit).
Song fails to teach only if the difference is less than a difference threshold.
However, Finch teaches only if the difference is less than a difference threshold (Finch [0061]: If the threshold is not exceeded, the sum is normalized with MAX_EXP to form the floating point result, as previously described).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the threshold as taught by Finch. One of ordinary skill in the art would be motivated to make this combination because it is desired to provide a scalable high speed, low power multiply-accumulate (MAC) apparatus and method operative to form dot products from the addition of large numbers of floating point multiplicands as taught by Finch (Finch [0004]).
Claims 8 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Song in view of Hilker et al. (US 20130060828 A1) hereinafter Hilker.
With regards to claim 8, Song teaches all of the limitations of claim 1 above. Song further teaches wherein the shifting circuit comprises: [a first stage configured to generate a plurality of intermediate data elements from the plurality of products based on the two least significant bits of the corresponding differences,] (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data).
Song fails to teach a first stage configured to generate a plurality of intermediate data elements from the plurality of products based on the two least significant bits of the corresponding differences, and a second stage configured to generate the plurality of shifted products from the plurality of intermediate data elements based on the other bits of the corresponding differences.
However, Hilker teaches a first stage configured to generate a plurality of intermediate data elements from the plurality of products based on the two least significant bits of the corresponding differences, (Hilker [0033]: the least significant bits (LSBs) of the exponent differences for all calculations are naturally available first and thus may be used first to optimize the aligner for lowest latency (i.e., the finest 1.times. shifting; the first two LSBs [0:1])
and a second stage configured to generate the plurality of shifted products from the plurality of intermediate data elements based on the other bits of the corresponding differences (Hilker [0033]: the least significant bits (LSBs) of the exponent differences for all calculations are naturally available first and thus may be used first to optimize the aligner for lowest latency (i.e., the finest 1.times. shifting; the first two LSBs [0:1], decoded by 00=shift 0, 10=shift 1, 01=shift 2 and 11=shift 3) may occur first, followed by the 4-bit shifting (the next two significant bits [2:3], decoded by 00=shift 0, 10=shift 4, 01=shift 8 and 11=shift 12), the 16-bit shifting (the next two significant bits [4:5], decoded by 00=shift 0, 10=shift 16, 01=shift 32 and 11=shift 48), and the 64-bit shifting (the next two significant bits [6:7], decoded by 00=shift 0, 10=shift 64, 01=shift 128 and 11=shift 192) as the more significant bits of the exponent differences become available).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the shifting stages as taught by Hilker. One of ordinary skill in the art would be motivated to make this combination because this would optimize the aligner for lowest latency as taught by Hilker (Hilker [0033]).
With regards to claim 17, Song teaches all of the limitations of claim 13 above. Song further teaches wherein the shifting each product of the plurality of products comprises: [generating an intermediate data element from the product based on the two least significant bits of the corresponding difference;] (Song [0005]: The pre-processing circuit may be configured to perform a shifting operation of shifting mantissa data of the multiplication data by a difference between first maximum exponent data having a greatest value among exponent data of the multiplication data and the exponent data of the multiplication data to output pre-processed mantissa data).
Song fails to teach generating an intermediate data element from the product based on the two least significant bits of the corresponding difference; and generating the corresponding shifted product from the intermediate data element based on the other bits of the corresponding difference.
However, Hilker teaches generating an intermediate data element from the product based on the two least significant bits of the corresponding difference; (Hilker [0033]: the least significant bits (LSBs) of the exponent differences for all calculations are naturally available first and thus may be used first to optimize the aligner for lowest latency (i.e., the finest 1.times. shifting; the first two LSBs [0:1])
and generating the corresponding shifted product from the intermediate data element based on the other bits of the corresponding difference (Hilker [0033]: the least significant bits (LSBs) of the exponent differences for all calculations are naturally available first and thus may be used first to optimize the aligner for lowest latency (i.e., the finest 1.times. shifting; the first two LSBs [0:1], decoded by 00=shift 0, 10=shift 1, 01=shift 2 and 11=shift 3) may occur first, followed by the 4-bit shifting (the next two significant bits [2:3], decoded by 00=shift 0, 10=shift 4, 01=shift 8 and 11=shift 12), the 16-bit shifting (the next two significant bits [4:5], decoded by 00=shift 0, 10=shift 16, 01=shift 32 and 11=shift 48), and the 64-bit shifting (the next two significant bits [6:7], decoded by 00=shift 0, 10=shift 64, 01=shift 128 and 11=shift 192) as the more significant bits of the exponent differences become available).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the shifting stages as taught by Hilker. One of ordinary skill in the art would be motivated to make this combination because this would optimize the aligner for lowest latency as taught by Hilker (Hilker [0033]).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Song in view of Zidan et al. (US 20230259282 A1) hereinafter Zidan.
With regards to claim 20, Song teaches all of the limitations of claim 19 above. Song further teaches wherein the buffer is configured to output the final stored accumulated sum [to a memory array of another CIM circuit] (Song [0137]: After MAC operations, the MAC operator may output MAC result data. The MAC result data may be stored in the data storage region 11 or output from the PIM device).
Song fails to teach [wherein the buffer is configured to output the final stored accumulated sum] to a memory array of another CIM circuit.
However, Zidan teaches [wherein the buffer is configured to output the final stored accumulated sum] to a memory array of another CIM circuit (Zidan [0111]: Referring now to FIG. 4, a memory processing unit, in accordance with aspects of the present technology, is shown. The memory processing unit 400 can include a first memory region and a plurality of processing region 410-414. The first memory can include a plurality of memory regions 402-408; Zidan Fig. 4: shows that the outputs of the processing regions are output to the memory of the next processing regions).
Therefore, it would have been obvious before the effective filing date of the claimed invention for one of ordinary skill in the art to combine the teachings of Song with the second CIM circuit as taught by Zidan. One of ordinary skill in the art would be motivated to make this combination because it would increase the efficiency of the system as it could process more data in parallel. Also, the wide plurality of first memory regions organized into a plurality of columns and rows, in accordance with aspects of the present technology, advantageously reduces the number of memory access cycles, which can smother the pipeline, improve arbitration and better latency hiding as taught by Zidan (Zidan [0176]).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jakob O Gudas whose telephone number is (571)272-0695. The examiner can normally be reached Monday-Thursday: 7:30AM-5:00PM Friday: 7:30AM-4:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571) 272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.O.G./Examiner, Art Unit 2151
/James Trujillo/Supervisory Patent Examiner, Art Unit 2151