Detailed Action
This Office Action is in response to the remarks entered on 06/06/2025. Claims 13-17 have been canceled. New Claims 21-25 have been added. Claims 1-12 and 18-25 are currently pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
Applicant’s arguments, see [Remarks, pages 7-8], filed 06/06/2025, with respect to 35 U.S.C. 101 rejection have been fully considered and are persuasive. The 35 U.S.C. 101 rejection of claims 1-12 and 18-25 have been withdrawn.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 7-12 and 18-25 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. Pub. No: US2020/0160222 A1(hereafter ZHANG) in view of Meng et al. Pub. No: US2020/0242468 A1(hereafter MENG) and further in view of Kawabe et al. Pub. Num: US2020/0192633 A1(hereafter KAWABE).
Regarding claim 1, ZHANG teaches the invention substantially as claimed, including:
A neural network processing unit, comprising: an operation circuit including a fixed-point (operation) and a floating-point (operation) ([Abstract]The computation device is configured to perform a machine learning computation, and includes an operation unit, a controller unit, and a conversion unit. [0133] As illustrated in Table 3, if the identifier of data type conversion is 00, the conversion manner of data type is converting the fixed-point data into fixed-point data. If the identifier of data type conversion is 01, the conversion manner of data type is converting the floating point data into floating point data). to perform tensor operations ([0113] and the data may be n-dimensional data, where n is an integer greater than or equal to one. For example, if n=1, the data is a one-dimensional data (that is, a vector). For another example, if n=2, the data is a two-dimensional data (that is, a matrix). If n=3 or more, the data is a multi-dimensional tensor.) of a neural network ([0056] FIG. 4 is a flow chart of a forward operation of a single-layer artificial neural network according to an example of the present disclosure) in a fixed-point number representation and a floating-point number representation, respectively, ; and ([0008] The conversion unit may be configured to convert the first input data into second input data according to the opcode and the opcode field of the data conversion instruction. [0009] The operation unit may be configured to perform operations on the second input data according to the plurality of operation instructions to obtain a computation result of the computation instruction. [0131] The conversion unit 13 may be configured to convert the first input data into the second input data according to the decimal point position and the identifier of data type conversion. [0132] Specifically, identifiers of data type conversion are in one-to-one correspondence with conversion manners of data type. Table 3 is a table illustrating correspondence relation between the identifier of data type conversion and the conversion manner of data type. [0133] As illustrated in Table 3, if the identifier of data type conversion is 00, the conversion manner of data type is converting the fixed-point data into fixed-point data. If the identifier of data type conversion is 01, the conversion manner of data type is converting the floating point data into floating point data)
bypass conversion of a first input operand and enable conversion of a second input operand to generate a converted second input operand having a same number representation as the first input operand, and send the first input operand and the converted second input operand to the selected functional unit; and ([0160] In an example, before the operation unit 12 of the computation device performs operations on data of an ith layer of a multi-layer neural network model, the controller unit 11 of the computation device acquires a configuration command, which may include a decimal point position and a data type of data involved in the operations. The controller unit 11 parses the configuration instruction to obtain the decimal point position and the data type of the data involved in the operations…. If the controller unit 11 has obtained the input data, it is determined whether the data type of the input data is consistent with that of the data involved in the operations. If it is determined that the data type of the input data is inconsistent with that of the data involved in the operations, the controller unit 11 sends the input data, the decimal point position, and the data type of the data involved in the operations to the conversion unit 13. The conversion unit 13 performs data type conversion on the input data according to the decimal point position and the data type of the data involved in the operations, such that the data type of the input data is consistent with that of the data involved in the operations. And then, the input data converted (i.e., second input operand) is transferred to the operation unit 12, and the primary processing circuit 101 and the secondary processing circuits 102 of the operation unit 12 perform operations on the input data converted. If it is determined that the data type of the input data (i.e., first input operand) is consistent with that of the data involved in the operations, the controller unit 11 transfers the input data to the operation unit 12, and the primary processing circuit 101 and the secondary processing circuits 102 of the operation unit 12 directly perform operations on the input data without performing data type conversion. [0012] The primary processing circuit may be configured to perform pre-processing on the second input data and to send data and the plurality of operation instructions between the plurality of secondary processing circuits and the primary processing circuit. [0013] The plurality of secondary processing circuits may be configured to perform an intermediate operation to obtain a plurality of intermediate results according to the second input data and the plurality of operation instructions sent from the primary processing circuit, and to transfer the plurality of intermediate results to the primary processing circuit)
selectively enable or bypass conversion of an output operand of the selected functional unit to generate a result of the given layer of the neural network, ([0160] In an example, before the operation unit 12 of the computation device performs operations on data of an ith layer of a multi-layer neural network model, the controller unit 11 of the computation device acquires a configuration command, which may include a decimal point position and a data type of data involved in the operations. The controller unit 11 parses the configuration instruction to obtain the decimal point position and the data type of the data involved in the operations…. If the controller unit 11 has obtained the input data, it is determined whether the data type of the input data is consistent with that of the data involved in the operations. If it is determined that the data type of the input data is inconsistent with that of the data involved in the operations, the controller unit 11 sends the input data, the decimal point position, and the data type of the data involved in the operations to the conversion unit 13. The conversion unit 13 performs data type conversion on the input data according to the decimal point position and the data type of the data involved in the operations, such that the data type of the input data is consistent with that of the data involved in the operations. And then, the input data converted is transferred to the operation unit 12, and the primary processing circuit 101 and the secondary processing circuits 102 of the operation unit 12 perform operations on the input data converted. If it is determined that the data type of the input data is consistent with that of the data involved in the operations, the controller unit 11 transfers the input data (i.e., first input operand) to the operation unit 12, and the primary processing circuit 101 and the secondary processing circuits 102 of the operation unit 12 directly perform operations on the input data without performing data type conversion)
… wherein the first input operand has one of the fixed-point number representation and the floating-point number representation, and the second input operand has the other one of the fixed-point number representation and the floating-point number representation ([0131] The conversion unit 13 may be configured to convert the first input data into the second input data according to the decimal point position and the identifier of data type conversion. [0132] Specifically, identifiers of data type conversion are in one-to-one correspondence with conversion manners of data type. Table 3 is a table illustrating correspondence relation between the identifier of data type conversion and the conversion manner of data type. [0133] As illustrated in Table 3, if the identifier of data type conversion is 00, the conversion manner of data type is converting the fixed-point data into fixed-point data. If the identifier of data type conversion is 01, the conversion manner of data type is converting the floating point data into floating point data. If the identifier of data type conversion is 10, the conversion manner of data type is converting the fixed-point data into floating point data. If the identifier of data type conversion is 11, the conversion manner of data type is converting the floating point data into fixed-point data)
While ZHANG teaches a processing unit that has an operation circuit and a conversion unit that perform tensor operations in one or more layers of a neural network in fixed-point and floating-point number representation, ZHANG does not explicitly disclose that
an operation circuit including a fixed-point circuit and a floating-point circuit … wherein one of the fixed-point circuit and the floating-point circuit is selected according to operating parameters of the neural network as a selected functional unit for use in a given layer of the neural network;
and a conversion circuit, according to the operating parameters, operative to:
However, in analogous art, MENG teaches:
and a conversion circuit, according to the operating parameters, operative to: ([0051] The controller unit 11 is connected to the operation unit 12 and the conversion unit 13 (the conversion unit may be set separately, or may be integrated in the controller unit or the operation unit. Figure 1 below shows the conversion unit next to the output of the operation unit, and the controller unit connected to the input of the operation unit)
PNG
media_image1.png
644
820
media_image1.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the filing date of the invention to have combined MENG’s teaching of a conversion unit coupled at the output and input port of the operation unit, with ZHANG’s teaching a processing unit that has an operation circuit perform tensor operations in one or more layers of a neural network in fixed-point and floating-point number representation, and a conversion unit that can convert between these two number representations, to realize, with a reasonable expectation of success, a system that has an operation unit that performs tensor operation in one or more layers of a neural network, and a conversion unit that converts between fixed-point and floating-point, as in ZHANG, and having the conversion unit coupled to the input and output port of the operation unit, as in MENG. A person of ordinary skill would have been motivated to make this combination to result in a processing unit that is flexible with data representation and is optimized for efficient data conversion and tensor operations at different processing stages.
While ZHANG and MENG teach a processing unit that has an operation circuit and a conversion unit that perform tensor operations in one or more layers of a neural network in fixed-point and floating-point number representation, ZHANG and MENG do not explicitly disclose that
an operation circuit including a fixed-point circuit and a floating-point circuit … wherein one of the fixed-point circuit and the floating-point circuit is selected according to operating parameters of the neural network as a selected functional unit for use in a given layer of the neural network;
However, in analogous art, KAWABE teaches:
an operation circuit including a fixed-point circuit … wherein one of the fixed-point circuit and the floating-point circuit is selected according to operating parameters of the neural network as a selected functional unit for use in a given layer of the neural network; ([0028] discloses that the selector 4 selects which operation to be performed: floating-point operator or fixed-point operator. [0025] The arithmetic processing device 100 includes a fixed-point operator 1 that executes an operation on a fixed-point number, a floating-point operator 2 that executes an operation on a floating-point number, a first converter 3, a selector 4, a statistical information acquirer 5, an update information generator 6, and a second converter 7. The fixed-point operator 1 includes, for example, a 16-bit multiply-accumulate operator and outputs an operation result DT1 (of, for example, 40 bits) to the selector 4) and a floating-point circuit ([0026] The floating-point operator 2 includes, for example, a 32-bit multiplier, a 32-bit divider, or the like and outputs a result DT2 (of, for example, 32 bits) of executing an operation on a floating-point number to the first converter 3. For example, the 32-bit floating-point number includes an 8-bit decimal part and may represent 256 decimal point positions)
PNG
media_image2.png
730
930
media_image2.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the filing date of the invention to have combined KAWABE’s teaching of a fixed-point operator that operates with fixed point data representation in one layer of a neural network and a floating-point operator that operates with floating point data representation in another layer of a neural network, with ZHANG and MENG’s teaching of a system comprises of one more processing unit to perform tensor operations in different layers of a neural network and conversion circuits coupled to the operation circuits, to realize, with a reasonable expectation of success, a system that has multiple circuits performs tensor operation in one or more layers of a neural network and conversion units coupled to at least one of the processing unit, as in Zhang and MENG, and uses fixed-point circuits for operations using fixed point data representation in one layer and floating-point circuits for operations using fixed point data representation, as in KAWABE. A person of ordinary skill would have been motivated to make this combination to result in an enhanced computation efficiency and flexibility when working with floating point and fixed point operations in different layers of a neural network, since some layer may benefit from the speed and power efficiency of fixed-point operations and some layer may benefits from the accuracy of floating-point operations.
Regarding claim 2, MENG further teaches:
wherein the conversion circuit, according to the operating parameters for the given layer of the neural network, is configurable to be coupled to one or both of an input port and an output port of the operation circuit ([0051] The controller unit 11 is connected to the operation unit 12 and the conversion unit 13 (the conversion unit may be set separately, or may be integrated in the controller unit or the operation unit (i.e., the conversion unit can be coupled at the output port as shown in Figure 1 below, or integrated with the operation unit at the input port ).
PNG
media_image1.png
644
820
media_image1.png
Greyscale
Regarding claim 3, ZHANG further teaches:
wherein the conversion circuit, according to the operating parameters for the given layer of the neural network, is configurable to be enabled or bypassed for one or both input conversion and output conversion ([0007] The controller unit may be configured to acquire first input data and a computation instruction, to parse the computation instruction to obtain at least one of a data conversion instruction and at least one operation instruction, where the data conversion instruction may include an opcode field and an opcode. The opcode may be configured to indicate information of a function of the data conversion instruction. The opcode field may include information of a decimal point position, a flag bit indicating a data type of the first input data, and an identifier of data type conversion. The controller unit may be further configured to transfer the opcode and the opcode field of the data conversion instruction and the first input data to the conversion unit, and to send a plurality of operation instructions to the operation unit. [0008] The conversion unit may be configured to convert the first input data into second input data according to the opcode and the opcode field of the data conversion instruction (i.e., according to the opcode and opcode field, the conversion could be bypassed or enabled)).
Regarding claim 4, ZHANG further teaches:
wherein the neural network processing unit is operative to perform hybrid-precision computing on the first input operand and the second input operand of the given layer, the first input operand and the second input operand having different number representations ([0093] The operation unit 12 may be configured to perform operations on the second input data according to the multiple operation instructions to obtain a computation result of the computation instruction. [0131] The conversion unit 13 may be configured to convert the first input data into the second input data according to the decimal point position and the identifier of data type conversion).
Regarding claim 5, MENG further teaches:
wherein the neural network processing unit is operative to perform mixed-precision computing in which computation in a first layer of the neural network and computation in a second layer of the neural network are performed in two different number representations. ([0160] In an example, before the operation unit 12 of the computation device performs operations on data of an ith layer of a multi-layer neural network model, the controller unit 11 of the computation device acquires a configuration command, which may include a decimal point position and a data type of data involved in the operations (i.e., the data type could be floating-point or fixed-point depending on the command received). The controller unit 11 parses the configuration instruction to obtain the decimal point position and the data type of the data involved in the operations)
Regarding claim 7, ZHANG further teaches:
further comprising: a buffer memory to buffer non-converted input for the converter circuit to determine, during operations of the given layer of the neural network ([0170] In an example, the operation unit 12 may be provided with a separate cache. As illustrated in FIG. 3G, the operation unit 12 may include a neuron cache unit 63 configured to buffer input neuron vector data and output neuron weight data of the secondary processing circuits 102. [0171] As illustrated in FIG. 3H, the operation unit 12 may further include a weight cache unit 64 configured to buffer weight data required by the secondary processing circuit 102 in the operation process); a scaling factor for conversion between the fixed-point number representation and the floating-point number representation ([0069] Examples of the present disclosure provide a data type. The data type may include an adjustment factor. The adjustment factor may be configured to indicate a value range and precision of the data type. [0070] The adjustment factor may include a first scaling factor. Optionally, the adjustment factor may further include a second scaling factor. The first scaling factor may be configured to indicate the precision of the data type, and the second scaling factor may be configured to adjust the value range of the data type. [0075] Scaling factors may be applied to any format of data (such as floating point data and discrete data), so as to adjust the size and precision of the data).
Regarding claim 8, ZHANG further teaches:
a buffer coupled between the converter circuit and the operation circuit (
PNG
media_image3.png
740
958
media_image3.png
Greyscale
As shown in the figure 3A above, the control unit has a Instruction Cache unit, and a storage queue unit, and the control unit is between the Conversion unit and the Operation unit. [0105] In an example, the controller unit 11 may include an instruction cache unit 110, an instruction processing unit 111, and a storage queue unit 113. [0106] The instruction cache unit 110 may be configured to store the computation instruction associated with artificial neural network operations. [0107] The instruction processing unit 111 may be configured to parse the computation instruction to obtain the data conversion instruction and the multiple operation instructions, and to parse the data conversion instruction to obtain the opcode and the opcode field of the data conversion instruction).
Regarding claim 9, while ZHANG teaches a processing unit that have multiple secondary operation circuits that can perform tensor operations in different layers of a neural network, and conversion units that convert between floating point and fixed point data, the combination of ZHANG and MENG does not explicitly disclose:
wherein the operation circuit includes the fixed-point circuit to compute a layer of the neural network in fixed-point and the floating-point circuit to compute another layer of the neural network in floating-point
However, in analogous art, KAWABE teaches:
wherein the operation circuit includes the fixed-point circuit to compute a layer of the neural network in fixed-point and the floating-point circuit to compute another layer of the neural network in floating-point (KAWABE figure 1 shows a fixed-point operator 1 that computes in fixed point and a floating-point operator 2 that computes in floating point. [0025] The arithmetic processing device 100 includes a fixed-point operator 1 that executes an operation on a fixed-point number, a floating-point operator 2 that executes an operation on a floating-point number, a first converter 3, a selector 4, a statistical information acquirer 5, an update information generator 6, and a second converter 7. The fixed-point operator 1 includes, for example, a 16-bit multiply-accumulate operator and outputs an operation result DT1 (of, for example, 40 bits) to the selector 4. [0026] The floating-point operator 2 includes, for example, a 32-bit multiplier, a 32-bit divider, or the like and outputs a result DT2 (of, for example, 32 bits) of executing an operation on a floating-point number to the first converter 3. For example, the 32-bit floating-point number includes an 8-bit decimal part and may represent 256 decimal point positions. [0021] For example, a deep neural network includes a layer for executing an operation using a fixed-point number and a layer for executing an operation using a floating-point number. In the case where the layer for executing the operation using the fixed-point number is coupled to and succeeding the layer for executing the operation using the floating-point number, a result of executing the operation on the floating-point number is converted to a fixed-point number and the fixed-point number after the conversion is input to the succeeding layer).
It would have been obvious to a person having ordinary skill in the art before the filing date of the invention to have combined KAWABE’s teaching of a fixed-point operator that operates with fixed point data representation in one layer of a neural network and a floating-point operator that operates with floating point data representation in another layer of a neural network, with the combination of ZHANG and MENG’s teaching of a processing unit that has an operation unit and multiple secondary processing unit that perform tensor operations in one or more layers of a neural network in fixed-point and floating-point number representation, to realize, with a reasonable expectation of success, a system that has an operation unit that performs tensor operation in one or more layers of a neural network using its secondary processing units, as in the combination of ZHANG and MENG, and uses fixed-point operators to compute operations in fixed point data representation in one layer and floating-point operators to compute operations in fixed point data representation in different layer of the neural network, as in KAWABE. A person of ordinary skill would have been motivated to make this combination to result in an enhanced computation efficiency in different layers of a neural network when working with floating point and fixed point data representations, since some layer may benefit from the speed and power efficiency of fixed-point operations and some layer may benefits from the accuracy of floating-point operations.
Regarding claim 10, Zhang further teaches that:
wherein the neural network processing unit is coupled to one or more processors that are operative to perform operations of one or more layers of the neural network ([0279] In an example, a main processor and a coprocessor are included in a system on chip (SOC), and the main processor may include the above-mentioned computation device. The coprocessor obtains the decimal point position of input data of the same type of the each layer in the multi-layer network model according to the above-identified method, and sends the decimal point position of the input data of the same type of the each layer in the multi-layer network model to the computation device. Alternatively, if the computation device needs to use the decimal point position of the input data of the same type of the each layer in the multi-layer network model, the decimal point position of the input data of the same type of the each layer in the multi-layer network model is obtained from the above-mentioned coprocessor).
Regarding claim 11, ZHANG further teaches:
and one or more of the conversion circuits coupled to the operation circuits ([0089] The controller unit 11 may be further configured to parse the computation instruction to obtain at least one of a data conversion instruction (i.e., one or more conversions) and at least one operation instruction, where the data conversion instruction may include an opcode field and an opcode. The opcode may be configured to indicate information of a function of the data conversion instruction. The opcode field may include information of a decimal point position, a flag bit indicating a data type of the first input data, and an identifier of data type conversion).
While the combination of ZHANG and MENG teaches a processing unit with multiple secondary processing unit to perform tensor operations in different data representations in different layers of a neural network, they do not explicitly disclose:
a plurality of operation circuits including one or more fixed-point circuits and floating- point circuits, different ones of the operation circuits operative to compute different layers of the neural network.
However, in analogous art, KAWABE teaches:
…fixed-point circuits and floating- point circuits, different ones of the operation circuits operative to compute different layers of the neural network ([0021] For example, a deep neural network includes a layer for executing an operation using a fixed-point number and a layer for executing an operation using a floating-point number. In the case where the layer for executing the operation using the fixed-point number is coupled to and succeeding the layer for executing the operation using the floating-point number, a result of executing the operation on the floating-point number is converted to a fixed-point number and the fixed-point number after the conversion is input to the succeeding layer [0025] The arithmetic processing device 100 includes a fixed-point operator 1 that executes an operation on a fixed-point number, a floating-point operator 2 that executes an operation on a floating-point number, a first converter 3, a selector 4, a statistical information acquirer 5, an update information generator 6, and a second converter 7. The fixed-point operator 1 includes, for example, a 16-bit multiply-accumulate operator and outputs an operation result DT1 (of, for example, 40 bits) to the selector 4. [0026] The floating-point operator 2 includes, for example, a 32-bit multiplier, a 32-bit divider, or the like and outputs a result DT2 (of, for example, 32 bits) of executing an operation on a floating-point number to the first converter 3. For example, the 32-bit floating-point number includes an 8-bit decimal part and may represent 256 decimal point positions).
It would have been obvious to a person having ordinary skill in the art before the filing date of the invention to have combined KAWABE’s teaching of a fixed-point operator that operates with fixed point data representation in one layer of a neural network and a floating-point operator that operates with floating point data representation in another layer of a neural network, with the combination of ZHANG and MENG’s teaching of a processing unit that has an operation unit and multiple secondary processing unit that perform tensor operations in one or more layers of a neural network in fixed-point and floating-point number representation, to realize, with a reasonable expectation of success, a system that has an operation unit with multiple secondary processing units that performs tensor operation in one or more layers of a neural network, as in the combination of ZHANG and MENG, and uses fixed-point circuits for operations in fixed point data representation in one layer and floating-point circuits for operations in fixed point data representation in different layer of the neural network, as in KAWABE. A person of ordinary skill would have been motivated to make this combination to result in an enhanced computation efficiency and flexibility when working with floating point and fixed point operations in different layers of a neural network, since some layer may benefit from the speed and power efficiency of fixed-point operations and some layer may benefits from the accuracy of floating-point operations.
Regarding claim 12, ZHANG further teaches that:
wherein the operation circuit further comprises one or more of: an adder, a subtractor, a multiplier, a function evaluator, and a multiply-and-accumulate (MAC) circuit ([0006] The operation unit may include a primary processing circuit and a plurality of secondary processing circuits. [0127] In an example, as illustrated in FIG. 3D, the primary processing circuit 101 illustrated in FIGS. 3A-3C may further include one or any combination of an activation processing circuit 1011 and an addition processing circuit 1012. [0129] The addition processing circuit 1012 may be configured to perform an addition operation or an accumulation operation. [0261] For example, the first input data may include the input data I1 and the input data I2, and the corresponding decimal point positions are P1 and P2, and P1>P2. If the operation type indicated by the above-mentioned operation instruction is an addition operation or a subtraction operation. [0293] Further, the operation instructions may be configured to perform neural network operations. The operation instruction may include a matrix operation instruction, a vector operation instruction, and a scalar operation instruction. [0294] Furthermore, the matrix operation instruction performs matrix operations in neural networks, which include operations of matrix multiplying vector, vector multiplying matrix, and matrix multiplying scalar, outer products, matrix adding matrix, matrix subtracting matrix. [0176] In an example, the operation unit 12 may include, but not limited to, one or more multipliers of a first part, one or more adders of a second part (more specifically, the adders of the second part may also constitute an addition tree), an activation function unit of a third part, and/or a vector processing unit of the fourth part.
Regarding claim 18, the combination of MENG and KAWABE teaches the limitations of claim 16. KAWABE further teaches that:
wherein output ports of one of the floating-point circuits and one of the fixed-point circuits are coupled, in parallel, to a multiplexer (KAWABE Figure 1 shows the floating-point operator and fixed-point operator coupled in parallel, and connected to a Selector 4. [0028] The selector 4 selects any of the operation results DT1 and DT3 based on a selection signal SEL and outputs the selected operation result DT1 or DT3 as an operation result DT4. The figure below shows the fixed-point operator and floating-pointer is connected by a selector).
PNG
media_image4.png
730
930
media_image4.png
Greyscale
Regarding claim 19, the combination of MENG and KAWABE teaches the limitation of claim 16. KAWABE further teaches that:
wherein the one or more conversion circuits includes a floating- point to fixed-point converter that is coupled to an input port of a fixed-point circuit or an output port of a floating-point circuit (KAWABE figure 1 shows that the first converter, a floating-point to fixed-point converter, is attached to the output of the floating point operator output. [0027] The first converter 3 converts the 32-bit operation result DT2 (floating-point number) obtained by the floating-point operator 2 to, for example, an operation result DT3 (fixed-point number) of 32 bits that are an example of a second bit width).
PNG
media_image5.png
730
930
media_image5.png
Greyscale
Regarding claim 20, MENG further teaches that:
wherein the one or more conversion circuits includes a fixed- point to floating-point converter that is coupled to an input port of a floating-point circuit or an output port of a fixed-point circuit ([0075] If artificial neural network operations have operations of multiple layers, input neurons and output neurons of the multi-layer operations do not refer to neurons in an input layer and in an output layer of the entire neural network. For any two adjacent layers in the network, neurons in a lower layer of the network forward operations are the input neurons, and neurons in an upper layer of the network forward operations are the output neurons. [0051] The controller unit 11 is connected to the operation unit 12 and the conversion unit 13 (the conversion unit may be set separately, or may be integrated in the controller unit or the operation unit). [0078] The conversion unit 13 is configured to convert the some fixed point forward output results between fixed point data and floating point data to obtain a first set of some floating point forward operation results, and send the first set of some floating point forward operation results to the operation unit).
Regarding claim 21, ZHANG teaches:
A device, comprising: one or more processors; a memory to store operating parameters of a neural network; a neural network processing unit coupled to the one or more processors and the memory, the neural network processing unit further comprising: ([Abstract]The computation device is configured to perform a machine learning computation, and includes an operation unit, a controller unit, and a conversion unit. [0015] further discloses that the computation device may further include a storage unit and a direct memory access unit)
an operation circuit including a fixed-point ([Abstract]The computation device is configured to perform a machine learning computation, and includes an operation unit, a controller unit, and a conversion unit. [0133] As illustrated in Table 3, if the identifier of data type conversion is 00, the conversion manner of data type is converting the fixed-point data into fixed-point data. If the identifier of data type conversion is 01, the conversion manner of data type is converting the floating point data into floating point data)
bypass conversion of a first input operand and enable conversion of a second input operand to generate a converted second input operand having a same number representation as the first input operand, and send the first input operand and the converted second input operand to the selected functional unit; and ([0160] In an example, before the operation unit 12 of the computation device performs operations on data of an ith layer of a multi-layer neural network model, the controller unit 11 of the computation device acquires a configuration command, which may include a decimal point position and a data type of data involved in the operations. The controller unit 11 parses the configuration instruction to obtain the decimal point position and the data type of the data involved in the operations…. If the controller unit 11 has obtained the input data, it is determined whether the data type of the input data is consistent with that of the data involved in the operations. If it is determined that the data type of the input data is inconsistent with that of the data involved in the operations, the controller unit 11 sends the input data, the decimal point position, and the data type of the data involved in the operations to the conversion unit 13. The conversion unit 13 performs data type conversion on the input data according to the decimal point position and the data type of the data involved in the operations, such that the data type of the input data is consistent with that of the data involved in the operations. And then, the input data converted (i.e., second input operand) is transferred to the operation unit 12, and the primary processing circuit 101 and the secondary processing circuits 102 of the operation unit 12 perform operations on the input data converted. If it is determined that the data type of the input data (i.e., first input operand) is consistent with that of the data involved in the operations, the controller unit 11 transfers the input data to the operation unit 12, and the primary processing circuit 101 and the secondary processing circuits 102 of the operation unit 12 directly perform operations on the input data without performing data type conversion. [0012] The primary processing circuit may be configured to perform pre-processing on the second input data and to send data and the plurality of operation instructions between the plurality of secondary processing circuits and the primary processing circuit. [0013] The plurality of secondary processing circuits may be configured to perform an intermediate operation to obtain a plurality of intermediate results according to the second input data and the plurality of operation instructions sent from the primary processing circuit, and to transfer the plurality of intermediate results to the primary processing circuit)
selectively enable or bypass conversion of an output operand of the selected functional unit to generate a result of the given layer of the neural network, ([0160] In an example, before the operation unit 12 of the computation device performs operations on data of an ith layer of a multi-layer neural network model, the controller unit 11 of the computation device acquires a configuration command, which may include a decimal point position and a data type of data involved in the operations. The controller unit 11 parses the configuration instruction to obtain the decimal point position and the data type of the data involved in the operations…. If the controller unit 11 has obtained the input data, it is determined whether the data type of the input data is consistent with that of the data involved in the operations. If it is determined that the data type of the input data is inconsistent with that of the data involved in the operations, the controller unit 11 sends the input data, the decimal point position, and the data type of the data involved in the operations to the conversion unit 13. The conversion unit 13 performs data type conversion on the input data according to the decimal point position and the data type of the data involved in the operations, such that the data type of the input data is consistent with that of the data involved in the operations. And then, the input data converted is transferred to the operation unit 12, and the primary processing circuit 101 and the secondary processing circuits 102 of the operation unit 12 perform operations on the input data converted. If it is determined that the data type of the input data is consistent with that of the data involved in the operations, the controller unit 11 transfers the input data (i.e., first input operand) to the operation unit 12, and the primary processing circuit 101 and the secondary processing circuits 102 of the operation unit 12 directly perform operations on the input data without performing data type conversion)
wherein the first input operand has one of the fixed-point number representation and the floating-point number representation, and the second input operand has the other one of the fixed-point number representation and the floating-point number representation. ([0131] The conversion unit 13 may be configured to convert the first input data into the second input data according to the decimal point position and the identifier of data type conversion. [0132] Specifically, identifiers of data type conversion are in one-to-one correspondence with conversion manners of data type. Table 3 is a table illustrating correspondence relation between the identifier of data type conversion and the conversion manner of data type. [0133] As illustrated in Table 3, if the identifier of data type conversion is 00, the conversion manner of data type is converting the fixed-point data into fixed-point data. If the identifier of data type conversion is 01, the conversion manner of data type is converting the floating point data into floating point data. If the identifier of data type conversion is 10, the conversion manner of data type is converting the fixed-point data into floating point data. If the identifier of data type conversion is 11, the conversion manner of data type is converting the floating point data into fixed-point data)
While ZHANG teaches a processing unit that has an operation circuit and a conversion unit that perform tensor operations in one or more layers of a neural network in fixed-point and floating-point number representation, ZHANG does not explicitly disclose that
an operation circuit including a fixed-point circuit and a floating-point circuit … wherein one of the fixed-point
However, in analogous art, MENG teaches:
a conversion circuit, according to the operating parameters, operative to: ([0051] The controller unit 11 is connected to the operation unit 12 and the conversion unit 13 (the conversion unit may be set separately, or may be integrated in the controller unit or the operation unit. Figure 1 below shows the conversion unit next to the output of the operation unit, and the controller unit connected to the input of the operation unit)
It would have been obvious to a person having ordinary skill in the art before the filing date of the invention to have combined MENG’s teaching of a conversion unit coupled at the output and input port of the operation unit, with ZHANG’s teaching a processing unit that has an operation circuit perform tensor operations in one or more layers of a neural network in fixed-point and floating-point number representation, and a conversion unit that can convert between these two number representations, to realize, with a reasonable expectation of success, a system that has an operation unit that performs tensor operation in one or more layers of a neural network, and a conversion unit that converts between fixed-point and floating-point, as in ZHANG, and having the conversion unit coupled to the input and output port of the operation unit, as in MENG. A person of ordinary skill would have been motivated to make this combination to result in a processing unit that is flexible with data representation and is optimized for efficient data conversion and tensor operations at different processing stages.
However, in analogous art, KAWABE teaches:
an operation circuit including a fixed-point circuit and a floating-point circuit … wherein one of the fixed-point circuit and the floating-point circuit is selected according to the operating parameters as a selected functional unit for use in a given layer of the neural network; (KAWABE figure 1 shows a fixed-point operator 1 that computes in fixed point and a floating-point operator 2 that computes in floating point. [0028] discloses that the selector 4 selects which operation to be performed: floating-point operator or fixed-point operator. [0025] The arithmetic processing device 100 includes a fixed-point operator 1 that executes an operation on a fixed-point number, a floating-point operator 2 that executes an operation on a floating-point number, a first converter 3, a selector 4, a statistical information acquirer 5, an update information generator 6, and a second converter 7. The fixed-point operator 1 includes, for example, a 16-bit multiply-accumulate operator and outputs an operation result DT1 (of, for example, 40 bits) to the selector 4) and a floating-point circuit ([0026] The floating-point operator 2 includes, for example, a 32-bit multiplier, a 32-bit divider, or the like and outputs a result DT2 (of, for example, 32 bits) of executing an operation on a floating-point number to the first converter 3. For example, the 32-bit floating-point number includes an 8-bit decimal part and may represent 256 decimal point positions)
It would have been obvious to a person having ordinary skill in the art before the filing date of the invention to have combined KAWABE’s teaching of a fixed-point operator that operates with fixed point data representation in one layer of a neural network and a floating-point operator that operates with floating point data representation in another layer of a neural network, with ZHANG and MENG’s teaching of a system comprises of one more processing unit to perform tensor operations in different layers of a neural network and conversion circuits coupled to the operation circuits, to realize, with a reasonable expectation of success, a system that has multiple circuits performs tensor operation in one or more layers of a neural network and conversion units coupled to at least one of the processing unit, as in Zhang and MENG, and uses fixed-point circuits for operations using fixed point data representation in one layer and floating-point circuits for operations using fixed point data representation, as in KAWABE. A person of ordinary skill would have been motivated to make this combination to result in an enhanced computation efficiency and flexibility when working with floating-point and fixed-point operations in different layers of a neural network, since some layer may benefit from the speed and power efficiency of fixed-point operations and some layer may benefits from the accuracy of floating-point operations.
Regarding claim 22, ZHANG in view of MENG and further in view of KAWABE teaches:
The device of claim 21, wherein output ports of the floating-point circuit and the fixed-point circuit are coupled, in parallel, to a multiplexer, which selects an output from one of the fixed-point circuit and the floating-point circuit. (KAWABE Figure 1 shows the floating-point operator and fixed-point operator coupled in parallel, and connected to a Selector 4. [0028] The selector 4 selects any of the operation results DT1 and DT3 based on a selection signal SEL and outputs the selected operation result DT1 or DT3 as an operation result DT4. The figure below shows the fixed-point operator and floating-pointer is connected by a selector)
Regarding claim 23, KAWABE teaches:
The device of claim 21, wherein the conversion circuit includes a floating-point to fixed-point converter that is coupled to an input port of the fixed-point circuit or an output port of the floating-point circuit. (KAWABE figure 1 shows that the first converter, a floating-point to fixed-point converter, is attached to the output of the floating point operator output. [0027] The first converter 3 converts the 32-bit operation result DT2 (floating-point number) obtained by the floating-point operator 2 to, for example, an operation result DT3 (fixed-point number) of 32 bits that are an example of a second bit width).
Regarding claim 24, MENG teaches:
The device of claim 21, wherein the conversion circuit includes a fixed-point to floating-point converter that is coupled to an input port of the floating-point circuit or an output port of the fixed-point circuit. ([0075] If artificial neural network operations have operations of multiple layers, input neurons and output neurons of the multi-layer operations do not refer to neurons in an input layer and in an output layer of the entire neural network. For any two adjacent layers in the network, neurons in a lower layer of the network forward operations are the input neurons, and neurons in an upper layer of the network forward operations are the output neurons. [0051] The controller unit 11 is connected to the operation unit 12 and the conversion unit 13 (the conversion unit may be set separately, or may be integrated in the controller unit or the operation unit). [0078] The conversion unit 13 is configured to convert the some fixed point forward output results between fixed point data and floating point data to obtain a first set of some floating point forward operation results, and send the first set of some floating point forward operation results to the operation unit).
Regarding claim 25, ZHANG teaches:
The device of claim 21, further comprising: a buffer memory to buffer non-converted input for the converter circuit to determine, during operations of the given layer of the neural network, ([0170] In an example, the operation unit 12 may be provided with a separate cache. As illustrated in FIG. 3G, the operation unit 12 may include a neuron cache unit 63 configured to buffer input neuron vector data and output neuron weight data of the secondary processing circuits 102. [0171] As illustrated in FIG. 3H, the operation unit 12 may further include a weight cache unit 64 configured to buffer weight data required by the secondary processing circuit 102 in the operation process) a scaling factor for conversion between the fixed-point number representation and the floating-point number representation. ([0069] Examples of the present disclosure provide a data type. The data type may include an adjustment factor. The adjustment factor may be configured to indicate a value range and precision of the data type. [0070] The adjustment factor may include a first scaling factor. Optionally, the adjustment factor may further include a second scaling factor. The first scaling factor may be configured to indicate the precision of the data type, and the second scaling factor may be configured to adjust the value range of the data type. [0075] Scaling factors may be applied to any format of data (such as floating point data and discrete data), so as to adjust the size and precision of the data).
Claim(s) 6 is rejected under 35 U.S.C. 103 as being unpatentable over ZHANG, in view of MENG, in view of KAWABE, as applied to claim 1 above, and in further view of Dawwd, “Time Sharing Based Parallel Implementation of CNN on Low Cost FPGA” (hereafter DAWWD).
Regarding claim 6, while ZHANG teaches the neural network processing unit that can perform tensor operations with floating point and fixed point data representations, the combination of ZHANG, MENG and KAWABE does not explicitly disclose:
the neural network processing unit is time-shared among multiple layers of the neural network by operating on one layer at a time
However, in analogous art, DAWWD teaches:
the neural network processing unit is time-shared among multiple layers of the neural network by operating on one layer at a time (Figure 4 below shows the time-sharing operation. (Section IV, A. Time Sharing Processing Flow) Therefore, in the proposed architecture of CNN, a repetitively utilized neuron circuits are performed using time-sharing operation. Each group of receptive field vectors are processed in a parallel manner spending a time slice that is enough to complete their respective computations, then another group which may be located in successive layers (except the feedforward layer) takes its role, and then return to the fist layer and so on, until the all neurons’ outputs in the last simple layer are calculated. The CNN architecture shown in Fig. 2 can share the time according to the above approach. Fig. 4 shows the time sharing operation of this architecture).
PNG
media_image6.png
490
1000
media_image6.png
Greyscale
It would have been obvious to a person having ordinary skill in the art before the filing date of the invention to have combined DAWWD’s teaching of operations being time-shared among the layers of a neural network and operates on one layer at a time, with the combination of ZHANG, MENG, and KAWABE’s teaching of a processing unit that has an operation circuit perform tensor operations in one or more layers of a neural network in fixed-point and floating-point number representation, and a conversion unit coupled to the input and output port of the operation circuit that can convert between these two number representations, to realize, with a reasonable expectation of success, a system that has an operation unit that performs tensor operation in one or more layers of a neural network, and a conversion unit that converts between fixed-point and floating-point, as in the combination of ZHANG, MENG, and KAWABE, and operations are time-shared among the layers by operating on one layer at a time, as in DAWWD. A person of ordinary skill would have been motivated to make this combination to reduce the resources required for implementation (DAWWD[Abstract]), and enable more efficient tensor operation while minimizing hardware requirement.
Response to Arguments
Response to Arguments under 35 U.S.C. 101
Applicant’s arguments, see [Remarks, pages 7-8], filed 06/06/2025, with respect to 35 U.S.C. 101 rejection have been fully considered and are persuasive. The 35 U.S.C. 101 rejection of claims 1-12 and 18-25 have been withdrawn.
Response to Arguments under 35 U.S.C. 103
Arguments: Applicant asserts that (a) Zhang does not disclose the amended elements, rather, Zhang discloses converting the first input data into the second input data, and transferring the second input data to the operating unit, and (b) Meng does not disclose “floating-point circuit” and “fixed-point circuit” and also does not disclose “bypass conversion …” and “selectively enable or bypass conversion of an output operand …” as recited in amended claim 1 [Remarks, page 8-9].
Examiner’s Response: Examiner respectfully disagrees. The applicant canceled claims 13-17 and incorporates the limitations of the canceled claim into Claim 1. Zhang in view of Meng teaches the canceled claims 13-14, Zhang in view of Meng and further in view of Dawwd teaches the canceled claim 15, Meng in view of Kawabe teaches canceled claim 16, and Meng in view of Kawabe and further in view of Shimokawa teaches the canceled claim 17. Therefore, Claim 1 is rejected under the same rationale as of the canceled claims.
Regarding (a), the examiner first asserts that Zhang clearly disclose “converts from a first number representation to a second number representation” in paragraphs [0160] and [0012]-[0013] and “bypass conversion of a first input operand and enable conversion of a second input operand to generate a converted second input operand having a same number of representations as the first input operand”, and “selectively enable or bypass conversion of an output operand of the selected functional unit to generate a result of the given layer of the neural network” in para [0160]. “If the controller unit 11 has obtained the input data, it is determined whether the data type of the input data is consistent with that of the data involved in the operations. If it is determined that the data type of the input data is inconsistent with that of the data involved in the operations, the controller unit 11 sends the input data, the decimal point position, and the data type of the data involved in the operations to the conversion unit 13 … If it is determined that the data type of the input data is consistent with that of the data involved in the operations, the controller unit 11 transfers the input data (i.e., first input operand) to the operation unit 12,” indicates that the selection of whether to convert the input data or not is performed by the controller unit. Even though Zhang does not specifically disclose the amended limitations of “one of the fixed-point functional unit and the floating-point functional unit is elected according to operating parameters of the neural network as a selected functional unit for use in a given layer of the neural network”, Kawabe teaches “one of the fixed-point functional unit and the floating-point functional unit is elected according to operating parameters of the neural network as a selected functional unit for use in a given layer of the neural network” in paragraphs [0028] and it is obvious for someone with ordinary skill in the art to combine the teachings of Zhang, Meng and Kawabe to implement the present invention. See 35 U.S.C. 103 rejection for details.
Regarding (b), although Zhang in view of Meng does not specifically disclose the fixed-point circuits and floating-point circuits, Kawabe et al. Pub. Num: US2020/0192633 A1 clearly disclose in paragraph [0025] the fixed-point operator including a 16-bit multiply-accumulate operator and floating-point operator including a 32-bit multiplier, a 32-bit divider, or the like, which are circuits. It is obvious for someone who knows the art to combine the teachings of Zhang, Meng and Kawabe to implement the present invention, because Zhang discloses converting the first number representation (fixed-point number representation) and the second number representation (floating-point representation), Meng discloses performing floating point operations and fixed point operations, and Kawabe discloses performing the operations using circuits to improve the power efficiency and speed of the neural network operations.
Accordingly, arguments regarding claim 1 are not persuasive. Similarly, arguments regarding claims 2-5, 7-8, 10 and 12 are not persuasive at least for the same reasons as claim 1.
Claim 6
Claim 6 depends from claim 1, and arguments regarding claim 1 are not persuasive, as discussed above. Accordingly, arguments regarding claim 6 are not persuasive.
Claims 9 and 11
Claims 9 and 11 depend from claim 1, and arguments regarding claim 1 are not persuasive, as discussed above. Kawabe clearly discloses selecting fixed-point and floating-point operations, and the limitations of “bypass conversion … and enable conversion …” are taught by the combination of Zhang and Meng. Accordingly, arguments regarding claims 9 and 11 are not persuasive.
Claims 13-17
Claims 13-17 have been canceled. 35 U.S.C. 103 rejection of claims 13-17 have been withdrawn.
Claims 18-20
Claims 18-20 depend from claim 1, and arguments regarding claim 1 are not persuasive, as discussed above. Kawabe clearly discloses selecting fixed-point and floating-point operations, and the limitations of “bypass conversion … and enable conversion …” are taught by the combination of Zhang and Meng. Accordingly, arguments regarding claim 18-20 are not persuasive.
New Claims 21-25
New claims 21-25 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Meng and further in view of Kawabe. See 35 U.S.C. 103 rejections above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US-20200202195-A1 (This prior art is pertinent as it discloses converting floating point representations to fixed point representations and vice versa to perform neural network operations)
US-20200210838-A1 (This prior art is pertinent as it discloses converting floating point representations to fixed point representations)
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUN KWON whose telephone number is (571)272-2072. The examiner can normally be reached Monday – Friday 8:00AM – 5:00PM ET
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JUN KWON/Examiner, Art Unit 2127
/ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127