Detailed Action
Claims 1-20 are currently pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. KR10-2024-0016687, filed on 02/02/2024.
Election/Restrictions
Restriction to one of the following inventions is required under 35 U.S.C. 121:
I. Claims 1-19, drawn to a method and a non-volatile computer-readable storage medium storing instructions for converting function of a first neural network model into at least one graph module, analyzing a relationship between inputs and outputs of the graph module, generating a second neural network model in a form of a directed acyclic graph using the graph module, adding markers to the graph module, generating calibration data for the graph module, and determining a scale value and an offset value of the graph module, classified in G06N20/00.
II. Claim 20, drawn to a neural processing unit comprising: a processing element circuitry configured to receive a first input feature map quantized to a first bitwidth and a first weight quantized to a second bitwidth, and a special function unit circuitry configured to receive the first output feature map as input, converts the first feature map to a second feature map, and output a third feature map, classified in G06N3/063.
The inventions are independent or distinct, each from the other because:
Inventions I and II are directed to related processes. The related inventions are distinct if: (1) the inventions as claimed are either not capable of use together or can have a materially different design, mode of operation, function, or effect; (2) the inventions do not overlap in scope, i.e., are mutually exclusive; and (3) the inventions as claimed are not obvious variants. See MPEP § 806.05(j). In the instant case, the inventions as claimed
(1) have different modes of operations and functions (Group I recites converting a first neural network into graph modules, analyzing relationship between input and output of the graph modules, generating a graph representation of the first NN, adding markers to the graph representation, and generating a new NN based on the marker and calibration data. Group II recites a neural processing unit comprising a processing element circuitry configured to receive a first input feature map and a special function unit circuitry to receive the first output feature map and converts it to a second feature map and a third feature map.)
(2) do not overlap in scope (each invention of groups I and II contain features which do not appear in any of the other inventions) and
(3) the inventions of groups I and II are not obvious variants of one another.
Furthermore, the inventions as claimed do not encompass overlapping subject matter and there is nothing of record to show them to be obvious variants.
Restriction for examination purposes as indicated is proper because all the inventions listed in this action are independent or distinct for the reasons given above and there would be a serious search and/or examination burden if restriction were not required because one or more of the following reasons apply:
Groups I and II have different, non-overlapping series of steps which would require a serious search and examination burden.
Applicant is reminded that upon the cancelation of claims to a non-elected invention, the inventorship must be corrected in compliance with 37 CFR 1.48(a) if one or more of the currently named inventors is no longer an inventor of at least one claim remaining in the application. A request to correct inventorship under 37 CFR 1.48(a) must be accompanied by an application data sheet in accordance with 37 CFR 1.76 that identifies each inventor by his or her legal name and by the processing fee required under 37 CFR 1.17(i).
Applicant is advised that the reply to this requirement to be complete must include (i) an election of an invention to be examined even though the requirement may be traversed (37 CFR 1.143) and (ii) identification of the claims encompassing the elected invention.
The election of an invention may be made with or without traverse. To reserve a right to petition, the election must be made with traverse. If the reply does not distinctly and specifically point out supposed errors in the restriction requirement, the election shall be treated as an election without traverse. Traversal must be presented at the time of election in order to be considered timely. Failure to timely traverse the requirement will result in the loss of right to petition under 37 CFR 1.144. If claims are added after the election, applicant must indicate which of these claims are readable upon the elected invention.
Should applicant traverse on the ground that the inventions are not patentably distinct, applicant should submit evidence or identify such evidence now of record showing the inventions to be obvious variants or clearly admit on the record that this is the case. In either instance, if the examiner finds one of the inventions unpatentable over the prior art, the evidence or admission may be used in a rejection under 35 U.S.C. 103 or pre-AIA 35 U.S.C. 103(a) of the other invention.
During a telephone conversation with JIHUN KIM, Reg. No. 80,730 on 22 July 2026, a provisional election was made without traverse to prosecute the invention of Group I, claims 1-19. Affirmation of this election must be made by applicant in replying to this Office action. Claim 20 is withdrawn from further consideration by the examiner, 37 CFR 1.142(b), as being drawn to a non-elected invention.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 6, 8, 13, and 17 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1, 9, and 10-12 of copending Application No. 18/639,195 (reference application). Although the claims at issue are not identical, they are not patentably distinct from each other because the claim limitations in Co-pending application 18/639,195 are substantially similar as highlighted in the table below. The claims of the instant application are anticipated by the claims of co-pending application 18/639,195.
This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented.
Instant Application: 18/603,346
Corresponding Application: 18/639,195
Claim 1. A method comprising:
converting at least one function or at least one function call instruction of a first neural network (NN) model into at least one graph module;
analyzing a relationship between one or more inputs and one or more outputs of the at least one graph module;
generating a second neural network (NN) model in a form of a directed acyclic graph (DAG) using the at least one graph module corresponding to the at least one function or the at least one function call instruction of the first NN model, by mapping the one or more inputs and the one or more outputs of the at least one graph module to each other based on the relationship;
adding at least one marker to the at least one graph module in the second NN model;
generating calibration data for each of the at least one graph module by collecting input values or output values of each of the at least one graph module using the at least one marker; and
determining, based on the calibration data, a scale value and an offset value of each of the at least one graph module applicable to the second NN model.
Claim 1. A method comprising:
converting a plurality of functions or function call instructions of a first neural network (NN) model into a plurality of graph modules;
analyzing a relationship between one or more inputs and one or more outputs of the plurality of graph modules;
generating a second neural network (NN) model in a form of a directed acyclic graph (DAG) using the plurality of graph modules corresponding to the first NN model, by mapping the one or more inputs and the one or more outputs of the plurality of graph modules to each other based on the relationship;
adding a plurality of markers to the plurality of graph modules in the second NN model;
generating calibration data by collecting input values and output values of each of the plurality of graph modules using the plurality of markers;
determining, based on the calibration data, a scale value and an offset value applicable to the second NN model; and
determining, for each graph module of the second NN model, an optimal value for the scale value or the offset value by performing a quantization simulation for one or more candidates among optimization candidates of the scale value or the offset value.
Claim 6: The method of claim 1,
wherein the scale value and the offset value are obtained by an equation below,
s
c
a
l
e
=
m
a
x
-
m
i
n
2
b
i
t
w
i
d
t
h
-
1
,
O
f
f
s
e
t
=
-
m
i
n
s
c
a
l
e
'
where max denotes a maximum value among the input values and output values collected for the calibration data, min denotes a minimum value among the input values and output values collected for the calibration data, and bitwidth denotes a target quantization bitwidth.
Claim 9: The method of claim 1,
wherein the scale value and the offset value are obtained by an equation below,
s
c
a
l
e
=
m
a
x
-
m
i
n
2
b
i
t
w
i
d
t
h
-
1
,
O
f
f
s
e
t
=
-
m
i
n
s
c
a
l
e
'
where max means a maximum value among the input values and output values collected for the calibration data, min means a minimum value among the input values and output values collected for the calibration data, and bitwidth means a target quantization bitwidth.
Claim 8: The method of claim 1, wherein a convolution operation in the second NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
f
p
=
[
f
e
a
t
u
r
e
i
n
f
p
-
o
f
s
f
]
×
s
f
+
o
f
⊗
[
w
e
i
g
h
t
f
p
s
w
]
×
s
w
where feature_infp represents an input feature map parameter in a form of floating-point, weightfp represents a weight parameter in a form of floating-point, of represents the offset value for an input feature map, sf represents the scale value for the input feature map, sw represents the scale value for a weight, and ⌊ ⌋ represents round and clip operations.
Claim 10: The method of claim 1, wherein a convolution operation in the second NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
f
p
=
[
f
e
a
t
u
r
e
i
n
f
p
-
o
f
s
f
]
×
s
f
+
o
f
⊗
[
w
e
i
g
h
t
f
p
s
w
]
×
s
w
where feature_infp represents an input feature map parameter in a form of floating-point, weightfp represents a weight parameter in a form of floating-point, of represents the offset value for an input feature map, sf represents the scale value for the input feature map, sw represents the scale value for a weight, and ⌊ ⌋ represents round and clip operations.
Claim 13: The method of claim 1,
further comprising: generating, based on the scale value, a third neural network (NN) model comprising a quantized weight parameter in a form of integer, based on the second NN model.
Claim 11: The method of claim 1,
further comprising: generating, based on optimal values of the scale value and the offset value, a third neural network (NN) model comprising a quantized weight parameter in a form of integer, based on the second NN model.
Claim 17: The method of claim 1, wherein a convolution operation in a third NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
i
n
t
=
f
e
a
t
u
r
e
_
i
n
i
n
t
⊗
w
e
i
g
h
t
i
n
t
where feature_outint denotes an output feature map parameter in a form of integer, feature_inint denotes an input feature map parameter in a form of integer, and weightint denotes a weight parameter in a form of integer.
Claim 12: The method of claim 11, wherein a convolution operation in the third NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
i
n
t
=
f
e
a
t
u
r
e
_
i
n
i
n
t
⊗
w
e
i
g
h
t
i
n
t
where feature_outint represents an output feature map parameter in a form of integer, feature_inint represents an input feature map parameter in a form of integer, and weightint represents a weight parameter in a form of integer.
Claim 19 is provisionally rejected on the ground of nonstatutory double patenting as being obvious under Claim 1 of 18/639,195 in view of Miret et al. (“Neuroevolution-Enhanced Multi-Objective Optimization for Mixed-Precision Quantization”, 2022, hereinafter ‘Miret’).
Claim 19. A non-volatile computer-readable storage medium storing instructions, when executed by one or more processors, causing the one or more processors to perform a method comprising:
converting at least one function or at least one function call instruction of a first neural network (NN) model into at least one graph module;
analyzing a relationship between one or more inputs and one or more outputs of the at least one graph module;
generating a second neural network (NN) model in a form of a directed acyclic graph (DAG) using the at least one graph module corresponding to the at least one function or the at least one function call instruction of the first NN model, by mapping the one or more inputs and the one or more outputs of the at least one graph module to each other based on the relationship;
adding at least one marker to the at least one graph module in the second NN model;
generating calibration data for each of the at least one graph module by collecting input values or output values of each of the at least one graph module using the at least one marker; and
determining, based on the calibration data, a scale value and an offset value of each of the at least one graph module applicable to the second NN model.
Claim 1. A method comprising:
converting a plurality of functions or function call instructions of a first neural network (NN) model into a plurality of graph modules;
analyzing a relationship between one or more inputs and one or more outputs of the plurality of graph modules;
generating a second neural network (NN) model in a form of a directed acyclic graph (DAG) using the plurality of graph modules corresponding to the first NN model, by mapping the one or more inputs and the one or more outputs of the plurality of graph modules to each other based on the relationship;
adding a plurality of markers to the plurality of graph modules in the second NN model;
generating calibration data by collecting input values and output values of each of the plurality of graph modules using the plurality of markers;
determining, based on the calibration data, a scale value and an offset value applicable to the second NN model; and
determining, for each graph module of the second NN model, an optimal value for the scale value or the offset value by performing a quantization simulation for one or more candidates among optimization candidates of the scale value or the offset value.
However, the reference application 18/639,195 does not suggest or disclose:
A non-volatile computer-readable storage medium storing instructions, when executed by one or more processors, causing the one or more processors to perform
Miret teaches:
A non-volatile computer-readable storage medium storing instructions, when executed by one or more processors, causing the one or more processors to perform ([Miret, ABSTRACT, lines 1-7] discloses reducing memory usage using the disclosed methods. This indicates that the method is performed using an ordinary computing device)
It would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to use the non-volatile computer-readable storage medium storing instructions, when executed by one or more processors, causing the one or more processors to perform a method of Miret to improve the performance of the neural network conversion method of the present invention. The suggestion and/or motivation for doing so is to enable the system to store input and output values of each graph modules and the neural networks.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6, 9-14 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Miret et al. (“Neuroevolution-Enhanced Multi-Objective Optimization for Mixed-Precision Quantization”, 2022, hereinafter ‘Miret’) in view of Lucas et al. (US 20250111201 A1, hereinafter ‘Lucas’) and further in view of Jang et al. (US 20220206698 A1, hereinafter ‘Jang’).
Regarding claim 1, Miret teaches:
A method comprising: converting at least one function or at least one function call instruction of a first neural network (NN) model into at least one graph module; ([Miret, page 1059, left col, 3 METHOD, lines 1-22] and [Figure 2] discloses transforming the workload into a sequential graph by declaring each quantizable operation in the workload as a node and build the edges by connecting the nodes sequentially)
analyzing a relationship [Miret, page 1059, left col, 3 METHOD, lines 1-22] and [Figure 2] discloses transforming the workload (i.e., first NN) into a sequential graph by declaring each quantizable operation in the workload as a node and build the edges by connecting the nodes sequentially (analyzing a relationship of nodes and edges))
generating a second neural network (NN) model in a form of a [Miret, page 1059, left col, 3 METHOD, lines 1-22] and [Figure 2] discloses transforming the workload (i.e., first NN) into a sequential graph (i.e., second neural network model) by declaring each quantizable operation in the workload as a node and build the edges by connecting the nodes sequentially (analyzing a relationship of nodes and edges). Each graph node represents computational operation in the targeted workload)
generating calibration data for each of the at least one graph module by collecting input values or output values of each of the at least one graph module [Miret, page 1059, right col, lines 2-19 and Equation (1)] teaches calculating q(x;b) by inputting the activation (input) tensor of a layer into the quantization function q which is the calibration data for the layer. The quantizer q maps the element x to a quantized value corresponding to one of the integers {-2^b-1, … 2^(b-1}-1} and is parametrized by the bit width b. [page 1059, right col, lines 24-31] The bit width is selected during the quantization)
determining, based on the calibration data, a scale value and an offset value of each of the at least one graph module applicable to the second NN model. ([Miret, page 1059, right col, lines 2-19 and Equation (1)] teaches calculating a scale and zero point z (i.e., offset) value by using xmax, xmin, weight (parameter) and activation (input) tensor of a layer, and scale values. [page 1059, right col, lines 24-31] The bit width b used to calculate the scale s and the zero point z is selected during the quantization)
However, Miret does not specifically disclose:
analyzing a relationship between one or more inputs and one or more outputs of the at least one graph module;
generating a second neural network (NN) model in a form of a directed acyclic graph (DAG) using the at least one graph module corresponding to the at least one function or the at least one function call instruction of the first NN model, by mapping the one or more inputs and the one or more outputs of the at least one graph module to each other based on the relationship;
adding at least one marker to the at least one graph module in the second NN model;
generating calibration data for each of the at least one graph module by collecting input values or output values of each of the at least one graph module using the at least one marker;
Lucas teaches:
analyzing a relationship between one or more inputs and one or more outputs of the at least one graph module; ([Lucas, 0042]-[0043] The subgraph generation module 321 converts individual layers within a neural network into a subgraph representation by constructing three types of nodes: input nodes, hidden nodes, and output nodes, labeling the nodes, identifying layer number by determining the shortest path from the neuron to an input neuron, and assigning weight values to the edges connecting corresponding nodes)
generating a second neural network (NN) model in a form of a directed acyclic graph (DAG) using the at least one graph module corresponding to the at least one function or the at least one function call instruction of the first NN model, by mapping the one or more inputs and the one or more outputs of the at least one graph module to each other based on the relationship; ([Lucas, 0042]-[0043] The subgraph generation module 321 converts individual layers within a neural network into a subgraph representation by constructing three types of nodes: input nodes, hidden nodes, and output nodes, labeling the nodes, identifying layer number by determining the shortest path from the neuron to an input neuron, assigning weight values to the edges connecting corresponding nodes, and generate bias nodes and incorporate bias information into the subgraphs, and [0046] combining generated subgraphs into a comprehensive graph representations. [0040] and [0042] collectively disclose that the neural network is feed-forward and directed, which indicates that the neural network is a directed acyclic graph. [0054] indicates that each node activations are functions of input/output and weight parameter)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Miret and Lucas to use the method of analyzing relationships between input and output of a plurality of graph modules (layers) of Lucas to implement the neural network conversion method of Miret. The suggestion and/or motivation to do so is to improve the efficiency of the neural network-graph conversion method, as such layer-by-layer approach allows simultaneous processing [Lucas, 0025].
However, Miret in view of Lucas do not specifically disclose:
adding at least one marker to the at least one graph module in the second NN model;
generating calibration data for each of the at least one graph module by collecting input values or output values of each of the at least one graph module using the at least one marker;
Jang teaches:
adding at least one marker to the at least one graph module in the second NN model; ([0077] The graph IR generator 210 generate the graph intermediate representations by converting the neural network. [0080]-[0081] The checkpoint generator 231 generates checkpoints (i.e., marker) in at least one of the layer included in the neural network, and indicate data remaining in the memory among intermediate result values calculated by the layer included in the neural network. The layer is interpreted as the module in the NN model)
generating calibration data for each of the at least one graph module by collecting input values or output values of each of the at least one graph module using the at least one marker; ([0087]-[0088] discloses training a neural network by using re-calculation, which is performed based on stored intermediate values. [0080]-[0081] The checkpoint generator 231 generates checkpoints (i.e., marker) in at least one of the layer included in the neural network, and indicate data remaining in the memory among intermediate result values calculated by the layer included in the neural network. The intermediate result values are the collected output values for each of the graph module)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Miret, Lucas and Jang to use the method of adding markers to the graph module of Jang to implement the neural network conversion method of Miret. The suggestion and/or motivation to do so is to improve the efficiency of the quantized neural network training process by increasing the utilization and throughput and thereby improving the overall learning rate [Jang, 0091].
Regarding claim 2, Miret in view of Lucas teaches:
The method of claim 1, wherein each of the at least one graph module included in the DAG of the second NN model is connected to a corresponding graph module based on the relationship. ([Lucas, 0042]-[0043] The subgraph generation module 321 converts individual layers within a neural network into a subgraph representation by constructing three types of nodes: input nodes, hidden nodes, and output nodes, labeling the nodes, identifying layer number by determining the shortest path from the neuron to an input neuron, assigning weight values to the edges connecting corresponding nodes, and generate bias nodes and incorporate bias information into the subgraphs, and [0046] combining generated subgraphs into a comprehensive graph representations. [0040] and [0042] collectively disclose that the neural network is feed-forward and directed, which indicates that the neural network is a directed acyclic graph. [0054] indicates that each node activations are functions of input/output and weight parameter)
Regarding claim 3, Miret in view of Lucas and further in view of Jang teaches:
The method of claim 1, wherein the adding the at least one marker further includes: adding each marker to one or more of the at least one graph module in the second NN model such that the one or more of the at least one graph module are connected to the respective marker. ([Jang, 0077] The graph IR generator 210 generate the graph intermediate representations by converting the neural network. [0080]-[0081] The checkpoint generator 231 generates checkpoints (i.e., marker) in at least one of the layer included in the neural network, and indicate data remaining in the memory among intermediate result values calculated by the layer included in the neural network. The layer is interpreted as the module in the NN model)
Regarding claim 4, Miret in view of Lucas and further in view of Jang teaches:
The method of claim 1, wherein the calibration data is generated by inputting a calibration dataset into the second NN model. ([Jang, 0077] The graph IR generator 210 generate the graph intermediate representations by converting the neural network. [0080]-[0081] The checkpoint generator 231 generates checkpoints (i.e., marker) in at least one of the layer included in the neural network, and indicate data remaining in the memory among intermediate result values (calibration data which is used to calibrate the neural network) calculated by the layer included in the neural network. The layer is interpreted as the module in the NN model)
Regarding claim 5, Miret teaches:
The method of claim 1, wherein the calibration data includes a maximum value and a minimum value. ([Miret, page 1059, right col, lines 2-19 and Equation (1)] teaches calculating a scale and zero point z (i.e., offset) value by using xmax, xmin and scale values. The values are determined based on quantized values generated by inputting the elements x of a tensor x to a quantization function q)
Regarding claim 6, Miret teaches:
The method of claim 1,
wherein the scale value and the offset value are obtained by an equation below,
s
c
a
l
e
=
m
a
x
-
m
i
n
2
b
i
t
w
i
d
t
h
-
1
,
O
f
f
s
e
t
=
-
m
i
n
s
c
a
l
e
'
where max denotes a maximum value among the input values and output values collected for the calibration data, min denotes a minimum value among the input values and output values collected for the calibration data, and bitwidth denotes a target quantization bitwidth. ([Miret, page 1059, right col, lines 2-19 and Equation (1)] teaches calculating a scale and zero point z (i.e., offset) value by using xmax, xmin and scale values. The values are determined based on quantized values generated by inputting the elements x of a tensor x to a quantization function q)
Regarding claim 9, Miret in view of Lucas teaches:
The method of claim 1, wherein a convolution operation in the second NN model is implemented using the at least one graph module only. ([Lucas, 0040] discloses that the network being converted may include convolutional layers. The architecture extraction module 310 extract parameters from the layer, and [0041] the neural network graph pipeline 320 transforms the layer to graph representations specific to each layer)
Regarding claim 10, Miret in view of Lucas teaches:
The method of claim 1, wherein the at least one function or the at least one function call instruction converted to the at least one graph module include: at least one of add function, subtract function, multiply function, divide function, slice function, concatenation function, tensor view function, reshape function, transpose function, softmax function, permute function, chunk function, split function, clamp function, flatten function, tensor mean function, and sum function. ([Lucas, 0040] discloses that the network being converted may include convolutional layers. The architecture extraction module 310 extract parameters from the layer, and [0041] the neural network graph pipeline 320 transforms the layer to graph representations specific to each layer. Convolution operation performs element-wise multiplication followed by summation. Therefore, converting the convolutional layer to a graph representation converts at least add and multiply functions included in the layer)
Regarding claim 11, Miret teaches:
The method of claim 1, wherein weight parameters and input feature map parameters of the first NN model and the second NN model are in a form of floating-points having a length of one of 16-bits to 32-bits. ([Miret, page 1059, right col, lines 2-29] discloses quantizing the input weight and activation tensors of 32-bit floating point precision by scaling the saturated tensor by s, adds the offset z and finally rounds the resulting tensor elements to the nearest integer. As a result, a configuration B consisting of weights and activations (i.e., third neural network) for each layer is determined)
Regarding claim 12, Miret in view of Lucas teaches:
The method of claim 1, wherein the first NN model and the second NN model are in PyTorchTM format. ([0065] shows that the machine-learning applications used by the system includes PyTorch)
Regarding claim 13, Miret teaches:
The method of claim 1, further comprising: generating, based on the scale value, a third neural network (NN) model comprising a quantized weight parameter in a form of integer, based on the second NN model. ([Miret, page 1059, right col, lines 2-29] discloses quantizing the input weight and activation tensors of 32-bit floating point precision by scaling the saturated tensor by s, adds the offset z and finally rounds the resulting tensor elements to the nearest integer. As a result, a configuration B consisting of weights and activations (i.e., third neural network) for each layer is determined)
Regarding claim 14, Miret teaches:
The method of claim 1, further comprising: generating, based on the scale value and the offset value, a third NN model comprising a quantized weight parameter in a form of integer, based on the second NN model. ([Miret, page 1059, right col, lines 2-29] discloses quantizing the input weight and activation tensors of 32-bit floating point precision by scaling the saturated tensor by s, adds the offset z and finally rounds the resulting tensor elements to the nearest integer. As a result, a configuration B consisting of weights and activations (i.e., third neural network) for each layer is determined)
Regarding claim 19, it is an apparatus claim which recites the similar features as the method claim 1, and is rejected under the same rationale as claim 1. Additional limitations of claim 19 not addressed in claim 1 are addressed below.
Miret teaches:
A non-volatile computer-readable storage medium storing instructions, when executed by one or more processors, causing the one or more processors to perform a method comprising: ([Miret, ABSTRACT, lines 1-7] discloses reducing memory usage using the disclosed methods. This indicates that the method is performed using an ordinary computing device)
Claims 7 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Miret in view of Lucas in view of Jang and further in view of Krishnamoorthi (“Quantizing deep convolutional networks for efficient inference: A whitepaper”, 2018)
Regarding claim 7, Miret in view of Lucas and further in view of Jang teaches:
The method of claim 1.
However, Miret in view of Lucas and further in view of Jang do not specifically disclose:
wherein a convolution operation in the first NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
f
p
=
f
e
a
t
u
r
e
_
i
n
f
p
⊗
w
e
i
g
h
t
f
p
where feature_outfp represents an output feature map parameter in a form of floating-point, feature_infp represents an input feature map parameter in a form of floating-point and weightfp represents a weight parameter in a form of floating-point.
Krishnamoorthi teaches:
wherein a convolution operation in the first NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
f
p
=
f
e
a
t
u
r
e
_
i
n
f
p
⊗
w
e
i
g
h
t
f
p
where feature_outfp represents an output feature map parameter in a form of floating-point, feature_infp represents an input feature map parameter in a form of floating-point and weightfp represents a weight parameter in a form of floating-point. ([Krishnamoorthi, page 20, Step 2] The first NN model is the NN model before the conversion and training process. First, the examiner notes that the equation is an obvious result of a convolution operation in a neural network. Krishnamoorthi is introduced to show the convolution operation before and after the convolution operation. The output feature y before the training is
y
=
c
o
n
v
(
Q
w
c
o
r
r
e
c
t
e
d
,
x
)
)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Miret, Lucas, Jang and Krishnamoorthi to use the quantization aware training, which models quantization during the training of Krishnamoorthi to implement the neural network conversion method of Miret. The suggestion and/or motivation to do so is to improve the accuracy of the quantized neural network. The quantization aware training provides significant improvement over conventional post training quantization schemes [Krishnamoorthi, page 13, 3.2 Quantization Aware Training, lines 1-8].
Regarding claim 17, Miret in view of Lucas in view of Jang and further in view of Krishnamoorthi teaches:
The method of claim 1, wherein a convolution operation in a third NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
i
n
t
=
f
e
a
t
u
r
e
_
i
n
i
n
t
⊗
w
e
i
g
h
t
i
n
t
where feature_outint denotes an output feature map parameter in a form of integer, feature_inint denotes an input feature map parameter in a form of integer, and weightint denotes a weight parameter in a form of integer. ([Krishnamoorthi, page 20, Step 3] The third NN model is the result of the calibration process. First, the examiner notes that the equation is an obvious result of a convolution operation in a neural network. Krishnamoorthi is introduced to show the convolution operation before and after the convolution operation. The output feature y after the sufficient training is
y
=
c
o
n
v
(
Q
w
c
o
r
r
e
c
t
e
d
,
x
)
)
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Miret in view of Lucas in view of Jang and further in view of Bhalgat et al. (“LSQ+: Improving low-bit quantization through learnable offsets and better initialization”, 2020, hereinafter ‘Bhalgat’).
Regarding claim 8,
The method of claim 1.
However, Miret in view of Lucas and further in view of Jang do not specifically disclose:
wherein a convolution operation in the second NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
f
p
=
[
f
e
a
t
u
r
e
i
n
f
p
-
o
f
s
f
]
×
s
f
+
o
f
⊗
[
w
e
i
g
h
t
f
p
s
w
]
×
s
w
where feature_infp represents an input feature map parameter in a form of floating-point, weightfp represents a weight parameter in a form of floating-point, of represents the offset value for an input feature map, sf represents the scale value for the input feature map, sw represents the scale value for a weight, and ⌊ ⌋ represents round and clip operations.
Bhalgat teaches:
wherein a convolution operation in the second NN model is expressed as:
f
e
a
t
u
r
e
_
o
u
t
f
p
=
[
f
e
a
t
u
r
e
i
n
f
p
-
o
f
s
f
]
×
s
f
+
o
f
⊗
[
w
e
i
g
h
t
f
p
s
w
]
×
s
w
where feature_infp represents an input feature map parameter in a form of floating-point, weightfp represents a weight parameter in a form of floating-point, of represents the offset value for an input feature map, sf represents the scale value for the input feature map, sw represents the scale value for a weight, and ⌊ ⌋ represents round and clip operations. ([page 2980, left col, lines 1-22]
β
denotes the offset, s denotes the scale, x^ denotes the output feature. The output feature x^ is calculated based on
x
×
s
+
β
and x is calculated based on
x
~
=
c
l
a
m
p
(
x
-
β
s
,
n
,
p
)
. The final feature output w^x^ is calculated based on the equation (5) which is
w
~
×
s
w
wherein
s
w
denotes the scale. The activation is quantized using asymmetric activation quantization, and the weight quantization is performed using symmetric signed quantization)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Miret, Lucas, Jang and Bhalgat to use the method of using symmetric signed quantization for weight quantization and asymmetric quantization for activation quantization of Bhalgat to implement the neural network conversion method of Miret. The suggestion and/or motivation to do so is to improve the efficiency of the quantized neural network by reducing additional cost during inference caused by using the additional offset term [Bhalgat, page 2980, left col, lines 1-22].
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Miret in view of Lucas in view of Jang and further in view of Eisenman et al. (“Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation Models”, 2022, hereinafter ‘Eisenman’).
Regarding claim 15,
Miret in view of Lucas teaches:
The method of claim 1, further comprising: generating a third NN model.
However, Miret in view of Lucas in view of Jang do not specifically disclose:
generating a third NN model based on the second NN model by removing the at least one marker of the second NN model.
Eisenman teaches:
generating a third NN model based on the second NN model by removing the at least one marker of the second NN model. ([Eisenman, page 934, right col, last para, line 1 – page 935, left col, line 15] At the end of the training stage, the controller removes an older checkpoint)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Miret, Lucas, Jang and Eisenman to use the method of generating a new neural network by removing checkpoints (markers) of Eisenman to implement the neural network conversion method of Miret. The suggestion and/or motivation to do so is to reduce the waste of storage space by updating the existing checkpoints information to new information related to the training of the neural network.
Claims 16 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Miret in view of Lucas in view of Jang and further in view of Loncar et al. (“QONNX: Representing Arbitrary-Precision Quantized Neural Networks”, 2022, hereinafter ‘Loncar’)
Regarding claim 16, Miret in view of Lucas teaches:
The method of claim 1, further comprising: generating, based on the second NN model, a third NN model comprising a weight parameter and an input feature map parameter ([Miret, page 1059, right col, lines 2-29] discloses quantizing the input weight (weight parameter) and activation (input feature map) tensors of 32-bit floating point precision by scaling the saturated tensor by s, adds the offset z and finally rounds the resulting tensor elements to the nearest integer. As a result, a configuration B consisting of weights and activations (i.e., third neural network) for each layer is determined)
However, Miret in view of Jang in view of Lucas do not specifically disclose:
generating, based on the second NN model, a third NN model comprising a weight parameter and an input feature map parameter in a form of integer having a length of one of 2-bits to 16-bits.
Loncar teaches:
generating, based on the second NN model, a third NN model comprising a weight parameter and an input feature map parameter in a form of integer having a length of one of 2-bits to 16-bits. ([Loncar, page 5, left col, lines 1-5] and [page 5, TABLE II, top box] Inputs are x (float32) tensor to be quantized, and the output is 8 bits tensor which is between 2 bits and 16 bits. [page 2, left col, lines 1-16] discloses that the quantization process performs neural network quantization to generate a quantized NN)
Before the effective filing date of the invention to a person of ordinary skill in the art, it would have been obvious, having the teachings of Miret, Lucas, Jang and Loncar to use the method of generating a new neural network in a form of integer having a length of one of 2-bits to 16-bits of Loncar to implement the neural network conversion method of Miret. The suggestion and/or motivation to do so is to improve the efficiency of the neural network by reducing the size of the neural network.
Regarding claim 18, Miret in view of Lucas in view of Jang and further in view of Loncar teaches:
The method of claim 1, further comprising: generating a third NN model based on the second NN model in an open neural network exchange (ONNX) format, wherein constant parameters of the third NN model that can be pre-calculated are stored as pre-calculated constant parameters. ([Loncar, page 5, left col, lines 1-5] and [page 5, TABLE II, top box] Inputs are x (float32) tensor to be quantized, and the output is 8 bits tensor which is between 2 bits and 16 bits. [page 2, left col, lines 1-16] discloses that the quantization process performs neural network quantization to generate a quantized NN. [page 3, right col, last para, lines 1-5] shows that the QONNX is an extension to the ONNX. [page 7, left col, 2nd para] shows that the quantization applied to the constants and quantization applied to the data flow are differentiated and the scale and offset are applied before the quantization (pre-calculated))
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUN KWON whose telephone number is (571)272-2072. The examiner can normally be reached Monday – Friday 8:00AM – 5:00PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Kawsar can be reached at (571)270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JUN KWON/Examiner, Art Unit 2127
/ABDULLAH AL KAWSAR/Supervisory Patent Examiner, Art Unit 2127