DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In response to Applicant’s claims filed on January 21, 2026 claims 1-20 are now pending for examination in the application.
Response to Arguments
This office action is in response to amendment filed 01/21/2026. In this action claim(s) 1, 6-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sather et al. (US Pub. No. 20210034982) in view of KIM (US pub. No. 20220137866). The Sather et al. reference has been added to address the amendment of encode the difference between the asymptotic value and the feature map by a quantization-based method wherein in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input.
Applicant’s arguments:
On Page 12, applicant argues “The Applicant submits that amended independent claim 1 is inextricably tied to the machine (Ex parte Steiner). For instance, amended independent claim 1 recites, for example, "a central processing unit (CPU) configured to: derive a difference between an asymptotic value of an activation function and a feature map, wherein the feature map is a processing result of a computational layer, the computational layer is subject to processing of a neural network, and the activation function is a function of the computational layer; and encode the difference between the asymptotic value and the feature map by a quantization-based method, wherein in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input." The Applicant submits that at least the above-recited features of amended independent claim 1 are essentially tied to a machine and do not represent mental process or mathematical concepts for mathematical relationships.”
Examiner’s Reply:
The examiner respectfully disagrees and would like to point out that human mind using a computer as a tool is fully capable of deriving values of functions which is a mathematical concept. A human would be able to follow these steps to calculate the values along with any needed additional elements (eg information processing).
Applicant’s arguments:
On Pages 15-16, applicant argues “Accordingly, the present claim integrates the alleged judicial exception into a practical application that reduces memory capacity requirements for storing the feature map by encoding differences between feature maps and asymptotic values of activation functions using a specific quantization-based method and decoding the encoded data to restore the feature maps, thereby reducing computational load, processing time, and complexity of managing data related to the feature map as compared to conventional technology.”
Examiner’s Reply:
Applicant argues that the amended claims comprises statutory subject matter. Examiner
respectfully disagrees. The examiner notes that the computer (being used as a generic tool) as recited in the claims is being used for managing memory storage. The use of mathematical functions does not improve the functioning of a computer. Therefore, the abstract idea recited in the claims is generally linking it to a computer environment, and does not integrate the abstract idea into a practical application.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1-20 is/are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The judicial exception is not integrated into a practical application. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. The eligibility analysis in support of these findings is provided below, in accordance with the 2019 Revised Patent Subject Matter Eligibility Guidance, hereinafter 2019 PEG.
Step 1. In accordance with Step 1 of the eligibility inquiry (as explained in MPEP 2106), it is noted that the system and methods of claims 1-20 are directed to one of the eligible categories of subject matter and therefore satisfy Step 1.
Step 2A. In accordance with Step 2A, prong one of the 2019 PEG, it is noted that the independent claims recite an abstract idea falling within the Mental Process & Mathematical Concepts enumerated groupings of abstract ideas set forth in the 2019 PEG. Examiner is of the position that independent claims 1, 10, 11, and 20 are directed towards the Mathematical Concepts and Mental Process Grouping of Abstract Ideas.
Independent claim 1 recites the following limitations directed towards a Mental Process & Mathematical Concepts:
derive a difference between an asymptotic value of an activation function and a feature map, wherein the feature map is a processing result of a computational layer, the computation layer is subject to processing of a neural network and the activation function is a function of the computational layer (The limitation recites a mathematical concept; deriving); and
encode the difference (The limitation recites a mental process of observation and/or evaluation capable of being performed by the human mind by using computer as a tool to encode a difference)
between the asymptotic value and the feature map by a quantization- based method wherein the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input (The limitation recites a mathematical concept; quantizing data).
Step 2A. In accordance with Step 2A, prong two of the 2019 PEG, the judicial exception is not integrated into a practical application because of the recitation in claim(s) 1:
a computational unit (i.e., as a generic processor performing a generic computer function)
Step 2B. Similar to the analysis under 2A Prong Two, the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Because the additional elements of the independent claims amount to insignificant extra solution activity and/or mere instructions, the additional elements do not add significantly more to the judicial exception such that the independent claims as a whole would be patent eligible.
Independent claim 10 recites the following limitations directed towards a Mental Process & Mathematical Concepts:
deriving a difference between an asymptotic value of an activation function and a feature map, wherein the feature map is a processing result of a computational layer, the computational layer is subject (The limitation recites a mathematical concept; deriving)
to processing of a neural network and the activation function, is a function of the computational layer (The limitation recites a mental process of observation and/or evaluation capable of being performed by the human mind by using computer as a tool to processing a neural network); and
encoding the difference between the asymptotic value and the feature map by a quantization-based method wherein in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input (The limitation recites a mental process of observation and/or evaluation capable of being performed by the human mind by using computer as a tool to encode a difference).
Step 2A. In accordance with Step 2A, prong two of the 2019 PEG, the judicial exception is not integrated into a practical application.
Step 2B. Similar to the analysis under 2A Prong Two, the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Because the additional elements of the independent claims amount to insignificant extra solution activity and/or mere instructions, the additional elements do not add significantly more to the judicial exception such that the independent claims as a whole would be patent eligible.
Independent claim(s) 11 and 20 recite(s) the following limitations directed towards a Mental Process & Mathematical Concepts:
decode encoded data (The limitation recites a mental process of observation and/or evaluation capable of being performed by the human mind by using computer as a tool to decode encoded data)
to generate a difference between an asymptotic value of an activation function and a first feature map¸ wherein the first feature map is a processing result of a first computational layer subject (The limitation recites a mathematical concept; calculating a difference)
to processing of a neural network¸ and the activation function is a function of the first computational layer (The limitation recites a mental process of observation and/or evaluation capable of being performed by the human mind by using computer as a tool to process a neural network);
derive the first feature map, based on the difference and the asymptotic value (The limitation recites a mathematical concept; deriving).
Step 2A. In accordance with Step 2A, prong two of the 2019 PEG, the judicial exception is not integrated into a practical application because of the recitation in claim(s) 11 and 20:
a first computational unit
Step 2B. Similar to the analysis under 2A Prong Two, the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Because the additional elements of the independent claims amount to insignificant extra solution activity and/or mere instructions, the additional elements do not add significantly more to the judicial exception such that the independent claims as a whole would be patent eligible.
Step 2A. In accordance with Step 2A, prong two of the 2019 PEG, the judicial exception is not integrated into a practical application.
Step 2B. Similar to the analysis under 2A Prong Two, the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Because the additional elements of the independent claims amount to insignificant extra solution activity and/or mere instructions, the additional elements do not add significantly more to the judicial exception such that the independent claims as a whole would be patent eligible.
Therefore, independent claims 1, 10, 11, and 20 are rejected under 35 U.S.C. 101.
With respect to claim(s) 2 and 17:
the CPU is further configured to derive the difference between the feature map and an asymptotic lower bound of the activation function, and the asymptotic value of the activation function is the asymptotic lower bound of the activation function (The limitation recites a mathematical concept; deriving).
Step 2A Prong Two Analysis:
This judicial exception is not integrated into a practical application because the claim as drafted recites insignificant extrasolution activity.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 3:
Step 2A Prong One Analysis:
the CPU is further configured to quantize the difference by a first method, and the first method sets the quantization level for the zero input to zero (The limitation recites a mathematical concept; quantizing).
Step 2A Prong Two Analysis:
This judicial exception is not integrated into a practical application because the claim as drafted recites insignificant extrasolution activity.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 4:
Step 2A Prong One Analysis:
the CPU is further configured to quantize the difference by a second method, and the second method sets a rounded value input is set as a quantization level for the input (The limitation recites a mathematical concept; quantizing).
Step 2A Prong Two Analysis:
This judicial exception is not integrated into a practical application because the claim as drafted recites insignificant extrasolution activity.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 5:
Step 2A Prong One Analysis:
derive the difference between the feature map and an asymptotic upper bound of the activation function, and the asymptotic value of the activation function is the asymptotic upper bound of the activation function (The limitation recites a mathematical concept; deriving).
Step 2A Prong Two Analysis:
the CPU is furthered configured to
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 6:
Step 2A Prong One Analysis:
Examiner is of the position the dependent claim is directed toward additional elements.
Step 2A Prong Two Analysis:
The CPU (i.e., as a generic processor performing a generic computer function) is further configured to control at least asymptotic value for at least one computational layer of the neural network (recites insignificant extra solution activity that amounts to controlling data), the at least one computational layer includes the computational layer, and
The at least one asymptotic value includes the asymptotic value.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 7:
Step 2A Prong One Analysis:
Examiner is of the position the dependent claim is directed toward additional elements.
Step 2A Prong Two Analysis:
The CPU (i.e., as a generic processor performing a generic computer function) is further configured to determine whether to encode the difference for each of at least one computational layer of the neural network, and
The at least one computational layer includes the computational layer (recites insignificant extra solution activity that amounts to controlling data).
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 8 and 18:
Step 2A Prong One Analysis:
Wherein the CPU is further configured to execute computation of the computational layer to generate the feature map (The limitation recites a mental process of observation and/or evaluation capable of being performed by the human mind by using computer as a tool to generate a map).
Step 2A Prong Two Analysis:
This judicial exception is not integrated into a practical application because the claim as drafted recites insignificant extrasolution activity.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 9:
Step 2A Prong One Analysis:
Examiner is of the position the dependent claim is directed toward additional elements.
Step 2A Prong Two Analysis:
a storage unit configured to store the encoded difference (recites insignificant extra solution activity that amounts to storing data).
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 10:
Step 2A Prong One Analysis:
deriving a difference between an asymptotic value of an activation function and a feature map¸ wherein the feature map is a processing result of a computational layer, the computational layer is subject to processing of a neural network, and the activation function is a function of the computational layer (The limitation recites a mathematical concept; deriving); and
encoding the difference between the asymptotic value and the feature map by a quantization-based method, wherein in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input (The limitation recites a mental process of observation and/or evaluation capable of being performed by the human mind by using computer as a tool to encode a difference).
Step 2A Prong Two Analysis:
This judicial exception is not integrated into a practical application because the claim as drafted recites insignificant extrasolution activity.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 12:
Step 2A Prong One Analysis:
Add an asymptotic lower bound of the activation function to the difference, wherein the asymptotic value of the activaqtion function is the asymptotic lower bound of the activation function (The limitation recites a mathematical concept; adding); and
derive the feature map by adding an asymptotic lower bound of the activation function to the difference (The limitation recites a mathematical concept; deriving).
Step 2A Prong Two Analysis:
This judicial exception is not integrated into a practical application because the claim as drafted recites insignificant extrasolution activity.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 13:
Step 2A Prong One Analysis:
subtract the difference from an asymptotic upper bound of the activation function, the asymptotic value of the activation function is the asymptotic upper bound of the activation function (The limitation recites a mathematical concept; subtracting); and
derive the first feature map based on the subtraction (The limitation recites a mathematical concept; deriving).
Step 2A Prong Two Analysis:
the first computational unit
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 14:
Step 2A Prong One Analysis:
Examiner is of the position the dependent claim is directed toward additional elements.
Step 2A Prong Two Analysis:
The CPU (i.e., as a generic processor performing a generic computer function) control at least one asymptotic value for at least one computational layer of the neural network (recites insignificant extra solution activity that amounts to controlling data),
the at least one computational layer includes the first computational layer, and
the at least one asymptotic value includes the asymptotic value.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 15:
Step 2A Prong One Analysis:
Examiner is of the position the dependent claim is directed toward additional elements.
Step 2A Prong Two Analysis:
The CPU (i.e., as a generic processor performing a generic computer function) is further configured to determine whether to decode the encoded data for each of at least one computational layer of the neural network (recites insignificant extra solution activity that amounts to controlling data), and the at least one computational layer includes the first computational layer.
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
With respect to claim(s) 16:
Step 2A Prong One Analysis:
Derive the difference between the first feature map and the asymptotic value (The limitation recites a mathematical concept; deriving);
Encode the difference to generate the encoded data, wherein the difference is encoded by a quantization-based method , and in the quantization-based method a midpoint of a quantization step size is not set as a quantization level for zero input (The limitation recites a mental process of observation and/or evaluation capable of being performed by the human mind by using computer as a tool to generate encoded data).
Step 2A Prong Two Analysis:
The CPU (i.e., as a generic processor performing a generic computer function).
Step 2B Analysis:
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim is not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 6-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sather et al. (US Pub. No. 20210034982) in view of KIM (US pub. No. 20220137866).
With respect to claim 1, Sather et al. teaches an information processing device comprising:
a central processing unit (CPU) (See Figure 7) configured to:
derive a difference between an asymptotic value of an activation function and a feature map, wherein the feature map is a processing result of a computational layer, the computational layer is subject to processing of a neural network and the activation function is a function of the
computational layer (Paragraph 34 discloses quality is maintained when activation quantization is introduced (without retraining network parameters). For example, the loss value should not increase significantly. Additionally, scaling and shifting values are initialized, in some embodiments, such that the information contained in feature maps produced throughout the network is maximally preserved by the value quantization). Sather et al. does not disclose encode the difference between the asymptotic value and the feature map by a quantization-based method wherein in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input.
However, KIM discloses encode the difference between the asymptotic value and the feature map by a quantization-based method wherein in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input (Paragraph 221 discloses artificial neural network model may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and an encoder-decoder structure, a deep neural network (DNN) such as SegNet, DeconvNet, DeepLAB V3+, or U-net, or SqueezeNet, Alexnet, ResNet18, MobileNet-v2, GoogLeNet, Resnet-v2, Resnet50, Resnet101, and Inception-v3 and Paragraph 157 discloses a quantization algorithm of the kernel, the feature map, or the like of the compiled artificial neural network model).
Therefore, it would have been obvious before the effective filing data of invention was made to a person having ordinary skill in the art to modify Sather with KIM et al. This would have facilitated training a deep neural network.
The Sather reference as modified by KIM teaches all the limitations of claim 1. With respect to claim 6, Sather discloses the information processing device according to claim 1,
wherein the CPU is further configured to control at least one asymptotic value for at least one
computational layer of the neural network, the at least one computational layer includes the computational layer, and the at least one asymptotic value includes the asymptotic value (Paragraph 19 discloses determining a distribution of values for each input of a layer of the neural network and selecting a set of scaling and shift values for each layer based on the determined distribution of values).
The Sather reference as modified by KIM teaches all the limitations of claim 1. With respect to claim 7, KIM discloses the information processing device according to claim 1,
wherein the CPU is further configured to determine whether to encode the difference for each of at least one computational layer of the neural network, and the at least one computational layer includes the computational layer (Paragraph 221 discloses artificial neural network model may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and an encoder-decoder structure, a deep neural network (DNN) such as SegNet, DeconvNet, DeepLAB V3+, or U-net, or SqueezeNet, Alexnet, ResNet18, MobileNet-v2, GoogLeNet, Resnet-v2, Resnet50, Resnet101, and Inception-v3 and Paragraph 157 discloses a quantization algorithm of the kernel, the feature map, or the like of the compiled artificial neural network model). The motivation to combine statement previously provided in the rejection of independent claim 1 provided above, combining the Sather et al. reference and the KIM reference is applicable to dependent claim 7.
The Sather reference as modified by KIM teaches all the limitations of claim 1. With respect to claim 8, KIM discloses the information processing device according to claim 1,
wherein the CPU is further configured to execute computation of the computational layer
to generate the feature map (Paragraph 923 discloses due to the characteristics of the artificial neural network model, when the input feature map and the kernel are convolved, an output feature map is generated, and the corresponding output feature map becomes the input feature map of the next layer). The motivation to combine statement previously provided in the rejection of independent claim 1 provided above, combining the Sather et al. reference and the KIM reference is applicable to dependent claim 8.
The Sather reference as modified by KIM teaches all the limitations of claim 1. With respect to claim 9, KIM discloses the information processing device according to claim 1,
further comprising a storage unit configured to store the encoded difference (Paragraph 462 discloses the artificial neural network memory controller may be configured to set the storage area of the memory).
10.
(Currently Amended) An information processing method comprising:
deriving a difference between an asymptotic value of an activation function and a feature map, wherein the feature map is a processing result of a computational layer, the computational layer is subject to processing of a neural network and the activation function is a function of the
computational layer (Paragraph 34 discloses quality is maintained when activation quantization is introduced (without retraining network parameters). For example, the loss value should not increase significantly. Additionally, scaling and shifting values are initialized, in some embodiments, such that the information contained in feature maps produced throughout the network is maximally preserved by the value quantization). Sather et al. does not disclose encode the difference between the asymptotic value and the feature map by a quantization-based method wherein in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input.
However, KIM discloses encoding the difference between the asymptotic value and the feature map by a quantization-based method wherein in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input (Paragraph 221 discloses artificial neural network model may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and an encoder-decoder structure, a deep neural network (DNN) such as SegNet, DeconvNet, DeepLAB V3+, or U-net, or SqueezeNet, Alexnet, ResNet18, MobileNet-v2, GoogLeNet, Resnet-v2, Resnet50, Resnet101, and Inception-v3 and Paragraph 157 discloses a quantization algorithm of the kernel, the feature map, or the like of the compiled artificial neural network model).
Therefore, it would have been obvious before the effective filing data of invention was made to a person having ordinary skill in the art to modify Sather with KIM et al. This would have facilitated training a deep neural network.
Claim(s) 11, 14-16, and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over KIM (US pub. No. 20220137866) in view of Sather et al. (US Pub. No. 20210034982).
With respect to claim 11, KIM discloses an information processing device comprising:
a central processing unit (CPU) (See Figure 7) configured to:
decode encoded data to generate a difference between an asymptotic value of an activation function and a first feature map, wherein the first feature map is a processing result of a first
computational layer, the first computation layer is subject to processing of a neural
network, and the activation function is a function of the first computational layer (Paragraph 221 discloses artificial neural network model may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and an encoder-decoder structure, a deep neural network (DNN) such as SegNet, DeconvNet, DeepLAB V3+, or U-net, or SqueezeNet, Alexnet, ResNet18, MobileNet-v2, GoogLeNet, Resnet-v2, Resnet50, Resnet101, and Inception-v3) that decodes encoded data to generate a difference between a feature map that is a processing result of a computational layer subject to processing of a neural network and an asymptotic value of an activation function of the computational layer subject to processing (Paragraph 251 discloses each layer is processed in one data access request unit. If the data size such as the weight value, the feature map, the kernel, the activation map, and the like of the artificial neural network model is larger than the available capacity of the cache memory of the processor, the corresponding data access request may be divided into a plurality of data access requests and in this case, the artificial neural network data locality of the artificial neural network model may be reconstructed). KIM does not disclose derive the first feature map, based on the difference and the asymptotic value.
However, Sather et al. discloses derive the first feature map, based on the difference and the asymptotic value (Paragraph 34 discloses quality is maintained when activation quantization is introduced (without retraining network parameters). For example, the loss value should not increase significantly. Additionally, scaling and shifting values are initialized, in some embodiments, such that the information contained in feature maps produced throughout the network is maximally preserved by the value quantization).
Therefore, it would have been obvious before the effective filing data of invention was made to a person having ordinary skill in the art to modify KIM et al. with Sather et al. This would have facilitated training a deep neural network.
With respect to claim 14, it is rejected on grounds corresponding to above rejected claim 6, because claim 14 is substantially equivalent to claim 6.
The KIM reference as modified by Sather et al. teaches all the limitations of claim 11. With respect to claim 15, KIM discloses the information processing device according to claim 11,
wherein the CPU is further configured to determine whether to decode the encoded data for each of at least one computational layer of the neural network, and the at least one computational layer includes the first computational layer (Paragraph 221 discloses artificial neural network memory controller).
The KIM reference as modified by Sather et al. teaches all the limitations of claim 11. With respect to claim 16, KIM discloses the information processing device according to claim 11,
wherein the CPU is further configured to:
derive the difference between the first feature map and the asymptotic value (Paragraph 251 discloses each layer is processed in one data access request unit. If the data size such as the weight value, the feature map, the kernel, the activation map, and the like of the artificial neural network model is larger than the available capacity of the cache memory of the processor, the corresponding data access request may be divided into a plurality of data access requests and in this case, the artificial neural network data locality of the artificial neural network model may be reconstructed); and
encode the difference to generate the encoded data, wherein the difference is encoded by a quantization-based method, and in the quantization-based method, a midpoint of a quantization step size is not set as a quantization level for zero input (Paragraph 221 discloses artificial neural network model may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and an encoder-decoder structure and Paragraph 157 discloses a quantization algorithm of the kernel, the feature map, or the like of the compiled artificial neural network model).
With respect to claim 18, it is rejected on grounds corresponding to above rejected claim 8, because claim 18 is substantially equivalent to claim 8.
The KIM reference as modified by Sather et al. teaches all the limitations of claim 11. With respect to claim 19, KIM discloses the information processing device according to claim 11,
further comprisinga storage unit configured to store the encoded data, wherein the CPU is further configured to decode the stored encoded data (Paragraph 462 discloses the artificial neural network memory controller may be configured to set the storage area of the memory).
With respect to claim 20, KIM discloses an information processing method comprising:
decoding encoded data to generate a difference between an asymptotic value of
an activation function and a feature map, wherein
the feature map is a processing result of a computational layer, the computational layer is subject to processing of a neural network, and the activation function is a function of the
computational layer (Paragraph 221 discloses artificial neural network model may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and an encoder-decoder structure, a deep neural network (DNN) such as SegNet, DeconvNet, DeepLAB V3+, or U-net, or SqueezeNet, Alexnet, ResNet18, MobileNet-v2, GoogLeNet, Resnet-v2, Resnet50, Resnet101, and Inception-v3) that decodes encoded data to generate a difference between a feature map that is a processing result of a computational layer subject to processing of a neural network and an asymptotic value of an activation function of the computational layer subject to processing (Paragraph 251 discloses each layer is processed in one data access request unit. If the data size such as the weight value, the feature map, the kernel, the activation map, and the like of the artificial neural network model is larger than the available capacity of the cache memory of the processor, the corresponding data access request may be divided into a plurality of data access requests and in this case, the artificial neural network data locality of the artificial neural network model may be reconstructed). KIM does not disclose derive the first feature map, based on the difference and the asymptotic value.
However, Sather et al. discloses deriving the feature map, based on the difference and the asymptotic value (Paragraph 34 discloses quality is maintained when activation quantization is introduced (without retraining network parameters). For example, the loss value should not increase significantly. Additionally, scaling and shifting values are initialized, in some embodiments, such that the information contained in feature maps produced throughout the network is maximally preserved by the value quantization).
Therefore, it would have been obvious before the effective filing data of invention was made to a person having ordinary skill in the art to modify KIM et al. with Sather et al. This would have facilitated training a deep neural network.
Claim(s) 2-5, 12-13, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sather et al. (US Pub. No. 20210034982) and KIM (US Pub. No. 20220137866) in further view of DAI et al. (US Pub. No. 20210241112).
The Sather reference as modified by KIM teaches all the limitations of claim 1. With respect to claim 2, Sather as modified KIM does not disclose the CPU is further configured to derive the difference between the feature map and an asymptotic lower bound of the activation function, and the asymptotic value of the activation function is the asymptotic lower bound of the activation function.
However, DAI et al. teaches the CPU is further configured to derive the
difference between the feature map and an asymptotic lower bound of the activation
function, and the asymptotic value of the activation function is the asymptotic lower bound of the activation function (Paragraphs 194-195 discloses to impel the neuron/feature map channel to further encode more information, its corresponding hyperparameter is tuned down to weaken its corresponding channel regularization term in the objective function. In actual operation, for example, the corresponding hyperparameter γ.sub.c.sup.l may be multiplied by a coefficient less than 1 (for example, 0.9). Further, in order to ensure the speed of tuning, a value range for each y.sub.c.sup.l may be set to, for example, the upper and lower bounds of 1×10.sup.−3 and 1×10.sup.−10, respectively).
Therefore, it would have been obvious before the effective filing data of invention was made to a person having ordinary skill in the art to modify Sather et al. and KIM with DAI et al. This would have facilitated training a deep neural network.
The Sather reference as modified by KIM and DAI et al. teaches all the limitations of claim 2. With respect to claim 3, Kim discloses the CPU is further configured to quantize the difference by a first method, and the first method sets the quantization level for the zero input to zero (Paragraph 221 discloses artificial neural network model may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and an encoder-decoder structure, a deep neural network (DNN) and Paragraph 951 discloses the compiler may compile an artificial neural network model with optimization algorithms (e.g., Quantization, Pruning, Retraining, Layer fusion, Model Compression, Transfer Learning, AI Based Model Optimization, and another Model Optimizations)). The motivation to combine statement previously provided in the rejection of independent claim 2 provided above, combining the Sather et al. reference and the KIM reference is applicable to dependent claim 3.
The Sather reference as modified by KIM and DAI et al. teaches all the limitations of claim 2. With respect to claim 4, Kim discloses the CPU is further configured to quantize the difference by a second method, and the second method sets a rounded value of an input as a quantization level for the input (Paragraph 221 discloses artificial neural network model may be a model such as a fully convolutional network (FCN) having VGG, VGG16, DenseNET, and an encoder-decoder structure, a deep neural network (DNN) and Paragraph 951 discloses the compiler may compile an artificial neural network model with optimization algorithms (e.g., Quantization, Pruning, Retraining, Layer fusion, Model Compression, Transfer Learning, AI Based Model Optimization, and another Model Optimizations) and Paragraph 951 discloses the compiler may compile an artificial neural network model with optimization algorithms (e.g., Quantization, Pruning, Retraining, Layer fusion, Model Compression, Transfer Learning, AI Based Model Optimization, and another Model Optimizations)). The motivation to combine statement previously provided in the rejection of independent claim 2 provided above, combining the Sather et al. reference and the KIM reference is applicable to dependent claim 4.
The Sather et al. reference as modified by KIM teaches all the limitations of claim 1. With respect to claim 5, Sather et al. as modified by KIM does not disclose wherein the CPU is further configured to derive the difference between the feature map and an asymptotic upper bound of the activation function, and the asymptotic value of the activation function is the asymptotic upper bound of the activation function.
However, DAI et al. teaches wherein the CPU is further configured to derive the difference between the feature map and an asymptotic upper bound of the activation function, and the asymptotic value of the activation function is the asymptotic upper bound of the activation function (Paragraphs 194-195 discloses to impel the neuron/feature map channel to further encode more information, its corresponding hyperparameter is tuned down to weaken its corresponding channel regularization term in the objective function. In actual operation, for example, the corresponding hyperparameter γ.sub.c.sup.l may be multiplied by a coefficient less than 1 (for example, 0.9). Further, in order to ensure the speed of tuning, a value range for each y.sub.c.sup.l may be set to, for example, the upper and lower bounds of 1×10.sup.−3 and 1×10.sup.−10, respectively).
Therefore, it would have been obvious before the effective filing data of invention was made to a person having ordinary skill in the art to modify Sather et al. and KIM with DAI et al. This would have facilitated training a deep neural network.
The Sather et al. reference as modified by KIM teaches all the limitations of claim 1. With respect to claim 12, Sather et al. as modified by KIM does not disclose add an asymptotic lower bound of the activation function to the difference, wherein the asymptotic value of the activation function is the asymptotic lower bound of the activation function.
However, DAI et al. teaches add an asymptotic lower bound of the activation function to the difference, wherein the asymptotic value of the activation function is the asymptotic lower bound of the activation function (Paragraphs 194-195 discloses to impel the neuron/feature map channel to further encode more information, its corresponding hyperparameter is tuned down to weaken its corresponding channel regularization term in the objective function. In actual operation, for example, the corresponding hyperparameter γ.sub.c.sup.l may be multiplied by a coefficient less than 1 (for example, 0.9). Further, in order to ensure the speed of tuning, a value range for each y.sub.c.sup.l may be set to, for example, the upper and lower bounds of 1×10.sup.−3 and 1×10.sup.−10, respectively); and
derive the first feature map based on the addition (Paragraphs 194-195 discloses to impel the neuron/feature map channel to further encode more information, its corresponding hyperparameter is tuned down to weaken its corresponding channel regularization term in the objective function. In actual operation, for example, the corresponding hyperparameter γ.sub.c.sup.l may be multiplied by a coefficient less than 1 (for example, 0.9). Further, in order to ensure the speed of tuning, a value range for each y.sub.c.sup.l may be set to, for example, the upper and lower bounds of 1×10.sup.−3 and 1×10.sup.−10, respectively).
Therefore, it would have been obvious before the effective filing data of invention was made to a person having ordinary skill in the art to modify Sather et al. and KIM with DAI et al. This would have facilitated training a deep neural network.
With respect to claim 13, it is rejected on grounds corresponding to above rejected claim 2, because claim 13 is substantially equivalent to claim 2.
With respect to claim 17, it is rejected on grounds corresponding to above rejected claim 2, because claim 17 is substantially equivalent to claim 2.
Relevant Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US Pub. No. 20220174328 is directed to High-Fidelity Generative Image Compression: Paragraphs 56-57 discloses he latent representation of the training data item using the hyper-encoder neural network and in accordance with current values of the hyper-encoder network parameters to generate a latent representation of the conditional entropy model, i.e., a “hyper-prior”. In one example, the hyper-encoder neural network is a convolutional neural network and the hyper-prior is a multi-channel feature map output by the final layer of the hyper-encoder neural network. The system quantizes and entropy encodes the hyper-prior. For example, the system can quantize the hyper-prior use a quantizing engine. Quantizing a value refers to mapping the value to a member of a discrete set of possible code symbols. For example, the set of possible code symbols may be integer values, and the system may perform quantization by rounding real-valued numbers to integer values. The system can entropy encode the quantized hyper-prior using, e.g., a pre-determined entropy model defined by one or more predetermined code symbol probability distributions. In one example, the predetermined entropy model may specify a respective predetermined code symbol probability distribution for each code symbol of the quantized hyper-prior. In this example, the system may entropy encode each code symbol of the quantized hyper-prior using the corresponding predetermined code symbol probability distribution. The system can use any appropriate entropy encoding technique, e.g., a Huffman encoding technique, or an arithmetic encoding technique.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS E ALLEN whose telephone number is (571)270-3562. The examiner can normally be reached Monday through Thursday 830-630.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Boris Gorney can be reached at (571) 270-5626. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BORIS GORNEY/Supervisory Patent Examiner, Art Unit 2154
/N.E.A/Examiner, Art Unit 2154