DETAILED ACTION
This action is in response to claims filed 16 March 2026 for application 18385871 filed 31 October 2023. Currently claims 30-49 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
The indicated allowability of claims 30-49 is withdrawn in view of the newly discovered reference to Mellempudi et al. (US 20190354846 A1). Rejections based on the newly cited reference follow.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 30-34, 38-40, 44-45 and 48 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Mellempudi et al. (US 20190354846 A1).
Regarding claims 30, 38, and 44, Mellempudi discloses:
A processor comprising:
One or more circuits to:
scale one or more neural network loss values to cause one or more gradients to be different from zero (see figure 6A, and paragraph 140, which describes that “FIG. 6A illustrates a scaling operation 600 to avoid information loss during 16-bit training, according to an embodiment. Embodiments described herein provide logic that creates a scaled FP16 tensor 602 for use during neural network training. The typical distribution of deep neural network (DNN) gradient tensors 606 lies within the FP32 dynamic range 607. However, converting the gradient tensors to FP16 may result in significant information loss 604, particularly for gradients very close to zero. In one embodiment a scaled FP16 tensor 602 may be created by scaling tensor data from the typical distribution of deep neural network (DNN) gradient tensors 606 into the FP16 dynamic range 605. After the scaling operation, the scaled FP16 tensor 602 has a scaled FP16 distribution 603 that allows training computations to be performed using 16-bit floating point values without significant information loss 604”; see steps 702, 703, at figure 6B, and paragraph 146, which describes that “FIG. 7A illustrates a flow diagram of logic 700 to perform scaled compute operations on a layer of a deep neural network, according to an embodiment. In one embodiment the logic 700 can analyze the dynamic range of input tensors at block 702 to determine whether to scale the input tensors, as shown at block 703”; also see paragraph 300, which describes “[t]he operations can additionally include performing, via the mixed precision tensor processor, tensor computations associated with a layer of a neural network to generate loss data, the loss data stored in a floating-point format and scaling the loss data generated via the tensor computations by a scaling factor to enable a data distribution of a gradient tensor generated based on the loss data to be represented by a scaled gradient tensor, the scaled gradient tensor stored as a 16-bit floating point data type”);
use the one or more gradients to perform a first adjustment of one or more weights of one or more neural networks (see steps 704, 706, 708, at figure 6B, and paragraph 146, which describes “[i]f the tensors are to be scaled, the logic 700 can compute an exponent bias using the ABS_MAX value of the input tensor and the input dynamic range, as shown at block 704. The logic 700 can then convert the weights and activations of the layer to scaled FP16 tensors at block 706 before performing the compute operations at block 708”); and
perform a second adjustment of the adjusted one or more weights to compensate for scaling the one or more neural network loss values (see step 710, at figure 6B, and paragraph 146, which describes that “[o]nce the compute operations are complete, the logic 700 can re-scale the weights and activations using the computed exponent bias and down-convert to a scaled FP16 values at block 710 before processing the next layer at block 712”).
Regarding claim 31, Mellempudi discloses: The processor of claim 30, wherein the one or more neural network loss values are scaled during a forward pass of a training iteration (“In one embodiment the exponent bias is dynamically adjusted based on the dynamic-range requirements at each neural network layer as the training progresses. The core computation can be performed in the scaled domain and the results can be scaled back using exponent bias before those results are passed to the next layer of the neural network or scaled to half-precision.” [0193]).
Regarding claim 32, Mellempudi discloses: The processor of claim 30, wherein the scaled one or more neural network loss values cause the one or more gradients to be scaled during a backward propagation of a training iteration (“In a further embodiment the mixed precision tensor processor is to perform tensor computations associated with a layer of a neural network to generate the loss data. The tensor computations can include a backward propagation operation associated with the layer of the neural network. The multiprocessor can dynamically adjust the scaling factor based on values within the gradient tensor. The scaling operations for the loss data can include performing operations to increase a minimum value of the gradient tensor to above the minimum value representable by the a 16-bit floating point data type and to select a scaling factor such that the maximum value of the gradient tensor is below the maximum value representable by the a 16-bit floating point data type. In one embodiment the graphics processor additionally includes a 3D memory stack coupled to the memory controller. The 3D memory stack can include high-bandwidth memory (e.g., HBM, HBM2, etc.).” [0299]).
Regarding claims 33 and 39, Mellempudi discloses:
The processor of claim 30, wherein the one or more neural network loss values are scaled based, at least in part, on a scaling factor that causes the one or more gradients to be scaled upward (see paragraph 141, which describes “[i]n one embodiment, scaled half-precision conversion is performed by multiplying the input floating point tensor with 2.sup.BIAS, where BIAS is a per-tensor scale factor that is shared among values within the tensor. The scaling can be configured to push the values into a higher magnitude, utilizing the previously unused range of the half-precision format”).
Regarding claims 34 and 40, Mellempudi discloses: The processor of claim 30, wherein the one or more neural network loss values are scaled by multiplying the one or more neural network loss values by a scaling factor. (see paragraph 141, which describes “[i]n one embodiment, scaled half-precision conversion is performed by multiplying the input floating point tensor with 2.sup.BIAS, where BIAS is a per-tensor scale factor that is shared among values within the tensor. The scaling can be configured to push the values into a higher magnitude, utilizing the previously unused range of the half-precision format”).
Regarding claim 45, Mellempudi discloses: The system of claim 44, wherein the one or more neural network loss values are scaled based, at least in part, on a scaling factor, and wherein the scaling factor is a constant selected by a user. (“The scaling operations for the loss data can include performing operations to increase a minimum value of the gradient tensor to above the minimum value representable by the a 16-bit floating point data type and to select a scaling factor such that the maximum value of the gradient tensor is below the maximum value representable by the a 16-bit floating point data type. In one embodiment the graphics processor additionally includes a 3D memory stack coupled to the memory controller. The 3D memory stack can include high-bandwidth memory (e.g., HBM, HBM2, etc.).” [0299])
Regarding claim 48, Mellempudi discloses: The system of claim 44, wherein the one or more neural network loss values are scaled during forward propagation of a training iteration to cause the one or more gradients to be scaled during backward propagation of the training iteration ([0193], [0299]).
.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ERIC NILSSON whose telephone number is (571)272-5246. The examiner can normally be reached M-F: 7-3.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, James Trujillo can be reached at (571)-272-3677. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ERIC NILSSON/Primary Examiner, Art Unit 2151