Prosecution Insights
Last updated: October 01, 2026
Application No. 18/680,995

DYNAMIC QUANTIZATION

Non-Final OA §102§103
Filed
May 31, 2024
Examiner
CHOI, YUK TING
Art Unit
Tech Center
Assignee
Advanced Micro Devices Inc.
OA Round
1 (Non-Final)
72%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 72% — above average
72%
Career Allowance Rate
481 granted / 673 resolved
+11.5% vs TC avg
Strong +36% interview lift
Without
With
+36.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
20 currently pending
Career history
698
Total Applications
across all art units

Statute-Specific Performance

§101
17.5%
-22.5% vs TC avg
§103
60.4%
+20.4% vs TC avg
§102
14.6%
-25.4% vs TC avg
§112
5.9%
-34.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 673 resolved cases

Office Action

§102 §103
DETAILED ACTION 1. The present application 18/680,995, filed on 05/31/2024, is being examined under the first inventor to file provisions of the AIA . Clams 1-20 are pending in this application. Drawings 2. The drawings received on 09/01/2025 are accepted by the Examiner. 35 USC § 101 Analysis: 3. Claims 1-20 are directed to a statutory category [e.g., machine]. The machine comprises a computing unit having memory and matrix-multiplier circuitry. Claims 1-20 fall within one of the groupings of abstract ideas [e.g., Mathematical operations] enumerated in the 2019 PEG. Although the claims include mathematical operations, including deriving a per-tile scale and performing quantization. These operations are applied in the claimed compute unit to quantize an input or a resulting matrix generated by the matrix-multiplier circuitry. When considered as a whole, the claims integrate the recited mathematical concept into a practical application involving an improvement to matrix-processing and quantization technology. Accordingly, claims 1-20 are eligible under 35 USC § 101. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-3, 5, 9-11, 13 and 16-19 are rejected under 35 U.S.C. 102 (a)(1) as being anticipated by over Khaliany et al. (US 2022/0067530 A1). Referring to claim 1, Khaliany discloses a compute unit, comprising: memory (See para. [0093] and para. [0094] and Figure 3B, a processor 375 is coupled to memory 360) configured to store a first input that includes subset of data in a tensor, a channel, or a head (See para. [0024], para. [0027], para. [0028] and Figure 1A, storing an input activation tenor 101 subdivided along the input-channel dimension into vectors of parameters 105, each vector representing a subset of data of the input activation tenor); and a matrix multiplier comprising circuitry configured to multiply the first input with a second input to generate a resulting matrix (See para. [0024], para. [0027], para. [0123], weights and/or activations values are quantized for input to a layer for a neural network. Weight-activation products are computed and summed up to produce vector dot-products), wherein the compute unit is configured to: derive a per tile scale (See para. [0024], para. [0061]-para. [0064] and Figure 2A, a scale factor is computed for each vector of parameters, including determining a maximum absolute value of the parameters in the vector and calculating a corresponding scale factor [e.g., Equations 7a-7j], the scale factors are computed at a per-vector granularity, each vector representing a subset of data of the tenor, note Khailany’s vector corresponds to the claimed tile since the vector represents a subset of data of the input activation tenor, Applicant’s specification [0015] and para. [0018] describes a tile as a subset of data in a tensor, channel, or head and further states that the term “tile” does not necessary imply a two-dimensional arrangement), and perform a quantization operation, using the per tile scale, on at least one of the first input or the resulting matrix (See para. [0024], para. [0031], performing a quantization using fine-gained per-vector scale factors to mitigate quantization-related accuracy loss). As to claims 2, 10 and 17, Khaliany discloses wherein determining the per tile scale is performed without the compute unit having to separately determine a scale for an entire tensor, an entire channel, or an entire head (See para. [0027], para. [0028], para. [0031] and para. [0032], the system subdivides the input-channel dimension into multiple vectors and determines a separate scale at the vector level, rather than requiring a scale for the entire tensor or channel). As to claims 3, 11 and 18, Khaliany discloses wherein the resulting matrix, and the quantization operation are part of activation or attention of a machine learning (ML) model (See para. [0024], para. [0026]-para.[0027], para. [0031], para. [0032] and para. [0062], input activations are quantified and scaled by a neural network and the weight-activation products are accumulated to produce output activations. The quantization unit receives a vector of output activations and computes the corresponding activation scale factor, the quantitation until includes a vector max unit 205 and a parameter quantization unit 215. Dynamic calibration is performed for each layer of a neural network, specifically, dynamic calibration is performed once for each vector of output activations of a multi-dimensional output). As to claims 5, 13 and 19, Khaliany discloses wherein determining the per tile scale comprises: deriving the per tile scale from either the first input or the resulting matrix (See para. [0031]-para. [0032] and Figure 2A, the quantization unit receives a vector of output activations and computes a corresponding activation scale factor based on the values of the output activation vector, including determining a maximum absolute value of the parameters of the vector and calculating the corresponding scale factor). Referring to claim 9, Khaliany discloses a hardware accelerator, comprising: a plurality of compute units, each comprising circuitry and registers, wherein the registers of each of the plurality of compute units (See para. [0093], para. [0094] and Figure 3B, processor 375 includes a plurality of microprocessor cores 370 and a plurality of matrix multiplying accumulators MMAs 365) are configured to store a respective first input comprises a different subset of data in a tensor, a channel, or a head (See para. [0093], para. [0094] and Figure 3B, processor 375 includes a plurality of microprocessor cores 370 and a plurality of matrix multiplying accumulators MMAs 365, wherein the MMAs 365 may comprise tensor cores, matrix multiply accelerators, or tensor processing units), wherein the circuitry in each of the plurality of compute units is configured to: perform an operation in a machine learning (ML) model using the first input to generate resulting data (See para. [0024], para. [0027], para. [0123], weights and/or activations values are quantized for input to a layer for a neural network. Weight-activation products are computed and summed up to produce vector dot-products), determine a per tile scale (See para. [0024], para. [0061]-para. [0064] and Figure 2A, a scale factor is computed for each vector of parameters, including determining a maximum absolute value of the parameters in the vector and calculating a corresponding scale factor [e.g., Equations 7a-7j], the scale factors are computed at a per-vector granularity, each vector representing a subset of data of the tenor, note Khailany’s vector corresponds to the claimed tile since the vector represents a subset of data of the input activation tenor, Applicant’s specification [0015] and para. [0018] describes a tile as a subset of data in a tensor, channel, or head and further states that the term “tile” does not necessary imply a two-dimensional arrangement), and perform a quantization operation, using the per tile scale, on at least one of the first input or the resulting data (See para. [0024], para. [0031], performing a quantization using fine-gained per-vector scale factors to mitigate quantization-related accuracy loss). Referring to claim 16, Khaliany discloses a computing system, comprising: a processor; memory (See para. [0093] and para. [0094] and Figure 3B, a processor 375 is coupled to memory 360) configured to store a training or inference application for a ML model (See para. [0093] and para. [0094], processor 375 and memory 360 implement a neural network using quantized wights and/or activations) ; and a compute unit, comprising: memory configured to store a first input that includes subset of data in a tensor, a channel, or a head (See para. [0024], para. [0027], para. [0028] and Figure 1A, storing an input activation tenor 101 subdivided along the input-channel dimension into vectors of parameters 105, each vector representing a subset of data of the input activation tenor); and a matrix multiplier comprising circuitry configured to multiply the first input with a second input to generate a resulting matrix as part of executing the training or inference application (See para. [0024], para. [0027], para. [0123], weights and/or activations values are quantized for input to a layer for a neural network. Weight-activation products are computed and summed up to produce vector dot-products), wherein the compute unit is configured to: determine a per tile scale (See para. [0024], para. [0061]-para. [0064] and Figure 2A, a scale factor is computed for each vector of parameters, including determining a maximum absolute value of the parameters in the vector and calculating a corresponding scale factor [e.g., Equations 7a-7j], the scale factors are computed at a per-vector granularity, each vector representing a subset of data of the tenor, note Khailany’s vector corresponds to the claimed tile since the vector represents a subset of data of the input activation tenor, Applicant’s specification [0015] and para. [0018] describes a tile as a subset of data in a tensor, channel, or head and further states that the term “tile” does not necessary imply a two-dimensional arrangement), and perform a quantization operation, using the per tile scale, on at least one of the first input or the resulting matrix (See para. [0024], para. [0031], performing a quantization using fine-gained per-vector scale factors to mitigate quantization-related accuracy loss). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 4 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Khaliany (US 2022/0067530 A1) and in view of Patwari (US 2023/0297824 A1). As to claims 4 and 12, Khaliany does not explicitly disclose the quantization operation is performed to de-quantize the resulting matrix in order to perform at least one of a softmax operation that is part of attention of the ML model. Patwari discloses the quantization operation is performed to de-quantize the resulting matrix in order to perform at least one of a softmax operation that is part of attention of the ML model (See para. [0028] and para. [0093], each Transformer and Bert neural networks uses an “attention” block. The attention block generally includes a linear matrix multiply (MatMul), followed by a non-linear GeLU/SoftMax/Erf functions depending on the particular architecture of the neural network, note in para. [0093] teaches both quantizer and/or dequantizer layers). Therefore, it would have been obvious to a person of ordinary skill in the computer art before the effective filing date of the claimed invention to modify Khailany’s quantized neural-network processing to include a dequantizer layer for de-quantizing the result matrix prior to perform a Softmax operation, as taught by Patwari, in order to provide the floating-point values used for subsequent neural-network processing while permitting quantized processing to reduce hardware and computational requirements (See Patwari, para. [0093]). Claims 6, 14 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Khaliany (US 2022/0067530 A1) and in view of Son (US 2025/0278615 A1). As to claims 6, 14 and 20, Khaliany discloses deriving a per-tile per vector scale based on values of the first input or resulting matrix but does not explicitly disclose deriving the per-tile scale by calculating a min-max of the first input or resulting matrix. Son discloses calculating a scale value based on maximum and maximum values of input data (See para. [0323]). Therefore, it would have been obvious to a person of ordinary skill in the computer art before the effective filing date of the claimed invention to modify Khailany’s scale-determination operation to derive the per-tile scale using minimum and maximum values of the input data, as taught by Son in order to provide a scale value that reflects the data distribution of the input data, thereby reducing deterioration of inference accuracy due to quantization errors. One of ordinary skill in the art would have been reasonably expected success because Son teaches that the scale value maybe calculated using the maximum and minimum values of the input parameters (See para. [0323]), which may be applied to Khailany’s scale determination without changing the underlying quantization operation. Claims 7 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Khaliany (US 2022/0067530 A1) and in view of Segall (US 2008/0253672A1). As to claims 7 and 15, Khaliany discloses deriving the per tile scale corresponding to the tensor, the channel, or the head and Khaliany does not explicitly disclose a historical scale. Segall discloses determining a scale factor for current data using previously determined scale factors (See para. [0069] and para. [0097], a scale factor for a current block may be predicted from scale factors of previously transmitted blocks, the scale and offset values associated with previous blocks are used in predicting and refining scales and offset values for a current block). Therefore, it would have been obvious to a person of ordinary skill in the computer art before the effective filing date of the claimed invention to modify the per-vector scaling technique scale of Khaliany to derive a scale factor for a current vector using previously determined factors, as taught by Segall, in order to use previously available scaling information to the scale for current data and thereby facilitate determination of the current scale factor. (See Segall, para. [0046] and para. [0047]). One of ordinary skill in the art would have reasonably expected success because Segall’s technique uses previously determined scale information to determine scaling information for subsequently processed data, which could be applied to Khailany’s successive subsets of tensor data without changing the underlying quantization function. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Khaliany (US 2022/0067530 A1) and in view of Pandit (CN 119585715 A). As to claim 8, Khaliany discloses determining the per-tile/per-vector scale and performing the quantitation operation using the determined scale. Pandit discloses executing a first kernel and a second kernel to perform machine-learning operations (See Figure 1, and Kernal 104 “a method includes receiving a first instruction by a virtual machine running on electronic hardware. The method includes parsing a first instruction using a virtual machine to determine a first kernel from a plurality of kernels coupled to the virtual machine. The method includes configuring, by a virtual machine, a first kernel with configuration data to perform an operation specified by a first instruction. The configuration data specifies a buffer containing input data of the first kernel and a buffer storing data generated by the first kernel. The method includes using a virtual machine to cause the first kernel to perform the configured operation. In the example of Figure 1, the kernel 104 may implement any of a variety of different machine learning functions. For example, the core 104-1 may be a GEMM core. The core 104-2 may be a BiasADD core. Another kernel 104 may be a ReLU kernel. A further kernel 104 may be a re-quantization kernel (e.g., a kernel capable of performing shifting and scaling functions). Other kernel 104 may implement other machine learning functions, such as Gaussian error linear unit (GELU), layer normalization, Softmax, and the like. The exemplary machine learning function provided herein is intended as a non-exhaustive list of machine learning functions that can be implemented by the kernel 104”). Therefore, it would have been obvious to a person of ordinary skill in the computer art before the effective filing date of the claimed invention to implement Khailany’s scale-determination operation and quantization operation using respective kernels, as taught by Pandit, such that a first kernel determines the per-tile scale and a second kernel performs the quantization operation, in order to provide separate configurable kernels for performing different machine-learning processing functions, thereby facilitating efficient implement and flexibility in executing the respective operation. One of ordinary skill in the art would have reasonably expected success because Pandit’s teaches configurable kernels capable of implementing different machine-learning functions, including a re-quantization kernel capable of performing scaling functions, and Khailany’s scale determine and quantization machine-learning quantization processing operations suitable for implementation using such kernels. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Weston et al. (US 2008/0215513 A1) discloses a pre-processing step prior to training a learning machine. The pre-processing step includes reducing the quantity of features to be processed using feature selection methods selected from the group consisting of recursive feature elimination (RFE), minimizing the number of non-zero parameters of the system (l.sub.0-norm minimization), evaluation of cost function to identify a subset of features that are compatible with constraints imposed by the learning set, unbalanced correlation score and transductive feature selection. The features remaining after feature selection are then used to train a learning machine for purposes of pattern classification, regression, clustering and/or novelty detection. Gokmen et al. (US 2023/0195832 A1) discloses a system comprises a processor, and a resistive processing resistive processing unit coupled to the processor. The resistive processing unit comprises an array of cells, wherein the cells respectively comprise resistive memory devices, wherein at least a portion of the resistive memory devices are programmable to store weight values of a given matrix in the array of cells. The processor is configured to store the given matrix in the array of cells of the resistive processing unit, and perform a calibration process to generate a first set of calibration parameters for calibrating forward pass matrix-vector multiplication operations performed on the stored matrix in the array of cells of the resistive processing unit, and a second set of calibration parameters for calibrating backward pass matrix-vector multiplication operations performed on a transpose of the stored matrix in the array of cells of the resistive processing unit. Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUK TING CHOI whose telephone number is (571)270-1637. The examiner can normally be reached Monday-Friday 9am-6pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, AMY NG can be reached at 5712701698. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YUK TING CHOI/ Primary Examiner, Art Unit 2164
Read full office action

Prosecution Timeline

May 31, 2024
Application Filed
Aug 19, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748954
DATA IMPUTATION USING AN INTERCONNECTED VARIATIONAL AUTOENCODER MODEL
3y 1m to grant Granted Sep 29, 2026
Patent 12749025
LOCAL EXPLANATION OF BLACK BOX MODEL BASED ON CONSTRAINED PERTURBATION AND ENSEMBLE-BASED SURROGATE MODEL
3y 0m to grant Granted Sep 29, 2026
Patent 12743448
RESEARCH AND INVESTIGATION SYSTEMS INCORPORATING GRAPH DATABASES
1y 11m to grant Granted Sep 22, 2026
Patent 12737397
DYNAMIC RESPONSE ENGINE
1y 7m to grant Granted Sep 15, 2026
Patent 12730817
SYSTEMS AND METHODS FOR LEARNING-BASED NETWORK
4y 0m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
72%
Grant Probability
99%
With Interview (+36.4%)
3y 2m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 673 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month