Prosecution Insights
Last updated: October 02, 2026
Application No. 18/622,237

METHOD AND APPARATUS FOR JUST-IN-TIME QUANTIZATION FOR MACHINE LEARNING

Non-Final OA §103
Filed
Mar 29, 2024
Examiner
PETRANEK, JACOB ANDREW
Art Unit
2183
Tech Center
2100 — Computer Architecture & Software
Assignee
Advanced Micro Devices Inc.
OA Round
3 (Non-Final)
80%
Grant Probability
Favorable
3-4
OA Rounds
1y 3m
Est. Remaining
88%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
624 granted / 781 resolved
+24.9% vs TC avg
Moderate +9% lift
Without
With
+8.6%
Interview Lift
resolved cases with interview
Typical timeline
3y 9m
Avg Prosecution
22 currently pending
Career history
817
Total Applications
across all art units

Statute-Specific Performance

§101
4.2%
-35.8% vs TC avg
§103
57.5%
+17.5% vs TC avg
§102
16.1%
-23.9% vs TC avg
§112
14.1%
-25.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 781 resolved cases

Office Action

§103
DETAILED ACTION Claims 1-20 are pending. A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 4/1/2026 has been entered. The office acknowledges the following papers: IDS filed on 5/8/2025, Claims and remarks filed on 4/1/2026. New Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 8, 14-15, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Yang (U.S. 2022/0137963), in view of Peng et al. (WO 2024/168514). As per claim 1: Yang disclosed an apparatus comprising: circuitry (Yang: Figure 1 element 10, paragraph 20) configured to: send a data value having a first data format to an accelerator (Yang: Figures 1 and 4 elements 12 and 130, paragraphs 22, 28, 30, 37, 48, 50, and 52-53)(The external memory sends data values to the type conversion data mover (i.e. accelerator) for selective conversion between data formats.); and cause circuitry of the accelerator to: replace the first data format of the data value with a second data format different from the first data format (Yang: Figures 1 4, and 6 elements 130 and S204, paragraphs 28, 48, 50-54, and 69-71)(The type conversion data mover performs a down conversion of input data from external memory.); and store the data value with the second data format in a memory accessible by a parallel data processing circuit for use during execution of a data model (Yang: Figures 1 4, and 6 elements 130 and S204, paragraphs 23, 28, 48, 50-54, and 69)(The type conversion data mover performs a down conversion of input data from external memory. The result is stored in the internal memory that is accessed by the polymorphic operator array during neural network model (i.e. data model) processing.). Yang failed to teach wherein the accelerator performs the data format replacement while the parallel data processing circuit executes an operation of the data model that is different from an operation of the data model that uses the data value. However, Peng combined with Yang disclosed wherein the accelerator performs the data format replacement while the parallel data processing circuit executes an operation of the data model that is different from an operation of the data model that uses the data value (Peng: Figures 4B and 5, translation pages 10-11)(Yang: Figure 1 elements 130 and 160, paragraphs 28 and 30-32)(The Polymorphic operator array (i.e. parallel data processing circuit) performs matrix operations that are physically distinct from the conversion operations done by the type converters (i.e. accelerators) using separate circuitry. The converters are configured to operate in parallel with the polymorphic operator array, but Yang doesn’t explicitly describe instances of parallel operation of these elements. Peng disclosed parallel data fetches for neural network layers with vector matrix multiplication (see in Peng: “Therefore, through the above-mentioned design of data access and calculation decoupling, data access and calculation execution are allowed to overlap in time, for example, the cache read data X2 and the macro unit 1 implement VMM1 (vector matrix multiplication) calculation are executed in parallel.” - Emphasis added). The combination allows for fetching/converting data inputs to be stored in the internal memory to be performed in parallel with execution operations on the Polymorphic Operator Array.). The advantage of performing parallel data fetches with execution operations is that the next set of data can be made available faster for the execution circuitry. Thus, it would have been obvious to one of ordinary skill in the art at the time of the effective filing date to implement the parallel fetching and execution method of Peng within the system of Yang for the above advantage. As per claim 8: Claim 8 essentially recites the same limitations of claim 1. Claim 8 additionally recites the following limitations: sending, by the circuitry to the accelerator, a first indication to cause circuitry of the accelerator (Yang: Figure 1 elements 110 and 130, paragraph 26)(The instruction analyzer determines changes in precision needed by an instruction and controls the type conversion data mover to perform a given conversion.). As per claim 14: Yang and Peng disclosed the apparatus as recited in claim 3, wherein the second data format has less precision than the first data format (Yang: Figures 14, and 6 elements 130 and S204, paragraphs 28, 48, 50-54, and 69)(The type conversion data mover performs a down conversion of input data from external memory). As per claim 15: Claim 15 essentially recites the same limitations of claim 1. Therefore, claim 15 is rejected for the same reasons as claim 1. As per claim 19: Yang and Peng disclosed the computing system as recited in claim 15, wherein at least one complete layer of the data model is executed by the parallel data processing circuit between initiation of the data format replacement and execution of the second layer that uses the data value (Peng: Figures 4B and 5, translation pages 10-11)(Yang: Figure 1 elements 130 and 160, paragraphs 28 and 30-32)(The Polymorphic operator array (i.e. parallel data processing circuit) performs matrix operations that are physically distinct from the conversion operations done by the type converters (i.e. accelerators) using separate circuitry. The converters are configured to operate in parallel with the polymorphic operator array, but Yang doesn’t explicitly describe instances of parallel operation of these elements. Peng disclosed parallel data fetches for neural network layers with vector matrix multiplication (see in Peng: “Therefore, through the above-mentioned design of data access and calculation decoupling, data access and calculation execution are allowed to overlap in time, for example, the cache read data X2 and the macro unit 1 implement VMM1 (vector matrix multiplication) calculation are executed in parallel.” - Emphasis added). The combination allows for fetching/converting data inputs to be stored in the internal memory to be performed in parallel with execution operations on the Polymorphic Operator Array. The overlapping fetches and execution also occurs for fetching and processing of different neural network layers. Additionally, execution of a first layer can be completed prior to execution of a subsequent layer (e.g. second, third, etc.).). Claims 2-7, 9-13, 16-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Yang (U.S. 2022/0137963), in view of Peng et al. (WO 2024/168514), in view of Official Notice. As per claim 2: Yang and Peng disclosed the apparatus as recited in claim 1, wherein the circuitry is further configured to allow the data value with the second data format to be overwritten in the memory, responsive to the parallel data processing circuit having completed use of the data value with the second data format (Yang: Figure 1 elements 140 and 160, paragraphs 28, 31, and 37)(The internal memory stores outputs from the type conversion data mover. The polymorphic operator array uses the stored data as inputs for matrix processing, which is later stored in the internal memory. Official notice is given that memory structures can use a least recently used model to overwrite cached data for the advantage of overwriting data that isn’t likely to be used. Thus, it would have been obvious to one of ordinary skill in the art to implement such overwriting in the internal memory. After the processing result from the polymorphic operator array is stored in internal memory, the converted input data is overwritten when it reaches a least recently used status and memory capacity for new conversions or processing outputs is requested.). As per claim 3: Yang and Peng disclosed the apparatus as recited in claim 1, wherein: the data model is a machine learning data model (Yang: Figure 1 element 100, paragraph 23); and the data value is one of a weight value, an activation value, or a gradient value (Yang: Figures 1 and 4 elements 12 and 130, paragraphs 3, 23, 28, and 31-32)(Official notice is given that neural networks process matrix operations using input activations and weight values for the advantage of performing training or inference operations. Thus, it would have been obvious to one of ordinary skill in the art that the input values being converted are weight and/or activation values.). As per claim 4: Yang and Peng disclosed the apparatus as recited in claim 1, wherein the second data format is selected based on a memory address range associated with a storage location of the data value (Yang: Figure 1 elements 12, 110, and 130, paragraphs 26 and 28)(The data mover receives operation data for adjusting precision data. Official notice is given that load operations can gather a range of data from memory for the advantage of reducing the number of load operations to gather a set of data. Thus, it would have been obvious to one of ordinary skill in the art that the operations adjusting precision on input data load a range of input data from memory.). As per claim 5: Yang and Peng disclosed the apparatus as recited in claim 3, wherein the circuitry is further configured to cause the accelerator to perform the data format replacement based on one or more of monitored activity levels of the accelerator and sizes of arrays processed by the data model (Yang: Figure 1 elements 110 and 130, paragraph 26)(The instruction analyzer determines changes in precision needed by an instruction and controls the type conversion data mover to perform a given conversion. Official notice is given that schedulers can be aware of if execution circuits are currently in use prior to issuing instructions to them for the advantage of ensuring correct processing results. Thus, it would have been obvious to one of ordinary skill in the art to implement a scheduler in Yang that is aware of processing loads on the converters and operator array.). As per claim 6: Yang and Peng disclosed the apparatus as recited in claim 1, wherein the accelerator is selected from a plurality of accelerators based at least in part on a current load of the accelerator (Yang: Figure 1 elements 12, 110, and 130, paragraphs 26 and 28)(The data mover and type converter both read upon the accelerators. Official notice is given that schedulers can be implemented to select between multiple circuits that can perform the same operation based on availability for the advantage of selecting an accelerator that is currently free. Thus, it would have been obvious to one of ordinary skill in the art to implement a selection method to select between the accelerators based on usage and the conversion needed.). As per claim 7: Yang and Peng disclosed the apparatus as recited in claim 1, wherein the circuitry is further configured to cause the accelerator to perform the data format replacement for a first layer of the machine learning data model while the parallel data processing circuit executes an operation corresponding to a second layer of the machine learning data model, the second layer being different from the first layer (Peng: Figures 4B and 5, translation pages 10-11)(Yang: Figure 1 elements 130 and 160, paragraphs 28 and 30-32)(The Polymorphic operator array (i.e. parallel data processing circuit) performs matrix operations that are physically distinct from the conversion operations done by the type converters (i.e. accelerators) using separate circuitry. The converters are configured to operate in parallel with the polymorphic operator array, but Yang doesn’t explicitly describe instances of parallel operation of these elements. Peng disclosed parallel data fetches for neural network layers with vector matrix multiplication (see in Peng: “Therefore, through the above-mentioned design of data access and calculation decoupling, data access and calculation execution are allowed to overlap in time, for example, the cache read data X2 and the macro unit 1 implement VMM1 (vector matrix multiplication) calculation are executed in parallel.” - Emphasis added). The combination allows for fetching/converting data inputs to be stored in the internal memory to be performed in parallel with execution operations on the Polymorphic Operator Array. The overlapping fetches and execution also occurs for fetching and processing of different neural network layers.). As per claim 9: The additional limitation(s) of claim 9 basically recite the additional limitation(s) of claim 2. Therefore, claim 9 is rejected for the same reason(s) as claim 2. As per claim 10: The additional limitation(s) of claim 10 basically recite the additional limitation(s) of claim 3. Therefore, claim 10 is rejected for the same reason(s) as claim 3. As per claim 11: The additional limitation(s) of claim 11 basically recite the additional limitation(s) of claim 4. Therefore, claim 11 is rejected for the same reason(s) as claim 4. As per claim 12: The additional limitation(s) of claim 12 basically recite the additional limitation(s) of claim 5. Therefore, claim 12 is rejected for the same reason(s) as claim 5. As per claim 13: The additional limitation(s) of claim 13 basically recite the additional limitation(s) of claim 6. Therefore, claim 13 is rejected for the same reason(s) as claim 6. As per claim 16: The additional limitation(s) of claim 16 basically recite the additional limitation(s) of claim 2. Therefore, claim 16 is rejected for the same reason(s) as claim 2. As per claim 17: The additional limitation(s) of claim 17 basically recite the additional limitation(s) of claim 3. Therefore, claim 17 is rejected for the same reason(s) as claim 3. As per claim 18: The additional limitation(s) of claim 18 basically recite the additional limitation(s) of claim 4. Therefore, claim 18 is rejected for the same reason(s) as claim 4. As per claim 20: The additional limitation(s) of claim 20 basically recite the additional limitation(s) of claim 6. Therefore, claim 20 is rejected for the same reason(s) as claim 6. Response to Arguments The arguments presented by Applicant in the response, received on 4/1/2026 are considered persuasive. Applicant argues for claims 1, 8, and 15: “Accordingly, Yang discloses type conversion in preparation for execution of the same neural network operation that consumes the converted data. Yang does not disclose or suggest that conversion for a first operation proceeds while the operator array executes a different operation of the data model that does not depend on the data value being converted. Stated differently, Yang's disclosure ties conversion to the data path of the operation that uses the converted data. There is no teaching that conversion of data for one Amended claim 1 expressly recites that the concurrently executed operation be different from the operation that uses the data value. This requirement excludes mere internal pipelining within a single neural network operation or layer. Internal pipelining processes sub-steps of the same model operation. In contrast, claim 1 requires that conversion occur while the parallel data processing circuit executes a different operation of the data model using data other than the data value being converted. Yang does not disclose or suggest such inter-operation overlap or decoupled scheduling. In the Office Action, the Examiner relied in part on Official Notice for the proposition that components "may operate concurrently." Applicant respectfully submits that Official Notice is not appropriate in view of amended claim 1. While it may be generally known that hardware components can be capable of concurrent operation, claim 1 does not recite mere structural capability. Rather, claim 1 affirmatively requires a specific execution relationship, namely: … This is not a general property of hardware concurrency, but a specific architectural and scheduling behavior. Such a limitation must be supported by an express teaching or a reasoned explanation grounded in the prior art. Yang does not disclose such behavior, and the Office Action has not identified any other reference that does. To the extent the Office maintains reliance on Official Notice for this limitation, Applicant respectfully traverses that reliance and requests citation of supporting evidence in accordance with MPEP §2144.03..” This argument is found to be persuasive for the following reason. The examiner agrees that Yang in view of the previous official notice failed to teach the newly claimed limitation. However, a new ground of rejection has been given due to the amendment. Conclusion The following is text cited from 37 CFR 1.111(c): In amending in reply to a rejection of claims in an application or patent under reexamination, the applicant or patent owner must clearly point out the patentable novelty which he or she thinks the claims present in view of the state of the art disclosed by the references cited or the objections made. The applicant or patent owner must also show how the amendments avoid such references or objections. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JACOB A. PETRANEK whose telephone number is (571)272-5988. The examiner can normally be reached on M-F 8:00-4:30. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jyoti Mehta can be reached on (571) 270-3995. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JACOB PETRANEK/Primary Examiner, Art Unit 2183
Read full office action

Prosecution Timeline

Show 1 earlier event
Apr 08, 2025
Non-Final Rejection mailed — §103
Sep 08, 2025
Response Filed
Dec 01, 2025
Final Rejection mailed — §103
Feb 13, 2026
Examiner Interview Summary
Feb 13, 2026
Applicant Interview (Telephonic)
Apr 01, 2026
Request for Continued Examination
Apr 07, 2026
Response after Non-Final Action
Aug 10, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748722
COMPUTE IN-MEMORY ARCHITECTURE FOR CONTINUOUS ON-CHIP LEARNING
2y 11m to grant Granted Sep 29, 2026
Patent 12748625
LOCAL LAUNCH IN WORKGROUP PROCESSORS
2y 9m to grant Granted Sep 29, 2026
Patent 12748592
PROCESSOR EMBEDDED STREAMING BUFFER
1y 9m to grant Granted Sep 29, 2026
Patent 12724610
SIMULATION APPARATUS, SIMULATION METHOD, AND NON-TRANSITORY COMPUTER READABLE MEDIUM
1y 7m to grant Granted Sep 01, 2026
Patent 12717583
PROCESSOR MACRO-OPERATION FUSION
2y 10m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
80%
Grant Probability
88%
With Interview (+8.6%)
3y 9m (~1y 3m remaining)
Median Time to Grant
High
PTA Risk
Based on 781 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month