Prosecution Insights
Last updated: October 02, 2026
Application No. 17/743,327

PERFORMING MATRIX VALUE INDICATION

Non-Final OA §103
Filed
May 12, 2022
Priority
May 13, 2021 — provisional 63/188,406
Examiner
DE LA GARZA, CARLOS HEBERTO
Art Unit
2182
Tech Center
2100 — Computer Architecture & Software
Assignee
NVIDIA Corporation
OA Round
3 (Non-Final)
68%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 68% — above average
68%
Career Allowance Rate
13 granted / 19 resolved
+13.4% vs TC avg
Strong +43% interview lift
Without
With
+42.9%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
19 currently pending
Career history
42
Total Applications
across all art units

Statute-Specific Performance

§101
14.4%
-25.6% vs TC avg
§103
46.3%
+6.3% vs TC avg
§102
12.4%
-27.6% vs TC avg
§112
26.4%
-13.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 19 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This Action is final and is in response to the claims filed 02/24/2026. Claims 1-32 are currently pending, of which claims 1-32 are currently rejected. Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 08/12/2026 has been entered. Response to Arguments Applicant’s arguments filed on 02/24/2026 have been fully considered. 35 U.S.C. 103: Applicant’s arguments regarding the 35 U.S.C. 103 rejection have been fully considered. Applicant argues in pages 9-10 that Kjolstad in view of Kulkarni and further in view of Berg do not teach amended claim 1. Applicant specifically argues: “At best, Kjolstad appears to disclose "a C++ library called taco," where "taco generates efficient code for simpler kernels such as sparse matrix-vector multiplication as well as complex kernels." Kjolstad at 77:2. However, Kjolstad fails to disclose any single instruction that is used to "indicate one or more non-zero values within one or more matrices of data, wherein the single instruction is to be compiled at runtime to be used by one or more accelerators." Kulkarni generally discusses JIT compilation policies for managed-language virtual machines such as Java and C#. See Kulkarni, § 1, Introduction. As such, Kulkarni is directed to generic VM runtime compilation and fails to disclose compilation of a single instruction for execution by accelerators. Nor does Berg cure the deficiency of Kjolstad and Kulkarni. Berg explains how to generate efficient sparse code for GPU accelerators and mixed sparse-dense operations that require tiling. See Berg: Pages 92-93, Section 6.7 Conclusion and Section 7.5. However, Berg does not disclose runtime compilation of the claimed single instruction for use by accelerators. Applicant’s arguments regarding amended claim 1 are found persuasive. However, see new grounds of rejection necessitated by amendments. Applicant further argues in pages 10-11 that Kjolstad in view of Kulkarni and further in view of Berg do not teach amended claim 25. Applicant specifically argues: “Applicant respectfully submits that claim 25 is allowable under 35 U.S.C. § 103 over Kjolstad in view of Kulkarni and further in view of Berg. In particular, amended claim 25 recites "generat[ing] an array indicating one or more indices of the one or more non-zero values within the one or more matrices of data, wherein the array is to be used as an operand to another instruction to generate a compressed representation of the one or more matrices of data." For at least the reasons discussed above in connection with claim 1, the proposed combination of Kjolstad, Kulkarni, and Berg fails to teach or suggest this claimed subject matter. At best, the cited references appear to discuss sparse tensor compilation, generic runtime/JIT compilation policies, and sparse code generation for accelerator execution separately. However, the proposed combination of Kjolstad in view of Kulkarni and further in view of Berg does not teach or suggest any use of "only the hardware-independent instruction to generate executable instructions" and "generat[ing] an array indicating one or more indices of the one or more non- zero values within the one or more matrices of data, wherein the array is to be used as an operand to another instruction to generate a compressed representation of the one or more matrices of data." Accordingly, Applicant respectfully submits that claim 25 is allowable under 35 U.S.C. § 103 over Kjolstad in view of Kulkarni and further in view of Berg. Applicant's arguments fail to comply with 37 CFR 1.111(b) because they amount to a general allegation that the claims define a patentable invention without specifically pointing out how the language of the claims patentably distinguishes them from the references. See new grounds of rejection below necessitated by amendments. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-6, 5-14, 16-22, 24-30, and 32 are rejected under 35 U.S.C. 103 as being unpatentable over Fredrik Kjolstad in NPL: “The Tensor Algebra Compiler” (dl.acm.org/doi/pdf/10.1145/3133901), hereinafter “Kjolstad”, in view of Prasad A. Kulkarni in NPL: “JIT Compilation Policy for Modern Machines” (www.ittc.ku.edu/~kulkarni/CARS/papers/oopsla11.pdf), hereinafter “Kulkarni”, in view of Fredrik Berg Kjølstad in NPL: “Sparse Tensor Algebra Compilation” (tensor-compiler.org/files/kjolstad-phd-thesis-taco-compiler.pdf), hereinafter “Berg”, further in view of Frazier et al. (U.S. Patent No.: US 10691459 B2), hereinafter “Frazier”. Regarding Claim 1, Kjolstad teaches: One or more processors, comprising: circuitry to use [an instruction] to indicate one or more non-zero values within one or more matrices of data, … (Page 77:17 Second paragraph, e.g., Compiler runs in Processor and memory (circuitry); Page 77:5, Sparse storing technique, e.g., Taco (Tensor Algebra Compiler) performs the operations of storing indices of non-zero values from a dimension (one or more matrices of data) in array idx;). Kjolstad does not teach: One or more processors, comprising: circuitry to use only a single instruction to indicate one or more non-zero values within one or more matrices of data, wherein the single instruction is to be compiled at runtime to be used by one or more accelerators. However, in the same field of endeavor, Kulkarni teaches using JIT compilation for applications written in managed languages, such as Java and C#. Kulkarni explains “interpreted execution is inherently slow, which makes dynamic or Just-in-Time (JIT) compilation essential to achieve efficient runtime performance for such applications” (Kulkarni: First page, 1. Introduction, First paragraph). Additionally, Kjolstad discloses in Page 77:26, “Runtime Support”, how future work involves taking advantage of JIT compilation to support properties available only at runtime. Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the JIT compilation for compiling applications as taught by Kulkarni with the tensor algebra compiler as taught by Kjolstad. One would have been motivated to combine these references because both references disclose JIT compilation, and Kulkarni enhances the model of Kjolstad because “Just-in-Time (JIT) compilation essential to achieve efficient runtime performance for such applications.” (Kulkarni: First page, 1. Introduction, First paragraph). Kjolstad in view of Kulkarni do not teach: One or more processors, comprising: circuitry to use only a single instruction to indicate one or more non-zero values within one or more matrices of data, wherein the single instruction is to be compiled at runtime to be used by one or more accelerators. However, in the same field of endeavor, Berg teaches performing sparse matrix multiplication using the Tensor Algebra Compiler in a GPU accelerator. Berg explains “And Section 7.5 shows how they let us generate efficient sparse code for GPU accelerators and mixed sparse-dense operations that require tiling.”. See Berg: Pages 92-93, Section 6.7 Conclusion. In addition, Kjolstad discloses that their future work includes removing restrictions and adapting code generation to target accelerators (e.g., GPUs and TPUs). See Kjolstad: Page 77:26 first paragraph. Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to modify the Tensor Algebra Compiler as taught by Kjolstad in view of Kulkarni to perform sparse matrix multiplication in a GPU accelerator as taught by Berg. One would have been motivated to combine these references because both references disclose the Tensor Algebra Compiler used for sparse matrix multiplication, and Berg enhances the model of Kjolstad in view of Kulkarni by allowing for the Tensor Algebra Compiler to "compile any sparse tensor algebra expression to CPU and GPU code that matches the performance of hand-optimized implementations" (Berg: Abstract). Kjolstad in view of Kulkarni in view of Berg do not teach: One or more processors, comprising: circuitry to use only a single instruction to indicate one or more non-zero values within one or more matrices of data, wherein the single instruction is to be compiled at runtime to be used by one or more accelerators. However, Frazier teaches converting multiple instructions into a single combined instruction. Frazier explains “The format may be selected based on the at least two program instructions determined to be converted into the single combined instruction. The format may indicate a portion of the single combined instruction that is an operation, the portion of the single combined instruction that is an operand, the portion of the single combined instruction that is a memory location for storing a final result of the single combined instruction, and/or the portion of the single combined instruction that is a program instruction.” (Frazier: Column 5, Lines 18-26) Kjolstad teaches a set of instructions written in C++ code. Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to modify the set of instructions written in C++ code as taught by Kjolstad to be a single combined instruction as taught by Frazier. One would have been motivated to combine these references because both references disclose executing program instructions, and Frazier enhances the model of Kjolstad in view of Kulkarni in view of Berg because “a length of the single combined instruction is less than a combined length of the two program instructions may be carried out by converting the at least two of the group of program instructions into a single combined instruction (322) that is less than the total length of the combination of the at least two program instructions.” (Frazier: Column 7 Line 63 – Column 8 Line 1). Regarding Claim 2, Kjolstad in view of Kulkarni in view of Berg in view of Frazier teach: The one or more processors of claim 1, wherein the circuitry is to indicate the one or more non-zero values by at least causing the one or more processors to store one or more indices corresponding to the one or more non-zero values in a memory (Kjolstad: Page 77:17 Second paragraph, e.g., Compiler runs in Processor and memory (one or more circuits); Page 77:5, Sparse storing technique, e.g., Taco (Tensor Algebra Compiler) performs the operations of storing indices of non-zero values in array idx) Berg further teaches: accessible to one or more graphics processing cores (Berg: Page 104, “Sparse Matrix-Vector Multiplication (SpMV)”, e.g., sparse matrix-vector multiplication (SPMV) is performed in a GPU (including gpu cores)). Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to modify the Tensor Algebra Compiler as taught by Kjolstad in view of Kulkarni in view of Berg to perform sparse matrix multiplication in a GPU as taught by Berg. One would have been motivated to combine these references because both references disclose the Tensor Algebra Compiler used for sparse matrix multiplication, and Berg enhances the model of Kjolstad in view of Kulkarni in view of Berg by allowing for the Tensor Algebra Compiler to "compile any sparse tensor algebra expression to CPU and GPU code that matches the performance of hand-optimized implementations" (Berg: Abstract). This modification would cause for the GPU as taught by Berg to compute sparse matrix multiplication in its cores. Hence, Kjolstad in view of Kulkarni in view of Berg teach Claim 2 in its entirety. Regarding Claim 3, Kjolstad in view of Kulkarni in view of Berg in view of Frazier teach: The one or more processors of claim 1, wherein the single instruction is to cause the one or more processors to store one or more indices of the one or more non-zero values in memory (Frazier: Column 5, Lines 18-26, e.g., converts multiple instructions into a single combined instruction; Kjolstad: Page 77:17 Second paragraph, e.g., Compiler runs in Processor and memory (one or more processors); Page 77:5, Sparse storing technique, e.g., Taco (Tensor Algebra Compiler) executes storing instructions to store indices of non-zero values in array idx) that is accessible to one or more threads when executing one or more sparse matrix multiplication operations in parallel (Kjolstad: Page 77:17, Section 8.2 Sparse Matrix-Vector Multiplication, e.g., Sparse matrix multiplication is performed in threads; Page 77:18, Top paragraph, e.g., Sparse matrix multiplication operations can be performed in parallel). The motivation to combine provided with respect to claim 1 applies equally to claim 3. Regarding Claim 4, Kjolstad in view of Kulkarni in view of Berg in view of Frazier teach: The one or more processors of claim 1, wherein the single instruction indicates the one or more non-zero values corresponding to a sparse matrix multiplication (Frazier: Column 5, Lines 18-26, e.g., converts multiple instructions into a single combined instruction; Kjolstad: Page 77:3, Fig. 3, e.g., shows an example of a sparse matrix multiplication instruction set using Tensor Algebra Compiler); and wherein the circuitry is to execute a compiler at a runtime to generate executable instructions for execution by the one or more accelerators (Kjolstad: Page 77:5, Sparse storing technique, e.g., Taco (Tensor Algebra Compiler) performs the operations of storing indices of non-zero values (by generating executable instructions); Kulkarni: First page, 1. Introduction, e.g., JIT compilation performs compilation at runtime; Berg: Pages 92-93, Section 6.7 Conclusion, e.g., sparse code is generated for GPU accelerators). The motivation to combine provided with respect to claim 1 applies equally to claim 4. Regarding Claim 5, Kjolstad in view of Kulkarni in view of Berg in view of Frazier teach: The one or more processors of claim 1, wherein the single instruction is to be compiled at a runtime by a compiler based on one or more first instructions with sparsity information of the one or more matrices of data (Frazier: Column 5, Lines 18-26, e.g., converts multiple instructions into a single combined instruction; Kjolstad: Page 77:5, Sparse storing technique, e.g., Taco (Tensor Algebra Compiler) executes storing instructions to store indices of non-zero values in array idx; Page 77:3, Fig. 2, e.g., shows array idx storing indices of non-zero elements (first instruction) for tensor B (one or more matrices)) and the one or more first instructions are to be compiled to generate one or more second instructions that are executable by a graphics processing unit (GPU) to perform a matrix multiplication operation with the sparsity information (Berg: Page 104 Sparse Matrix-Vector Multiplication (SpMV), e.g., Tensor Algebra Compiler generates kernel (second instruction) for Sparse matrix-vector multiplication in a GPU). The motivation to combine provided with respect to claim 2 applies equally to claim 5. Regarding Claim 6, Kjolstad in view of Kulkarni in view of Berg in view of Frazier teach: The one or more processors of claim 1, wherein the single instruction includes a half-precision matrix multiply and accumulate (HMMA) operation, integer matrix multiplication and accumulate (IMMA) operation, single-precision matrix multiplication operation, or a floating point multiplication and accumulate operation (Frazier: Column 5, Lines 18-26, e.g., converts multiple instructions into a single combined instruction; Kjolstad: Page 77:5, Fig. 12, e.g., shows matrices being a “double” data type (float). Line 20 performs matrix multiplication). The motivation to combine provided with respect to claim 1 applies equally to claim 6. Regarding Claim 8, Kjolstad in view of Kulkarni in view of Berg in view of Frazier teach: The one or more processors of claim 1, wherein to indicate the one or more non-zero values within the one or more matrices of data includes to cause the circuitry to cause a compiler to generate an operand (Kjolstad: Page 77:3, fourth paragraph, e.g., Only values that are in matching non-zero locations (operands) of matrices B and C are computed) that is to be used by one or more graphics processing cores to perform one or more matrix multiplication operations (Kjolstad: Page 77:3, Fig. 3, e.g., shows Tensor Algebra Compiler performing matrix multiplication; Berg: Page 104, “Sparse Matrix-Vector Multiplication (SpMV)”, e.g., sparse matrix-vector multiplication (SPMV) is performed in a GPU (including gpu cores)), and wherein the operand includes index information of the one or more non-zero values (Kjolstad: Page 77:5, Sparse storing technique, e.g., Taco (Tensor Algebra Compiler) performs the operations of storing indices of non-zero values in array idx). With regards to Claims 9-14 and 16, they are directed to instructions stored in a memory to be executed by the claimed processor above (claims 1-6 and 8 respectively), wherein all claim limitations also have been addressed and/or covered in cited areas. Thus, accordingly, these claims are rejected for at least the same reasons therein. Regarding Claims 17-22 and 24, they are media claims practiced by the apparatus of claims 1-6 and 8 respectively. They are rejected for the same reasons as claims 1-6 and 8. Regarding Claims 25-30 and 32, they are method claims practiced by the apparatus of claims 1-6 and 8 respectively. They are rejected for the same reasons as claims 1-6 and 8. Regarding additional limitations of Claim 25 that are distinct from claim 1, Kjolstad in view of Kulkarni in view of Berg in view of Frazier teach: generate an array indicating one or more indices of the one or more non-zero values within the one or more matrices of data, wherein the array is to be used as an operand to another instruction to generate a compressed representation of the one or more matrices of data (Kjolstad: Page 77:5, Sparse storing technique, e.g., Taco (Tensor Algebra Compiler) performs the operations of storing indices of non-zero values from a dimension (one or more matrices of data) in array idx; Page 77:12, 6.1 Code Generation Algorithm, e.g., array idx is used to merge idx sparse idx variables). The motivation to combine provided with respect to claim 1 applies equally to claim 25. Claims 7, 15, 23, and 31 are rejected under 35 U.S.C. 103 as being unpatentable over Kjolstad in view of Kulkarni in view of Berg in view of Frazier, further in view of Liao (U.S. Patent Application Publication No.: US 20160299874 A1), hereinafter “Liao”. Regarding Claim 7, Kjolstad in view of Kulkarni in view of Berg in view of Frazier teach the one or more processors of Claim 1. Kjolstad in view of Kulkarni in view of Berg in view of Frazier do not teach: wherein performing the single instruction causes a compiler to modify a Directed Acyclic Graph (DAG) interface to receive one or more instructions with sparsity information of the one or more matrices of data. However, in the same field of endeavor, Liao teaches how a Directed Acyclic Graph (DAG) can be used for matrix multiplication in software. Liao explains “A PLASMA implementation itself relies on runtime scheduling of parallel subtasks (i.e., functions such as matrix multiplications). However, some subtasks depend on others, and the relationships between the subtasks can be complex. The relationships may be expressed through a task graph, typically shown as a directed acyclic graph (DAG), which can be explored at runtime”. See Liao: ¶0100. Therefore, it would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to which said subject matter pertains to combine the Directed Acyclic Graph in software for matrix multiplications as taught by Liao with the Tensor Algebra Compiler for computing sparse matrix multiplication as taught by Kjolstad in view of Kulkarni in view of Berg in view of Frazier. One would have been motivated to combine these references because both references disclose matrix multiplication in software, and Liao enhances the model of Kjolstad in view of Kulkarni in view of Berg in view of Frazier by improving load balancing. See Liao: ¶0013. Combination of Kjolstad in view of Kulkarni in view of Berg in view of Frazier in view of Liao teach claim 7 in it’s entirety. With regards to Claim 15, it is directed to instructions stored in a memory to be executed by the claimed processor above (claim 7), wherein all claim limitations also have been addressed and/or covered in cited areas. Thus, accordingly, this claim is rejected for at least the same reasons therein. Regarding Claim 23, it is a media claim practiced by the apparatus of claim 7. It is rejected for the same reasons as claim 7. Regarding Claim 31, it is a method claim practiced by the apparatus of claims 7. It is rejected for the same reasons as claim 7. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARLOS H DE LA GARZA whose telephone number is (571)272-0474. The examiner can normally be reached Monday-Friday 9:30AM-6PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Caldwell can be reached at (571) 272-3702. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.H.D./ Carlos H. De La GarzaExaminer, Art Unit 2182 (571)272-0474 /EMILY E LAROCQUE/Primary Examiner, Art Unit 2182
Read full office action

Prosecution Timeline

Show 3 earlier events
Nov 25, 2025
Applicant Interview (Telephonic)
Feb 24, 2026
Response Filed
May 12, 2026
Final Rejection mailed — §103
Aug 07, 2026
Applicant Interview (Telephonic)
Aug 07, 2026
Examiner Interview Summary
Aug 12, 2026
Request for Continued Examination
Aug 14, 2026
Response after Non-Final Action
Aug 25, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737154
COMPUTING DEVICE AND METHOD USING MULTIPLIER-ACCUMULATOR
4y 4m to grant Granted Sep 15, 2026
Patent 12717863
CALCULATION DEVICE
4y 6m to grant Granted Aug 25, 2026
Patent 12688041
Vectorized Operations for Sparse Kernels
4y 2m to grant Granted Jul 21, 2026
Patent 12681695
WEIGHT STATIONARY IN-MEMORY-COMPUTING NEURAL NETWORK ACCELERATOR WITH LOCALIZED DATA MULTIPLEXING
4y 0m to grant Granted Jul 14, 2026
Patent 12675548
BITWISE PRODUCT-SUM ACCUMULATIONS WITH SKIP LOGIC
4y 4m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
68%
Grant Probability
99%
With Interview (+42.9%)
4y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 19 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month