Prosecution Insights
Last updated: October 02, 2026
Application No. 18/602,951

COMPRESSION OF MACHINE LEARNING MODELS VIA SPARSIFICATION AND QUANTIZATION

Non-Final OA §101§103§112
Filed
Mar 12, 2024
Priority
Sep 14, 2023 — provisional 63/538,465
Examiner
PHAKOUSONH, DARAVANH
Art Unit
Tech Center
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
25%
Grant Probability
At Risk
1-2
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants only 25% of cases
25%
Career Allowance Rate
1 granted / 4 resolved
-35.0% vs TC avg
Strong +100% interview lift
Without
With
+100.0%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
25 currently pending
Career history
41
Total Applications
across all art units

Statute-Specific Performance

§101
52.8%
+12.8% vs TC avg
§103
13.7%
-26.3% vs TC avg
§102
19.9%
-20.1% vs TC avg
§112
12.4%
-27.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 4 resolved cases

Office Action

§101 §103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 7-31 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 7 recites the limitation "the machine learning model.” There is insufficient antecedent basis for this limitation in the claim. Claims 8-31 are depend from claim 7 and therefore inherit the indefiniteness. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-51 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. 101 Subject Matter Eligibility Analysis Step 1: Claims 1-51 are within the four statutory categories (a process, machine, manufacture or composition of matter). Step 2A Prong One, Step 2A Prong Two, and Step 2B Analysis: Step 2A Prong One asks if the claim recites a judicial exception (abstract idea, law of nature, or natural phenomenon). If the claim recites a judicial exception, analysis proceeds to Step 2A Prong Two, which asks if the claim recites additional elements that integrate the abstract idea into a practical application. If the claim does not integrate the judicial exception, analysis proceeds to Step 2B, which asks if the claim amounts to significantly more than the judicial exception. If the claim does not amount to significantly more than the judicial exception, the claim is not eligible subject matter under 35 U.S.C. 101. None of the claims represent an improvement to technology. Claims 1-31 are directed to a method consisting of a series of steps, meaning that it is directed to the statutory category of process. Claims 32-51 are directed to storage mediums and processors which are machines. Regarding claim 1, the following claim elements are abstract ideas: to generate a plurality of sparse data structures (This is an abstract idea of a mental process. The limitation involves observing numerical values, evaluating how these values should be organized, and exercising judgement to arrange the values into multiple representations having designated zero values or patterns of zero values. A person could review the values, determine which entries should be represented as zero, and manually organize the values into separate lists or tables according to the selected sparsity pattern. This type of observation, evaluation, and judgement can be practically performed in the human mind with the aid of pen and paper or basic computational tools, and therefore falls within the mental process grouping of abstract ideas. See MPEP 2106.04(a)(2)(III).) quantizing the first data structure…to a first bit width (This is an abstract idea of a mental process. The limitation involves evaluating numerical values and using judgement to convert values to a selected numerical precision represented by the first bit width. A person could review the numerical values and manually round or otherwise represent the values using the selected precision with the aid of pen and paper or basic computational tools. This type of evaluation and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.), quantizing the second data structure… to a second bit width that is different from the first bit width (This is an abstract idea of a mental process. The limitation involves evaluating numerical values and using judgement to round or otherwise convert values to a representation having a selected bit width that differs from the bit width used for the first data structure. A person could review the numerical values and manually perform the rounding or conversion using pen and paper or basic computational tools. This type of evaluation and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.) The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: compressing a machine learning model having a plurality of values to reduce at least one of a size of the machine learning model or computation requirements of the machine learning model (This limitation merely states the intended result of performing the recited abstract processing and amounts to an instruction to apply the judicial exception to achieve the result. The limitation does not impose a meaningful limitation on the judicial exception.), processing the machine learning model to (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).) storing inlier values of the machine learning model in a first data structure, and storing outlier values of the machine learning model in a second data structure (The steps are merely generic data storage operations that amount to storing information in memory, which is well-understood, routine, and conventional activity. See MPEP 2106.05(d)(II)(iv).), wherein at least one of the first data structure or the second data structure has a structured sparse pattern (This limitation adds insignificant extra-solution activity by merely specifying the arrangement of the stored data associated with the abstract processing.); non-uniformly quantizing the machine learning model (This limitation merely provides an instruction to apply the recited abstract quantization operations and does not impose a meaningful limitation on the judicial exception.) Regarding claim 2, the rejection of claim 1 is incorporated herein. Further, claim 2 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the inlier values and the outlier values are weights of the machine learning model (This limitation merely limits the recited abstract processing to a particular type of data, namely weights of a machine learning model, and does not impose a meaningful limitation on the judicial exception.). Regarding claim 3, the rejection of claim 1 is incorporated herein. Further, claim 3 recites the following abstract ideas: wherein the inlier values and the outlier values are determined according to a defined threshold metric (This is an abstract idea of a mental process. The limitation involves observing and evaluating numerical values against a defined threshold metric and exercising judgement to determine whether the values are inliers or outliers. A person could review each value, compare it to a defined threshold, and classify the value accordingly using pen and paper or basic computational tools. This type of observation, evaluation, and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 4, the rejection of claim 1 is incorporated herein. Further, claim 4 recites the following abstract ideas: wherein the first data structure has a first structured sparse pattern that has less sparsity than a second structured sparse pattern of the second data structure (This is an abstract idea of a mental process. The limitation involves evaluating numerical values and exercising judgement to organize values into respective data structures having different degrees of sparsity. A person could arrange the values in separate tables or lists with different numbers or patterns of zero values using pen and paper or basic computational tools. This type of evaluation, organization, and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 5, the rejection of claim 1 is incorporated herein. Further, claim 5 recites the following abstract ideas: wherein the first data structure has a first structured sparse pattern that is the same as a second structured sparse pattern of the second data structure (This is an abstract idea of a mental process. The limitation involves evaluating the arrangement of values in two data structures and determining that the structures use the same sparsity pattern. A person could compare the number and placement of zero values in each structure and determine that the patterns are the same using pen and paper or basic computational tools. This type of observation and evaluation can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 6, the rejection of claim 1 is incorporated herein. Further, claim 6 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the first bit width and the second bit width are supported by different hardware accelerators (This limitation adds insignificant extra-solution activity by merely specifying the hardware used to support the different bit-width representations associated with the abstract processing.). Regarding claim 7, the following claim elements are abstract ideas: apportioning a plurality of different subsets of values of the machine learning model into a plurality of data structures at least one of which has a defined structured sparse pattern (This is an abstract idea of a mental process. The limitation involves observing and evaluating numerical values and exercising judgement to separate values into different groups and organize at least one group according to a defined sparsity pattern. A person could review the values, sort them into separate lists or tables, and arrange selected entries according a particular pattern of zero values using pen and paper or basic computational tools. The type of observation, evaluation, organization, and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.); and changing a data representation of at least one data structure of the plurality of data structures, wherein at least two data structures of the plurality of data structures have different data representations (This is an abstract idea of a mental process. The limitation involves evaluating how numerical values are represented and exercising judgement to change the representation of at least one group of values so that the groups have different representations. A person could review the values and manually represent one group differently from another, such as by using a different format or level of detail, with the aid of pen and paper or basic computational tools.). The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: the machine learning model (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) Regarding claim 8, the rejection of claim 7 is incorporated herein. Further, claim 8 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the machine learning model is a deep neural network (This limitation constitutes mere instructions to apply the abstract idea and insignificant extra-solution activity. See MPEP 2106.05(f) and 2106.05(g).). Regarding claim 9, the rejection of claim 7 is incorporated herein. Further, claim 9 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the machine learning model is a large language model (LLM) (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).). Regarding claim 10, the rejection of claim 7 is incorporated herein. Further, claim 10 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the values of the machine learning model are weights of the machine learning model (This limitation merely specifies the type of numerical data used in performing the recited abstract processing and does not impose a meaningful limitation.). Regarding claim 11, the rejection of claim 7 is incorporated herein. Further, claim 11 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the plurality of data structures are tensors (This limitation merely specifies the form of data structures used to implement the recited abstract processing and does not impose a meaningful limitation on the judicial exception.). Regarding claim 12, the rejection of claim 7 is incorporated herein. Further, claim 12 recites the following abstract ideas: wherein at least two of the plurality of data structures have different defined structured sparse patterns (This is an abstract idea of a mental process. The limitation involves observing and evaluating the arrangement of values in different data structures and determining that structures have different sparsity patterns. A person could compare the placement or number of zero values in the respective structures and determine that the patterns differ using pen and paper or basic computational tools.). Regarding claim 13, the rejection of claim 12 is incorporated herein. Further, claim 13 recites the following abstract ideas: wherein the different defined structured sparse patterns include at least: a first defined structured sparse pattern having a first sparsity degree, and a second defined structured sparse pattern having a second sparsity degree, wherein the first sparsity degree is different from the second sparsity degree (This is an abstract idea of a mental process. The limitation involves observing and evaluating the degree of sparsity associated with two different patterns and determining that the respective sparsity degrees are different. A person could compare the number or proportion of zero values associated with each pattern using pen and paper or basic computational tools. The type of observation, evaluation, and comparison can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 14, the rejection of claim 7 is incorporated herein. Further, claim 14 recites the following abstract ideas: wherein the plurality of different subsets of values of the machine learning model include: at least one subset comprised of at least a portion of inlier values of the machine learning model, and at least another subset comprised of at least a portion of outlier values of the machine learning model (This is an abstract idea of a mental process. The limitation involves evaluating model values, such as model weights, according to a criterion and exercising judgement to classify the values into respective inlier and outlier subsets. A person could review a written list of model weights, compare the weights to the applicable criterion, and manually separate the weights into inlier and outlier groups using pen and paper or basic computational tools. This type of evaluation, classification, and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: the machine learning model (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) Regarding claim 15, the rejection of claim 14 is incorporated herein. Further, claim 15 recites the following abstract ideas: wherein the inlier values and the outlier values are determined according to a defined threshold metric (This is an abstract idea of a mental process. The limitation involves evaluating model values, such as model weights, against a defined threshold metric and exercising judgement to determine whether each value is an inlier or an outlier. A person could compare each weight to the applicable threshold and classify the weight accordingly using pen and paper or basic computational tools.). Regarding claim 16, the rejection of claim 15 is incorporated herein. Further, claim 16 recites the following abstract ideas: wherein the defined threshold metric is a magnitude of weight (This is an abstract idea of a mental process. The limitation involves evaluating a model weight based on its magnitude. A person could determine the magnitude of the weight, for example by taking the absolute value of the weight, and use that value for comparison to the threshold using pen and paper or basic computational tools. This type of evaluation and comparison can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 17, the rejection of claim 15 is incorporated herein. Further, claim 17 recites the following abstract ideas: wherein the defined threshold metric is an error after quantization for inlier and outlier (This is an abstract idea of a mental process. The limitation involves evaluating the difference between an original value and its quantized value and using the resulting error to determine whether the value is an inlier or an outlier. A person could compare the original and quantized values, calculate the difference, and evaluate the error using pen and paper or basic computational tools. This type of evaluation and comparison can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 18, the rejection of claim 15 is incorporated herein. Further, claim 18 recites the following abstract ideas: wherein the defined threshold metric is a product of a corresponding weight and activation (This is an abstract idea of a mental process. The limitation involves evaluating numerical model values by multiplying a weight by its corresponding activation and using the result value as a metric. A person could perform the multiplication and evaluate the resulting product using pen and paper or basic computational tools.). Regarding claim 19, the rejection of claim 14 is incorporated herein. Further, claim 19 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein at least a portion of the inlier values are stored with a first structured sparse pattern that has less sparsity than a second structured sparse pattern used to store at least a portion of the outlier values (This limitation merely recites storing information in particular data structures, which amounts to generic data storage activity that is well-understood, routine, and conventional.). Regarding claim 20, the rejection of claim 17 is incorporated herein. Further, claim 20 recites the following abstract ideas: sparsifying the machine learning model by pruning values from the machine learning model, to form a sparse machine learning model, wherein the plurality of different subsets of values of the machine learning model are determined from the sparse machine learning model (This is an abstract idea of a mental process. The limitation involves evaluating model values, exercising judgement to select values for pruning, and organizing the remaining or zeroed values into different subsets. A person could review a written list or table of model weights, cross out or replace selected weights with zero to represent the sparse model, and then separate the resulting values into different groups using pen and paper or basic computational tools. This type of evaluation, selection, organization, and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.) Regarding claim 21, the rejection of claim 20 is incorporated herein. Further, claim 21 recites the following abstract ideas: wherein the values of the machine learning model are selected for pruning according to a defined threshold metric (This is an abstract idea of a mental process. The limitation involves evaluating model values against a defined threshold metric and exercising judgement to determine which values should be pruned. A person could review a written list of model weights, compare each weight to the applicable threshold metric, and identify values to be removed or replaced with zero using pen and paper or basic computational tools.). Regarding claim 22, the rejection of claim 21 is incorporated herein. Further, claim 22 recites the following abstract ideas: wherein the defined threshold metric is a magnitude of weight (This is an abstract idea of a mental process. The limitation involves evaluating a model weight based on its magnitude. A person could determine the magnitude of a weight, such as by considering its absolute value, and compare that magnitude to the applicable threshold using pen and paper or basic computational tools. This type of evaluation and comparison can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 23, the rejection of claim 21 is incorporated herein. Further, claim 23 recites the following abstract ideas: wherein the defined threshold metric is an error after pruning (This is an abstract idea of a mental process. The limitation involves evaluating the difference or error resulting after selected model values are pruned. A person could compare the model values or resulting output before and after pruning and determine a resulting error using pen and paper or basic computational tools. This type of evaluation and comparison can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 24, the rejection of claim 21 is incorporated herein. Further, claim 24 recites the following abstract ideas: wherein the defined threshold metric is a product of a corresponding weight and activation obtained with training or validation data (This is an abstract idea of a mental process. The limitation involves evaluating numerical values by multiplying a model weight by a corresponding activation and using the resulting product as the threshold metric. A person could obtain the corresponding activation from training or validation data, perform the multiplication, and evaluate the resulting product using pen and paper or basic computational tools. This type of calculation and evaluation can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.). Regarding claim 25, the rejection of claim 20 is incorporated herein. Further, claim 25 recites the following abstract ideas: wherein the machine learning model is sparsified to a defined degree of sparsity (This is an abstract idea of a mental process. The limitation involves evaluating model values and exercising judgement to select or designate values as zero until a defined degree of sparsity is reached. A person could review a written list of model weights, mark selected weights as zero, and count or calculate the resulting proportion of zero values using pen and paper or basic computational tools.). Regarding claim 26, the rejection of claim 20 is incorporated herein. Further, claim 26 recites the following abstract ideas: wherein the machine learning model is sparsified with a defined structured sparse pattern (This is an abstract idea of a mental process. The limitation involves evaluating model values and exercising judgement to select or designate values as zero until a defined degree of sparsity is reached. A person could review a written list of model weights, mark selected weights as zero, and count or calculate the resulting proportion of zero values using pen and paper or basic computational tools.). Regarding claim 27, the rejection of claim 7 is incorporated herein. Further, claim 27 recites the following abstract ideas: wherein changing the data representation of the at least one data structure includes quantizing the at least one data structure (This is an abstract idea of a mental process. The limitation involves evaluating numerical values and converting the values to a reduced set of representable values, such as by rounding. A person could review the numerical values and manually round or otherwise convert the values using pen and paper or basic computational tools.). Regarding claim 28, the rejection of claim 7 is incorporated herein. Further, claim 28 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the data representation of the plurality of data structures includes a bit width of the plurality of data structures (This limitation merely specifies the format used to represent the data in performing the abstract idea and does not impose a meaningful limitation on the judicial exception.). Regarding claim 29, the rejection of claim 7 is incorporated herein. Further, claim 29 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the data representation of the plurality of data structures includes a data type of the plurality of data structures (This limitation merely specifies the format used to represent the data in performing the recited abstract processing and does impose a meaningful limitation on the judicial exception.). Regarding claim 30, the rejection of claim 7 is incorporated herein. Further, claim 30 recites the following additional elements, which taken alone or in combination with other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the different data representations are supported by a single hardware accelerator or multiple different hardware accelerators (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).). Regarding claim 31, the rejection of claim 7 is incorporated herein. Further, claim 31 recites the following abstract ideas: wherein at least two of the plurality of data structures have different defined structured sparse patterns (This is an abstract idea of a mental process. The limitation involves evaluating and comparing the arrangement of values in different data structure and determining that the structures have different sparsity patterns. A person could compare the number or placement of zero values in the respective structure using pen and paper or basic computational tools. This type of evaluation and comparison can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.), The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: wherein the different defined structured sparse patterns are supported by a single hardware accelerator or multiple different hardware accelerators (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).). Regarding claim 32, the following claim elements are abstract ideas: apportioning a plurality of different subsets of values of the machine learning model into a plurality of data structures at least one of which has a defined structured sparse pattern (This is an abstract idea of a mental process. The limitation involves observing and evaluating numerical values and exercising judgement to separate values into different groups and organize at least one group according to a defined sparsity pattern. A person could review the values, sort them into separate lists or tables, and arrange selected entries according a particular pattern of zero values using pen and paper or basic computational tools. The type of observation, evaluation, organization, and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.); and changing a data representation of at least one data structure of the plurality of data structures, wherein at least two data structures of the plurality of data structures have different data representations (This is an abstract idea of a mental process. The limitation involves evaluating how numerical values are represented and exercising judgement to change the representation of at least one group of values so that the groups have different representations. A person could review the values and manually represent one group differently from another, such as by using a different format or level of detail, with the aid of pen and paper or basic computational tools.). The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: the machine learning model (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) a non-transitory memory storage (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) one or more processors (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) Regarding claim 33, the rejection of claim 32 is incorporated herein. The claim recites similar limitations corresponding to claim 8. Therefore, the same subject matter analysis that was utilized for claim 8, as described above, is equally applicable to claim 33. Therefore, claim 33 is ineligible. Regarding claim 34, the rejection of claim 32 is incorporated herein. The claim recites similar limitations corresponding to claim 9. Therefore, the same subject matter analysis that was utilized for claim 9, as described above, is equally applicable to claim 34. Therefore, claim 34 is ineligible. Regarding claim 35, the rejection of claim 32 is incorporated herein. The claim recites similar limitations corresponding to claim 10. Therefore, the same subject matter analysis that was utilized for claim 10, as described above, is equally applicable to claim 35. Therefore, claim 35 is ineligible. Regarding claim 36, the rejection of claim 32 is incorporated herein. The claim recites similar limitations corresponding to claim 12. Therefore, the same subject matter analysis that was utilized for claim 12, as described above, is equally applicable to claim 36. Therefore, claim 36 is ineligible. Regarding claim 37, the rejection of claim 32 is incorporated herein. The claim recites similar limitations corresponding to claim 14. Therefore, the same subject matter analysis that was utilized for claim 14, as described above, is equally applicable to claim 37. Therefore, claim 37 is ineligible. Regarding claim 38, the rejection of claim 37 is incorporated herein. The claim recites similar limitations corresponding to claim 19. Therefore, the same subject matter analysis that was utilized for claim 19, as described above, is equally applicable to claim 38. Therefore, claim 38 is ineligible. Regarding claim 39, the rejection of claim 32 is incorporated herein. The claim recites similar limitations corresponding to claim 27. Therefore, the same subject matter analysis that was utilized for claim 27, as described above, is equally applicable to claim 39. Therefore, claim 39 is ineligible. Regarding claim 40, the rejection of claim 32 is incorporated herein. The claim recites similar limitations corresponding to claim 28. Therefore, the same subject matter analysis that was utilized for claim 28, as described above, is equally applicable to claim 40. Therefore, claim 40 is ineligible. Regarding claim 41, the rejection of claim 32 is incorporated herein. The claim recites similar limitations corresponding to claim 29. Therefore, the same subject matter analysis that was utilized for claim 29, as described above, is equally applicable to claim 41. Therefore, claim 41 is ineligible. Regarding claim 42, the following claim elements are abstract ideas: apportioning a plurality of different subsets of values of the machine learning model into a plurality of data structures at least one of which has a defined structured sparse pattern (This is an abstract idea of a mental process. The limitation involves observing and evaluating numerical values and exercising judgement to separate values into different groups and organize at least one group according to a defined sparsity pattern. A person could review the values, sort them into separate lists or tables, and arrange selected entries according a particular pattern of zero values using pen and paper or basic computational tools. The type of observation, evaluation, organization, and judgement can be practically performed in the human mind and therefore falls within the mental process grouping of abstract ideas.); and changing a data representation of at least one data structure of the plurality of data structures, wherein at least two data structures of the plurality of data structures have different data representations (This is an abstract idea of a mental process. The limitation involves evaluating how numerical values are represented and exercising judgement to change the representation of at least one group of values so that the groups have different representations. A person could review the values and manually represent one group differently from another, such as by using a different format or level of detail, with the aid of pen and paper or basic computational tools.). The following claim elements are additional elements which, taken alone or in combination with the other elements, do not integrate the judicial exception into a practical application nor amount to significantly more than the judicial exception: the machine learning model (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) non-transitory computer-readable media (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) one or more processors (This a high-level recitation of generic computer components for performing the abstract idea. See MPEP 2106.05(f).) Regarding claim 43, the rejection of claim 42 is incorporated herein. The claim recites similar limitations corresponding to claim 8. Therefore, the same subject matter analysis that was utilized for claim 8, as described above, is equally applicable to claim 43. Therefore, claim 43 is ineligible. Regarding claim 44, the rejection of claim 42 is incorporated herein. The claim recites similar limitations corresponding to claim 9. Therefore, the same subject matter analysis that was utilized for claim 9, as described above, is equally applicable to claim 44. Therefore, claim 44 is ineligible. Regarding claim 45, the rejection of claim 42 is incorporated herein. The claim recites similar limitations corresponding to claim 10. Therefore, the same subject matter analysis that was utilized for claim 10, as described above, is equally applicable to claim 45. Therefore, claim 45 is ineligible. Regarding claim 46, the rejection of claim 42 is incorporated herein. The claim recites similar limitations corresponding to claim 12. Therefore, the same subject matter analysis that was utilized for claim 12, as described above, is equally applicable to claim 46. Therefore, claim 46 is ineligible. Regarding claim 47, the rejection of claim 42 is incorporated herein. The claim recites similar limitations corresponding to claim 14. Therefore, the same subject matter analysis that was utilized for claim 14, as described above, is equally applicable to claim 47. Therefore, claim 47 is ineligible. Regarding claim 48, the rejection of claim 47 is incorporated herein. The claim recites similar limitations corresponding to claim 19. Therefore, the same subject matter analysis that was utilized for claim 19, as described above, is equally applicable to claim 48. Therefore, claim 48 is ineligible. Regarding claim 49, the rejection of claim 42 is incorporated herein. The claim recites similar limitations corresponding to claim 27. Therefore, the same subject matter analysis that was utilized for claim 27, as described above, is equally applicable to claim 49. Therefore, claim 49 is ineligible. Regarding claim 50, the rejection of claim 42 is incorporated herein. The claim recites similar limitations corresponding to claim 28. Therefore, the same subject matter analysis that was utilized for claim 28, as described above, is equally applicable to claim 50. Therefore, claim 50 is ineligible. Regarding claim 51, the rejection of claim 42 is incorporated herein. The claim recites similar limitations corresponding to claim 29. Therefore, the same subject matter analysis that was utilized for claim 29, as described above, is equally applicable to claim 51. Therefore, claim 51 is ineligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, and 7-17, 19-20, and 25-51 are rejected under the 35 U.S.C. 103 as being unpatentable over Dettmers et al., (NPL: SpQR: A Sparse-Quantized Representation for Near-LossLess LLM Weight Compression” (Published: June 2023)) in view of Boo et al.,(NPL: “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations (Published: 2017)). Regarding claim 1, Dettmers teaches the following limitations: A method, comprising: at a device, compressing a machine learning model having a plurality of values to reduce at least one of a size of the machine learning model or computation requirements of the machine learning model, by: processing the machine learning model to generate a plurality of sparse data structures including: storing inlier values of the machine learning model in a first data structure (Dettmers, [Abstract] ”By compressing such LLMs via quantization to 3-4 bits per parameter, they can fit into memory-limited devices such as laptops and mobile phones…To address this accuracy issue, we introduce the Sparse-Quantized Representation (SpQR), a new compressed format and quantization technique which enables for the first time near-lossless compression of LLMs across model scales, while reaching similar compression levels to previous methods. SpQR works by identifying and isolating outlier weights, which cause particularly large quantization errors, and storing them in higher precision, while compressing all other weights to 3-4 bits, and achieves relative accuracy losses of less than 1% in perplexity for highly-accurate LLaMA and Falcon LLMs.” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions. Overall, the representation consists of (1) quantized weights, (2) first level quantized quantization statistics, second level quantization statistics, and (3) the CSR outlier indices and values.” & “Storing quantized groups. All non-outlier weights are encoded as a structure that contains: a b w -bit individual weight;” – Dettmer teaches compressing an LLM having a plurality of weight values to 3-4 bits per parameter so that the compressed model can fit within memory-limited devices such as laptops and mobile phones. Dettmers identifies SpQR as a Sparse-Quantized Representation and processes the model by identifying and isolating outlier weights while compressing the remaining non-outlier weights. Dettmers further converts the weights in several data structures of different sizes and precisions and encodes the non-outlier weights in a structure. Under the broadest reasonable interpretation, the non-outlier weights correspond to the claimed inlier values, and the structure encoding the non-outlier weights corresponds to the claimed first data structure.), and storing outlier values of the machine learning model in a second data structure (Dettmers, [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions. Overall, the representation consists of (1) quantized weights, (2) first level quantized quantization statistics, second level quantization statistics, and (3) the CSR outlier indices and values.” & “Storing outliers. Recall that our outliers are unstructured; for storage, we sort them by their row first and column second, so that outliers in the same row are contiguous in memory. For each outlier, we store two scalars: the 16-bit weight value and the 16-bit column index.” – teaches storing outlier weight values and corresponding indices in a CSR outlier representation. Under BRI, the outlier weights correspond to the claimed outlier values, and the CSR outlier representation corresponds to the claimed second data structure.), non-uniformly quantizing the machine learning model, including: quantizing the first data structure storing the inlier values to a first bit width (Dettmers, [page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure…(2) quantize the non-outlier “base” weights into 3-4 bit and transfer the remaining quantization into the the 16-bit outliers weights.” [section 4.2] “Storing quantized groups. All non-outlier weights are encoded as a structure that contains: a b w -bit individual weight…As a particular example for a SpQR representation, consider bw=bq=3 and Bw = Bq = 16. The weight matrix is split into groups of Bq × Bw = 256 weights. A group contains 256 individual bw = 3-bit codes.” – Dettmers teaches quantizing the non-outlier base weights to a 3-4 bit representation and encoding those non-outlier weights in a structure containing b w -bit individual weight, including an example using 3-bit codes. Under BRI, the non-outlier weights correspond to the claimed inlier values, the structure encoding those weights correspond to the claimed inlier values, the structure encoding those weights correspond to the first data structure, and the 3-4 bit representation corresponds to the claimed first bit width.);, and quantizing the second data structure storing the outlier values to a second bit width that is different from the first bit width (Dettmers, [page 6] “Since these weights appear to lead to high, irreducible error, we choose to keep these outliers in high precision (16-bit).” & “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit and transfer the remaining quantization into the the 16-bit outliers weights.” [section 4.2] “Storing outliers. Recall that our outliers are unstructured; for storage, we sort them by their row first and column second, so that outliers in the same row are contiguous in memory. For each outlier, we store two scalars: the 16-bit weight value and the 16-bit column index.” – Dettmers teaches storing the outlier weight values in the outlier representation at 16-bit precision, while the non-outlier base weights are quantized to 3-4 bits. Under BRI, the outlier representation corresponds to the claimed second data structure, the 16-bit representation corresponds to the claimed second bit width, and that second bit width is different from the 3-4 bit first bit width used for the non-outlier weights.). However, Dettmers does not teach but Dettmers in view of Boo teaches the following limitation: wherein at least one of the first data structure or the second data structure has a structured sparse pattern (Boo, [page 2, section II B] “The proposed structured ternary quantization divides a weight into many one-dimensional sub-vectors with the size of N, and each sub-vector is represented by a ternary vector with a limited number, K, of +1 or -1.” & “For example, the (4; 1) structured sparse coding denotes that the sub-vector length is 4 and only one position is allowed to be +1 or -1.” – Boo teaches a structured sparse pattern where weights are divided into sub-vectors of a defined size N, with each sub-vector limited to a defined number K of non-zero values. Boo further provides a structured sparse coding example in which each sub-vector has a length of four and only one position is permitted to contain a non-zero value. Under BRI, the defined pattern of sparsity corresponds to the claimed structured sparse pattern.); and Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Dettmers and Boo before them, to incorporate the structured sparse coding of Boo into the quantized non-outlier weight structure of Dettmers. One would have been motivated to make such a combination in order to further reduced the amount of storage required for the quantized non-outlier weights while providing a representation having low decoding overhead. This would allow the compressed machine learning model to require less memory for storing the non-outlier weights while preserving the higher-precision representation of outlier weights to reduce quantization error. Regarding claim 2, Dettmers in view of Boo teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Dettmer in view of Boo further teaches: wherein the inlier values and the outlier values are weights of the machine learning model (Dettmers, [page 6] “Since these weights appear to lead to high, irreducible error, we choose to keep these outliers in high precision (16-bit).” & “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit and transfer the remaining quantization into the the 16-bit outliers weights” – Dettmers teaches separating the machine learning model weights into outlier weights and non-outlier base weights. Under BRI, the non-outlier base weights correspond to the claimed inlier values and the outlier weights correspond to the claimed outlier values. Thus, both the inlier values and the outlier values are weights of the machine learning model.). Regarding claim 3, Dettmers in view of Boo teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Dettmer in view of Boo further teaches: wherein the inlier values and the outlier values are determined according to a defined threshold metric (Dettmers, [page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit and transfer the remaining quantization into the the 16-bit outliers weights. For the outlier isolation step, the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights. Globally, for each matrix, the algorithm aims to pick a sensitivity threshold τ to obtain the desired number of outliers across the whole model, usually around 1% of weights. Specifically, a particular weight is considered an outlier if keeping the weight in 16-bit reduces the error in Eq. (2) by at least τ .” – Dettmers teaches isolating and separating outlier weights from non-outlier base weights using a sensitivity criterion having a defined threshold τ . Under BRI, the non-outlier base weights correspond to the claimed inlier values and the outlier weights correspond to the claimed outlier values, which are therefore determined according to the defined threshold metric.). Regarding claim 4, Dettmers in view of Boo teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Dettmer in view of Boo further teaches: wherein the first data structure has a first structured sparse pattern that has less sparsity than a second structured sparse pattern of the second data structure (Boo, [page 2, section II, B] “The proposed structured ternary quantization divides a weight into many one-dimensional sub-vectors with the size of N, and each sub-vector is represented by a ternary vector with a limited number, K, of +1 or -1. We first prune each sub-vector of the floating-point weight matrix to have only K non-zero values.” [page2, table 1] Boo discloses structured sparce patterns including (16,4), (16,3) and (16,2). [page 4, section II, C] “At the first iteration, the network is retrained according to the proposed algorithm with a low sparse (N,K), which means a large K… At the next iteration, instead of W(q), W is pruned with higher sparsity by decrementing K.” – Boo teaches structured sparse patterns defined by (N,K), where K specifies the number of non-zero values permitted in a sub-vector of length N, and expressly teaches that a larger K corresponds to lower sparsity while decreasing K results in higher sparsity. Thus, a first structured sparse pattern (16,4) has less sparsity than a second structured sparse pattern of (16,2), since the first permits four non-zero values within the 16-value sub-vector while the second permits only two non-zero values.). Regarding claim 5, Dettmers in view of Boo teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. Dettmer in view of Boo further teaches: wherein the first data structure has a first structured sparse pattern that is the same as a second structured sparse pattern of the second data structure (Boo, [page 2, section II, B] “The proposed structured ternary quantization divides a weight into many one-dimensional sub-vectors with the size of N, and each sub-vector is represented by a ternary vector with a limited number, K, of +1 or -1.” [page 4, section IV, B] “The performances of the networks are shown in TABLE III. (N,K) is (16; 4) for the structured sparse networks.” – Boo teaches a structured sparse pattern defined by the same (N,K) configuration across the structured sparse network, including use of (16,4). Thus, the same structured sparse pattern can be used for both the first and second data structures.). Regarding claim 7, Dettmers teaches the following limitation: changing a data representation of at least one data structure of the plurality of data structures, wherein at least two data structures of the plurality of data structures have different data representations (Dettmers, [page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions.” & “For each outlier, we store two scalars: the 16-bit weight value and the 16-bit column index.” – Dettmers changes the data representation of the non-outlier weight data structure by quantizing the non-outlier weights to a 3-4 bit representation, while the outlier data structure stores the outlier weights in a different 16-bit representation. Thus, at least two of the data structures have different data representations.). However, Dettmers does not teach but Dettmers in view of Boo teaches the following limitation: A method, comprising: at a device: apportioning a plurality of different subsets of values of the machine learning model into a plurality of data structures at least one of which has a defined structured sparse pattern (Dettmers, [section 4.1] “Existing LLM quantization algorithms “ & “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions. Overall, the representation consists of (1) quantized weights, (2) first level quantized quantization statistics, second level quantization statistics, and (3) the CSR outlier indices and values.” & “Inference with SpQR. To illustrate the practicality of our approach, we design an efficient GPU-based decoding implementation for the SpQR format” Boo, [page 2, section II, B] “The proposed structured ternary quantization divides a weight into many one-dimensional sub-vectors with the size of N, and each sub-vector is represented by a ternary vector with a limited number, K, of +1 or -1. We first prune each sub-vector of the floating-point weight matrix to have only K non-zero values… For example, the (4, 1) structured sparse coding denotes that the sub-vector length is 4 and only one position is allowed to be +1 or -1.” – Dettmers implements SpQR on a GPU device and apportions different subsets of model weights, including outlier weights and non-outlier base weights, into different data structures. Boo provides a defined structured sparse pattern for at least one such data structure through its (N,K) structured sparsity configuration.); and Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Dettmers and Boo before them, to incorporate the structured sparse coding of Boo into the quantized non-outlier weight structure of Dettmers. One would have been motivated to make such a combination in order to further reduced the amount of storage required for the quantized non-outlier weights while providing a representation having low decoding overhead. This would allow the compressed machine learning model to require less memory for storing the non-outlier weights while preserving the higher-precision representation of outlier weights to reduce quantization error. Regarding claim 8, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the machine learning model is a deep neural network (Boo, [Abstract] “We propose a weight compression method for deep neural networks, which allows values of +1 or -1 only at predetermined positions of the weights so that decoding using a table can be conducted easily.”). Regarding claim 9, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the machine learning model is a large language model (LLM) (Dettmers, [Abstract] “To address this accuracy issue, we introduce the Sparse-Quantized Representation (SpQR), a new compressed format and quantization technique which enables for the first time near-lossless compression of LLMs across model scales, while reaching similar compression levels to previous methods.”). Regarding claim 10, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the values of the machine learning model are weights of the machine learning model (Dettmers, [section 4.1] “Existing LLM quantization algorithms…“ &“The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4” – Dettmers teaches that the values being processed are weights of the machine learning model.). Regarding claim 11, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the plurality of data structures are tensors (Dettmers, [page 7, Fig. 3, Algorithm 1] “Figure 3: A high-level overview of the SpQR representation for a single weight tensor. The right side of the image depicts all stored data types and their dimensions.” & “Q := int_matrix(m, n) // quantized weight… Wsparse = gather_outlier_matrix(W,O)” – Dettmers represents the machine learning weights as a weight tensor and generates separate matrix data structures for the quantized weights and sparse outlier weights. Under BRI, these matrix data structures are two-dimensional tensors.). Regarding claim 12, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein at least two of the plurality of data structures have different defined structured sparse patterns (Boo, [page 5] “TABLE IV shows the performance and the weight storage compression ratio of networks with various choices of (N,K).” [Table IV] “N is sub-vector length, and K is the number of non-zeros in a sub-vector.” & “(16,4)… (16,3)… (16,2)… (8,1)” – As previously mapped in claim 7, Dettmers provides the plurality of data structures containing different subsets of model weights. Boo defines structured sparse patterns according to an (N,K) configuration and discloses different structured sparse patterns, including (16,4) and (16,2). Applying these different structured sparse patterns to respective data structures results in at least two of the plurality of data structures having different structured sparse patterns.) Regarding claim 13, Dettmers in view of Boo teaches all the elements of claim 12, therefore is rejected for the same reasons as those presented for claim 12. Dettmer in view of Boo further teaches: wherein the different defined structured sparse patterns include at least: a first defined structured sparse pattern having a first sparsity degree, and a second defined structured sparse pattern having a second sparsity degree, wherein the first sparsity degree is different from the second sparsity degree (Boo, [page 4, section III, C] “At the first iteration, the network is retrained according to the proposed algorithm with a low sparse (N,K), which means a large K. Floating-point weights W and fixed-point weights W(q) are both retrained at this iteration. At the next iteration, instead of W(q), W is pruned with higher sparsity by decrementing K.” [page 5, Table IV] “N is sub-vector length, and K is the number of non-zeros in a sub-vector.” & “(16,4)… (16,3)… (16,2)” – Boo teaches that K represents the number of non-zero values in a sub-vector, with a larger K corresponding to lower sparsity and a smaller K corresponding to the higher sparsity. Thus, a first defined structured sparse pattern such as (16,4) has a first degree of sparsity, while a second defined structured sparse pattern such as (16,2) has a different, greater sparsity degree.). Regarding claim 14, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the plurality of different subsets of values of the machine learning model include: at least one subset comprised of at least a portion of inlier values of the machine learning model, and at least another subset comprised of at least a portion of outlier values of the machine learning model (Dettmers, [section 4.1, page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4…the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights.” – Dettmers separates the model weights into a subset of non-outlier base weights and a subset of outlier weights. Under BRI, the non-outlier base weights correspond to the claimed inlier values, while the isolated outlier weights correspond to the claimed outlier values.). Regarding claim 15, Dettmers in view of Boo teaches all the elements of claim 14, therefore is rejected for the same reasons as those presented for claim 14. Dettmer in view of Boo further teaches: wherein the inlier values and the outlier values are determined according to a defined threshold metric (Dettmers, [section 4.1, page 6] “For the outlier isolation step, the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights. Globally, for each matrix, the algorithm aims to pick a sensitivity threshold τ to obtain the desired number of outliers across the whole model, usually around 1% of weights. Specifically, a particular weight is considered an outlier if keeping the weight in 16-bit reduces the error in Eq. (2) by at least τ .” – Dettmers determines the outlier weights using a defined sensitivity threshold τ and separates those outlier weights from the remaining base weights. Under BRI, the remaining base weights are the claimed inlier values, such that both the inlier and the outlier values are determined according to the defined threshold metric.). Regarding claim 16, Dettmers in view of Boo teaches all the elements of claim 15, therefore is rejected for the same reasons as those presented for claim 15. Dettmer in view of Boo further teaches: wherein the defined threshold metric is a magnitude of weight (Dettmers, [page 6] “the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights.” Boo, [page 3, section III, A] “Each sub-vector of the floating-point weight matrix is pruned in order of magnitude so that every sub-vector only has K non-zero elements.” – As mapped in claim 15, Dettmers separates outlier values from non-outlier base values using a defined threshold metric. Boo teaches evaluating and selecting weight values according to weight magnitude. Applying Boo’s magnitude-based selection to Dettmer’s threshold determination results in the defined threshold metric being a magnitude of weight.). Regarding claim 17, Dettmers in view of Boo teaches all the elements of claim 15, therefore is rejected for the same reasons as those presented for claim 15. Dettmer in view of Boo further teaches: wherein the defined threshold metric is an error after quantization for inlier and outlier (Dettmers, [page 6] “For the outlier isolation step, the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights. Globally, for each matrix, the algorithm aims to pick a sensitivity threshold τ to obtain the desired number of outliers across the whole model, usually around 1% of weights. Specifically, a particular weight is considered an outlier if keeping the weight in 16-bit reduces the error in Eq. (2) by at least τ .” – Dettmers determines whether a weight is an outlier according to the reduction in quantization error obtained when the weight is retained in 16-bit rather than treated with the quantized base weights. Thus, the defined threshold metric τ is applied to an error resulting from the alternative inlier and outlier quantization treatment.). Regarding claim 19, Dettmers in view of Boo teaches all the elements of claim 14, therefore is rejected for the same reasons as those presented for claim 14. Dettmer in view of Boo further teaches: wherein at least a portion of the inlier values are stored with a first structured sparse pattern that has less sparsity than a second structured sparse pattern used to store at least a portion of the outlier values (Boo, [page 4, section III, C] “At the first iteration, the network is retrained according to the proposed algorithm with a low sparse (N,K), which means a large K… At the next iteration, instead of W(q), W is pruned with higher sparsity by decrementing K.” [Table IV] “N is sub-vector length, and K is the number of non-zeros in a sub-vector.” Boo provides, for examples, both (16,4) and (16,2) for structured sparse patterns. As previously mapped in claim 14, Dettmers provides a subset of non-outlier base values corresponding to the claimed inlier values and a separate subset of outlier values. Boo teaches structured sparse patterns having different degrees of sparsity, when a larger K provides less sparsity and a smaller K provides greater sparsity. Applying a lower-sparsity structured sparse pattern to at least a portion of the inlier values and a higher-sparsity structured sparse pattern to at least a portion of the outlier values meets the claimed relationship.). Regarding claim 20, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the machine learning model is further compressed by: sparsifying the machine learning model by pruning values from the machine learning model, to form a sparse machine learning model, wherein the plurality of different subsets of values of the machine learning model are determined from the sparse machine learning model (Boo, [page 2, section II, B] “We first prune each sub-vector of the floating-point weight matrix to have only K non-zero values. Then, quantization and retraining are performed with pruned weights kept to zero.” [page 3, section III, A] “Each sub-vector of the floating-point weight matrix is pruned in order of magnitude so that every sub-vector only has K non-zero elements… The quantization step size ∆ is calculated using the pruned weight matrix instead of the original one.” Dettmers, [page 6] “For the outlier isolation step, the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights.” – Boo teaches sparsifying the machine learning model by pruning weight values to produce a pruned weight matrix having only K non-zero values in each sub-vector, and then performing subsequent processing using that pruned weight matrix. Dettmers determines different subsets of model weights by separating outlier weights from non-outlier base weights. Applying Dettmers’ subset determination to Boo’s pruned weight matrix results in the plurality of different subsets being determined from the sparse machine learning model.). Regarding claim 25, Dettmers in view of Boo teaches all the elements of claim 20, therefore is rejected for the same reasons as those presented for claim 20. Dettmer in view of Boo further teaches: wherein the machine learning model is sparsified to a defined degree of sparsity (Boo, [page 4, section III, C] “W is pruned with higher sparsity by decrementing K… After the network is pruned for the target sparsity, the final W(q) is used for the inference.” [page 2] “We first prune each sub-vector of the floating-point weight matrix to have only K non-zero values.” – Boo teaches sparsifying the machine learning model by pruning weights until a target sparsity is reached, with K defining the number of non-zero values retained in each sub-vector. Thus, the model is sparsified to a defined degree of sparsity.). Regarding claim 26, Dettmers in view of Boo teaches all the elements of claim 20, therefore is rejected for the same reasons as those presented for claim 20. Dettmer in view of Boo further teaches: wherein the machine learning model is sparsified with a defined structured sparse pattern (Boo, [page 2] “The proposed structured ternary quantization divides a weight into many one-dimensional sub-vectors with the size of N, and each sub-vector is represented by a ternary vector with a limited number, K, of +1 or -1… For example, the (4, 1) structured sparse coding denotes that the sub-vector length is 4 and only one position is allowed to be +1 or -1.” – Boo teaches sparsifying the machine learning model according to an expressly defined structured sparse pattern, where an (N,K) pattern specifies a sub-vector of N weights having only K permitted non-zero values, such as the disclosed (4,1) structured sparse pattern.). Regarding claim 27, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein changing the data representation of the at least one data structure includes quantizing the at least one data structure (Dettmers, [page 4, section 3.1] “Not all parameters in a neural network are equally important. Intuitively, a weight could be seen as sensitive to quantization if its rounding error is large, i.e. it is not close to a quantization point, and/or the inputs it is usually multiplied with a large, amplifying even a small rounding error…We define the sensitivity sij of some weight wij in the layer’s weight matrix W as the minimum squared difference between the original predictions on X and those of any weight matrix W’ where this weight is quantized…” – As previously mapped in claim 7, Dettmers changes the data representation of at least one data structure containing values of the machine learning model. Dettmers further teaches quantizing a weight within the data structure. Thus, changing the data representation of the at least one data structure includes quantizing the at least one data structure.) Regarding claim 28, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the data representation of the plurality of data structures includes a bit width of the plurality of data structures (Dettmers, [page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” – As previously mapped in claim 7, Dettmers apportions different subsets of model weights into different data structures having different data representations. Dettmers further teaches representing the outlier weights at 16-bit and the non-outlier base weights at 3-4 bit. Thus, the data representation of the plurality of data structures includes a bit width.). Regarding claim 29, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the data representation of the plurality of data structures includes a data type of the plurality of data structures (Dettmers, [page 7, Fig. 3] “Figure 3: A high-level overview of the SpQR representation for a single weight tensor. The right side of the image depicts all stored data types and their dimensions.” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions.” – As previously mapped in claim 7, Dettmers represents different subsets of model weights using a plurality of data structures. Dettmers further identifies the representations of those structures as including stored data types. Thus, the data representation of the plurality of data structures includes a data type of the plurality of data structures.). Regarding claim 30, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein the different data representations are supported by a single hardware accelerator or multiple different hardware accelerators (Dettmers, [page 8] “Inference with SpQR. To illustrate the practicality of our approach, we design an efficient GPU-based decoding implementation for the SpQR format… At a high level, our algorithm loads group statistics and the quantized weights into shared memory (SRAM), dequantizes to 16-bits, and then performs matrix multiplication with 16-bit inputs. For handling outliers, we design a sparse matrix algorithm that takes advantage of outliers that occur in rows.” – As previously mapped in claim 7, Dettmers provides different data representations for the different weight data structures. Dettmers further teaches a GPU-based implementation that processes both the quantized weights and the separately represented outliers. Thus, the different data representations are supported by a single hardware accelerator, namely the GPU.). Regarding claim 31, Dettmers in view of Boo teaches all the elements of claim 7, therefore is rejected for the same reasons as those presented for claim 7. Dettmer in view of Boo further teaches: wherein at least two of the plurality of data structures have different defined structured sparse patterns, and wherein the different defined structured sparse patterns are supported by a single hardware accelerator or multiple different hardware accelerators (Boo, [page 2] “The proposed structured ternary quantization divides a weight into many one-dimensional sub-vectors with the size of N, and each sub-vector is represented by a ternary vector with a limited number, K, of +1 or -1. We first prune each sub-vector of the floating-point weight matrix to have only K non-zero values… The process is described in Fig. 1 when (N,K) is (8; 4)… For example, the (4, 1) structured sparse coding denotes that the sub-vector length is 4 and only one position is allowed to be +1 or -1.” Dettmers, [page 8] “Inference with SpQR. To illustrate the practicality of our approach, we design an efficient GPU-based decoding implementation for the SpQR format…” – Boo teaches defined structured sparse patterns in which N defines the sub-vector size and K limits the number of non-zero values, and discloses different patterns including (8, 4) and (4,1). Dettmers teaches processing compressed machine learning weight structures using a GPU-based implementation. Applying Boo’s different structured sparse patterns to the weight data structures processed by Dettmers results in at least two data structures having different defined structured sparse patterns supported by a single hardware accelerator, namely the GPU.). Regarding claim 32, Dettmers teaches the following limitation: A system, comprising: a non-transitory memory storage comprising instructions; and one or more processors in communication with the memory, wherein the one or more processors execute the instructions to at least one of compress a machine learning model or reduce a computation of the machine learning model by (Dettmers, [page 2] “To convert a given pretrained LLM into SpQR format, we adopt an extended version of the post-training quantization (PTQ) approach recently introduced by GPTQ [FAHA22]. Specifically, the method passes calibration data through the uncompressed model; to compress each layer, it applies a layer-wise solver with respect to the L2 error between the outputs of the uncompressed model, and those of the quantized weights.” [page 8] “Then, each GPU core (thread block) (2) loads a large slice of outliers into shared memory (SRAM), and each GPU core (3) determines if outliers are part of the segment or not. The corresponding weights are (4) loaded from main memory; finally, the matrix multiplication is performed.”) changing a data representation of at least one data structure of the plurality of data structures, wherein at least two data structures of the plurality of data structures have different data representations (Dettmers, [page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions.” & “For each outlier, we store two scalars: the 16-bit weight value and the 16-bit column index.” – Dettmers changes the data representation of the non-outlier weight data structure by quantizing the non-outlier weights to a 3-4 bit representation, while the outlier data structure stores the outlier weights in a different 16-bit representation. Thus, at least two of the data structures have different data representations.). However, Dettmers does not teach but Dettmers in view of Boo teaches the following limitation: apportioning a plurality of different subsets of values of the machine learning model into a plurality of data structures at least one of which has a defined structured sparse pattern (Dettmers, [section 4.1] “Existing LLM quantization algorithms “ & “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions. Overall, the representation consists of (1) quantized weights, (2) first level quantized quantization statistics, second level quantization statistics, and (3) the CSR outlier indices and values.” & “Inference with SpQR. To illustrate the practicality of our approach, we design an efficient GPU-based decoding implementation for the SpQR format” Boo, [page 2, section II, B] “The proposed structured ternary quantization divides a weight into many one-dimensional sub-vectors with the size of N, and each sub-vector is represented by a ternary vector with a limited number, K, of +1 or -1. We first prune each sub-vector of the floating-point weight matrix to have only K non-zero values… For example, the (4, 1) structured sparse coding denotes that the sub-vector length is 4 and only one position is allowed to be +1 or -1.” – Dettmers implements SpQR on a GPU device and apportions different subsets of model weights, including outlier weights and non-outlier base weights, into different data structures. Boo provides a defined structured sparse pattern for at least one such data structure through its (N,K) structured sparsity configuration.); and Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Dettmers and Boo before them, to incorporate the structured sparse coding of Boo into the quantized non-outlier weight structure of Dettmers. One would have been motivated to make such a combination in order to further reduced the amount of storage required for the quantized non-outlier weights while providing a representation having low decoding overhead. This would allow the compressed machine learning model to require less memory for storing the non-outlier weights while preserving the higher-precision representation of outlier weights to reduce quantization error. Regarding claim 33, Dettmers in view of Boo teaches all the elements of claim 32, therefore is rejected for the same reasons as those presented for claim 32. Dettmer in view of Boo further teaches: wherein the machine learning model is a deep neural network (Boo, [Abstract] “We propose a weight compression method for deep neural networks, which allows values of +1 or -1 only at predetermined positions of the weights so that decoding using a table can be conducted easily.”). Regarding claim 34, Dettmers in view of Boo teaches all the elements of claim 32, therefore is rejected for the same reasons as those presented for claim 32. Dettmer in view of Boo further teaches: wherein the machine learning model is a large language model (LLM) (Dettmers, [Abstract] “To address this accuracy issue, we introduce the Sparse-Quantized Representation (SpQR), a new compressed format and quantization technique which enables for the first time near-lossless compression of LLMs across model scales, while reaching similar compression levels to previous methods.”). Regarding claim 35, Dettmers in view of Boo teaches all the elements of claim 32, therefore is rejected for the same reasons as those presented for claim 32. Dettmer in view of Boo further teaches: wherein the values of the machine learning model are weights of the machine learning model (Dettmers, [section 4.1] “Existing LLM quantization algorithms…“ &“The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4” – Dettmers teaches that the values being processed are weights of the machine learning model.). Regarding claim 36, Dettmers in view of Boo teaches all the elements of claim 32, therefore is rejected for the same reasons as those presented for claim 32. Dettmer in view of Boo further teaches: wherein at least two of the plurality of data structures have different defined structured sparse patterns (Boo, [page 5] “TABLE IV shows the performance and the weight storage compression ratio of networks with various choices of (N,K).” [Table IV] “N is sub-vector length, and K is the number of non-zeros in a sub-vector.” & “(16,4)… (16,3)… (16,2)… (8,1)” – As previously mapped in claim 32, Dettmers provides the plurality of data structures containing different subsets of model weights. Boo defines structured sparse patterns according to an (N,K) configuration and discloses different structured sparse patterns, including (16,4) and (16,2). Applying these different structured sparse patterns to respective data structures results in at least two of the plurality of data structures having different structured sparse patterns.) Regarding claim 37, Dettmers in view of Boo teaches all the elements of claim 32, therefore is rejected for the same reasons as those presented for claim 32. Dettmer in view of Boo further teaches: wherein the plurality of different subsets of values of the machine learning model include: at least one subset comprised of at least a portion of inlier values of the machine learning model, and at least another subset comprised of at least a portion of outlier values of the machine learning model (Dettmers, [section 4.1, page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4…the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights.” – Dettmers separates the model weights into a subset of non-outlier base weights and a subset of outlier weights. Under BRI, the non-outlier base weights correspond to the claimed inlier values, while the isolated outlier weights correspond to the claimed outlier values.). Regarding claim 38, Dettmers in view of Boo teaches all the elements of claim 37, therefore is rejected for the same reasons as those presented for claim 37. Dettmer in view of Boo further teaches: wherein at least a portion of the inlier values are stored with a first structured sparse pattern that has less sparsity than a second structured sparse pattern used to store at least a portion of the outlier values (Boo, [page 4, section III, C] “At the first iteration, the network is retrained according to the proposed algorithm with a low sparse (N,K), which means a large K… At the next iteration, instead of W(q), W is pruned with higher sparsity by decrementing K.” [Table IV] “N is sub-vector length, and K is the number of non-zeros in a sub-vector.” Boo provides, for examples, both (16,4) and (16,2) for structured sparse patterns. As previously mapped in claim 37, Dettmers provides a subset of non-outlier base values corresponding to the claimed inlier values and a separate subset of outlier values. Boo teaches structured sparse patterns having different degrees of sparsity, when a larger K provides less sparsity and a smaller K provides greater sparsity. Applying a lower-sparsity structured sparse pattern to at least a portion of the inlier values and a higher-sparsity structured sparse pattern to at least a portion of the outlier values meets the claimed relationship.). Regarding claim 39, Dettmers in view of Boo teaches all the elements of claim 32, therefore is rejected for the same reasons as those presented for claim 32. Dettmer in view of Boo further teaches: wherein changing the data representation of the at least one data structure includes quantizing the at least one data structure (Dettmers, [page 4, section 3.1] “Not all parameters in a neural network are equally important. Intuitively, a weight could be seen as sensitive to quantization if its rounding error is large, i.e. it is not close to a quantization point, and/or the inputs it is usually multiplied with a large, amplifying even a small rounding error…We define the sensitivity sij of some weight wij in the layer’s weight matrix W as the minimum squared difference between the original predictions on X and those of any weight matrix W’ where this weight is quantized…” – As previously mapped in claim 32, Dettmers changes the data representation of at least one data structure containing values of the machine learning model. Dettmers further teaches quantizing a weight within the data structure. Thus, changing the data representation of the at least one data structure includes quantizing the at least one data structure.) Regarding claim 40, Dettmers in view of Boo teaches all the elements of claim 32, therefore is rejected for the same reasons as those presented for claim 32. Dettmer in view of Boo further teaches: wherein the data representation of the plurality of data structures includes a bit width of the plurality of data structures (Dettmers, [page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” – As previously mapped in claim 32, Dettmers apportions different subsets of model weights into different data structures having different data representations. Dettmers further teaches representing the outlier weights at 16-bit and the non-outlier base weights at 3-4 bit. Thus, the data representation of the plurality of data structures includes a bit width.). Regarding claim 41, Dettmers in view of Boo teaches all the elements of claim 32, therefore is rejected for the same reasons as those presented for claim 32. Dettmer in view of Boo further teaches: wherein the data representation of the plurality of data structures includes a data type of the plurality of data structures (Dettmers, [page 7, Fig. 3] “Figure 3: A high-level overview of the SpQR representation for a single weight tensor. The right side of the image depicts all stored data types and their dimensions.” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions.” – As previously mapped in claim 32, Dettmers represents different subsets of model weights using a plurality of data structures. Dettmers further identifies the representations of those structures as including stored data types. Thus, the data representation of the plurality of data structures includes a data type of the plurality of data structures.). Regarding claim 42, Dettmers teaches the following limitation: A non-transitory computer-readable media storing computer instructions which when executed by one or more processors of a device cause the device to at least one of compress a machine learning model or reduce a computation of the machine learning model by: (Dettmers, [page 2] “To convert a given pretrained LLM into SpQR format, we adopt an extended version of the post-training quantization (PTQ) approach recently introduced by GPTQ [FAHA22]. Specifically, the method passes calibration data through the uncompressed model; to compress each layer, it applies a layer-wise solver with respect to the L2 error between the outputs of the uncompressed model, and those of the quantized weights.” [page 8] “Then, each GPU core (thread block) (2) loads a large slice of outliers into shared memory (SRAM), and each GPU core (3) determines if outliers are part of the segment or not. The corresponding weights are (4) loaded from main memory; finally, the matrix multiplication is performed.”) changing a data representation of at least one data structure of the plurality of data structures, wherein at least two data structures of the plurality of data structures have different data representations (Dettmers, [page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions.” & “For each outlier, we store two scalars: the 16-bit weight value and the 16-bit column index.” – Dettmers changes the data representation of the non-outlier weight data structure by quantizing the non-outlier weights to a 3-4 bit representation, while the outlier data structure stores the outlier weights in a different 16-bit representation. Thus, at least two of the data structures have different data representations.). However, Dettmers does not teach but Dettmers in view of Boo teaches the following limitation: apportioning a plurality of different subsets of values of the machine learning model into a plurality of data structures at least one of which has a defined structured sparse pattern (Dettmers, [section 4.1] “Existing LLM quantization algorithms “ & “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions. Overall, the representation consists of (1) quantized weights, (2) first level quantized quantization statistics, second level quantization statistics, and (3) the CSR outlier indices and values.” & “Inference with SpQR. To illustrate the practicality of our approach, we design an efficient GPU-based decoding implementation for the SpQR format” Boo, [page 2, section II, B] “The proposed structured ternary quantization divides a weight into many one-dimensional sub-vectors with the size of N, and each sub-vector is represented by a ternary vector with a limited number, K, of +1 or -1. We first prune each sub-vector of the floating-point weight matrix to have only K non-zero values… For example, the (4, 1) structured sparse coding denotes that the sub-vector length is 4 and only one position is allowed to be +1 or -1.” – Dettmers implements SpQR on a GPU device and apportions different subsets of model weights, including outlier weights and non-outlier base weights, into different data structures. Boo provides a defined structured sparse pattern for at least one such data structure through its (N,K) structured sparsity configuration.); and Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, having a combination of Dettmers and Boo before them, to incorporate the structured sparse coding of Boo into the quantized non-outlier weight structure of Dettmers. One would have been motivated to make such a combination in order to further reduced the amount of storage required for the quantized non-outlier weights while providing a representation having low decoding overhead. This would allow the compressed machine learning model to require less memory for storing the non-outlier weights while preserving the higher-precision representation of outlier weights to reduce quantization error. Regarding claim 43, Dettmers in view of Boo teaches all the elements of claim 42, therefore is rejected for the same reasons as those presented for claim 42. Dettmer in view of Boo further teaches: wherein the machine learning model is a deep neural network (Boo, [Abstract] “We propose a weight compression method for deep neural networks, which allows values of +1 or -1 only at predetermined positions of the weights so that decoding using a table can be conducted easily.”). Regarding claim 44, Dettmers in view of Boo teaches all the elements of claim 42, therefore is rejected for the same reasons as those presented for claim 42. Dettmer in view of Boo further teaches: wherein the machine learning model is a large language model (LLM) (Dettmers, [Abstract] “To address this accuracy issue, we introduce the Sparse-Quantized Representation (SpQR), a new compressed format and quantization technique which enables for the first time near-lossless compression of LLMs across model scales, while reaching similar compression levels to previous methods.”). Regarding claim 45, Dettmers in view of Boo teaches all the elements of claim 42, therefore is rejected for the same reasons as those presented for claim 22. Dettmer in view of Boo further teaches: wherein the values of the machine learning model are weights of the machine learning model (Dettmers, [section 4.1] “Existing LLM quantization algorithms…“ &“The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4” – Dettmers teaches that the values being processed are weights of the machine learning model.). Regarding claim 46, Dettmers in view of Boo teaches all the elements of claim 42, therefore is rejected for the same reasons as those presented for claim 42. Dettmer in view of Boo further teaches: wherein at least two of the plurality of data structures have different defined structured sparse patterns (Boo, [page 5] “TABLE IV shows the performance and the weight storage compression ratio of networks with various choices of (N,K).” [Table IV] “N is sub-vector length, and K is the number of non-zeros in a sub-vector.” & “(16,4)… (16,3)… (16,2)… (8,1)” – As previously mapped in claim 42, Dettmers provides the plurality of data structures containing different subsets of model weights. Boo defines structured sparse patterns according to an (N,K) configuration and discloses different structured sparse patterns, including (16,4) and (16,2). Applying these different structured sparse patterns to respective data structures results in at least two of the plurality of data structures having different structured sparse patterns.) Regarding claim 47, Dettmers in view of Boo teaches all the elements of claim 42, therefore is rejected for the same reasons as those presented for claim 42. Dettmer in view of Boo further teaches: wherein the plurality of different subsets of values of the machine learning model include: at least one subset comprised of at least a portion of inlier values of the machine learning model, and at least another subset comprised of at least a portion of outlier values of the machine learning model (Dettmers, [section 4.1, page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4…the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights.” – Dettmers separates the model weights into a subset of non-outlier base weights and a subset of outlier weights. Under BRI, the non-outlier base weights correspond to the claimed inlier values, while the isolated outlier weights correspond to the claimed outlier values.). Regarding claim 48, Dettmers in view of Boo teaches all the elements of claim 47, therefore is rejected for the same reasons as those presented for claim 47. Dettmer in view of Boo further teaches: wherein at least a portion of the inlier values are stored with a first structured sparse pattern that has less sparsity than a second structured sparse pattern used to store at least a portion of the outlier values (Boo, [page 4, section III, C] “At the first iteration, the network is retrained according to the proposed algorithm with a low sparse (N,K), which means a large K… At the next iteration, instead of W(q), W is pruned with higher sparsity by decrementing K.” [Table IV] “N is sub-vector length, and K is the number of non-zeros in a sub-vector.” Boo provides, for examples, both (16,4) and (16,2) for structured sparse patterns. As previously mapped in claim 42, Dettmers provides a subset of non-outlier base values corresponding to the claimed inlier values and a separate subset of outlier values. Boo teaches structured sparse patterns having different degrees of sparsity, when a larger K provides less sparsity and a smaller K provides greater sparsity. Applying a lower-sparsity structured sparse pattern to at least a portion of the inlier values and a higher-sparsity structured sparse pattern to at least a portion of the outlier values meets the claimed relationship.). Regarding claim 49, Dettmers in view of Boo teaches all the elements of claim 42, therefore is rejected for the same reasons as those presented for claim 42. Dettmer in view of Boo further teaches: wherein changing the data representation of the at least one data structure includes quantizing the at least one data structure (Dettmers, [page 4, section 3.1] “Not all parameters in a neural network are equally important. Intuitively, a weight could be seen as sensitive to quantization if its rounding error is large, i.e. it is not close to a quantization point, and/or the inputs it is usually multiplied with a large, amplifying even a small rounding error…We define the sensitivity sij of some weight wij in the layer’s weight matrix W as the minimum squared difference between the original predictions on X and those of any weight matrix W’ where this weight is quantized…” – As previously mapped in claim 42, Dettmers changes the data representation of at least one data structure containing values of the machine learning model. Dettmers further teaches quantizing a weight within the data structure. Thus, changing the data representation of the at least one data structure includes quantizing the at least one data structure.) Regarding claim 50, Dettmers in view of Boo teaches all the elements of claim 42, therefore is rejected for the same reasons as those presented for claim 42. Dettmer in view of Boo further teaches: wherein the data representation of the plurality of data structures includes a bit width of the plurality of data structures (Dettmers, [page 6] “The procedure for detecting the outliers is described in detail in Alg. 1. If follows a rough two-step procedure: (1) find and isolate outliers as 16-bit weights, (2) quantize the non-outlier “base” weights into 3-4 bit” – As previously mapped in claim 42, Dettmers apportions different subsets of model weights into different data structures having different data representations. Dettmers further teaches representing the outlier weights at 16-bit and the non-outlier base weights at 3-4 bit. Thus, the data representation of the plurality of data structures includes a bit width.). Regarding claim 51, Dettmers in view of Boo teaches all the elements of claim 42, therefore is rejected for the same reasons as those presented for claim 42. Dettmer in view of Boo further teaches: wherein the data representation of the plurality of data structures includes a data type of the plurality of data structures (Dettmers, [page 7, Fig. 3] “Figure 3: A high-level overview of the SpQR representation for a single weight tensor. The right side of the image depicts all stored data types and their dimensions.” [section 4.2] “Our algorithm converts homogeneous weights into several data structures of various sizes and precisions.” – As previously mapped in claim 42, Dettmers represents different subsets of model weights using a plurality of data structures. Dettmers further identifies the representations of those structures as including stored data types. Thus, the data representation of the plurality of data structures includes a data type of the plurality of data structures.). Claim 6 is rejected under the 35 U.S.C. 103 as being unpatentable over Dettmers et al., (NPL: SpQR: A Sparse-Quantized Representation for Near-LossLess LLM Weight Compression” (Published: June 2023)) in view of Boo et al.,(NPL: “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations (Published: 2017)) further in view of Guo et al., (NPL: “OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization” (Published: April 2023). Regarding claim 6, Dettmers in view of Boo teaches all the elements of claim 1, therefore is rejected for the same reasons as those presented for claim 1. However, Dettmers in view of Boo does not teach but Dettmers in view of Boo further in view of Guo teaches: wherein the first bit width and the second bit width are supported by different hardware accelerators (Guo, [Abstract] “This enables a memory-aligned OVP encoding scheme, which can be efficiently integrated to the existing hardware accelerators like systolic array and tensor core. “[page 8, section 4.5] “For the systolic array, our architecture naturally supports 8-bit computation with four 4-bit PEs [72].” & [page 8, section 4.6] “For 4-bit tensor cores, the Turing GPU architecture adopts the instruction mma.s32.s4.s4.s32”) – Guo identifies the systolic array and tensor core as hardware accelerators, with the systolic array supporting 8-bit computation and the tensor core supporting 4-bit.). Accordingly, it would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the combination of Dettmers and Boo to support different bit widths using the respective hardware accelerators taught by Guo in order to reduce hardware overhead and improve performance of processing the differently quantized values (Guo, [Abstract]). Claim 18, 21-24 is rejected under the 35 U.S.C. 103 as being unpatentable over Dettmers et al., (NPL: SpQR: A Sparse-Quantized Representation for Near-LossLess LLM Weight Compression” (Published: June 2023)) in view of Boo et al.,(NPL: “Structured Sparse Ternary Weight Coding of Deep Neural Networks for Efficient Hardware Implementations (Published: 2017)) further in view of Sun et al., (NPL: “A Simple and Effective Pruning Approach for Large Language Models” (Published: June 2023). Regarding claim 18, Dettmers in view of Boo teaches all the elements of claim 15, therefore is rejected for the same reasons as those presented for claim 15. However, Dettmers in view of Boo does not teach but Dettmers in view of Boo further in view of Sun teaches: wherein the defined threshold metric is a product of a corresponding weight and activation (Dettmers, [page 6] “the algorithm implements a filtering technique based on the sensitivity criterion in Eq. (2), which is used to isolate and separate outliers from base weights. Globally, for each matrix, the algorithm aims to pick a sensitivity threshold τ to obtain the desired number of outliers across the whole model” Sun, [page 3] “Consider a fully connected layer with weight W of shape   ( C o u t , C i n ) . For language models, this linear layer takes in input activations X … For each individual weight, we propose to evaluate its importance by the product of its magnitude and the corresponding input feature norm. Specifically, the score for the current weight W i j is defined by: S i j = W i j ⋅ X j 2 ” – As mapped in claim 15, Dettmers uses a defined metric to determine whether weights are outliers or non-outlier base weights. Sun teaches a metric for evaluating a weight that multiplies the magnitude of the corresponding weight W i j by magnitude of its corresponding input activation feature X j . Using Sun’s weight-and-activation metric for Dettmers’ weight determination therefore results in the defined threshold metric being a product of a corresponding weight and activation.). It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention, to further modify the combination of Dettmers and Boo to determine the threshold metric using Sun’s product of corresponding weight and input activation, in order to account for the influence of both the weight and its corresponding activation when evaluating the weight, thereby preserving weights that may be relatively low magnitude but a greater effect on the model output. Regarding claim 21, Dettmers in view of Boo teaches all the elements of claim 20, therefore is rejected for the same reasons as those presented for claim 20. However, Dettmer in view of Boo does not teach but Dettmers in view of Boo further in view of Sun teaches: wherein the defined threshold metric is a magnitude of weight (Sun, [page 2] “Magnitude Pruning [25, 26] is a standard pruning technique to induce sparsity in neural networks. Different from structured pruning approaches [40, 69], magnitude pruning removes individual weights based on their magnitudes, where weights with magnitudes below a certain threshold are removed. In practice, this threshold is typically determined by comparing weights globally [41] or layer-wise [77].” – As previously mapped in claim 20, the machine learning model is sparsified by pruning values for the model. Sun teaches selecting which weight values are pruned according to a defined threshold metric, where the weight values are evaluated according to magnitude and values having magnitudes below the determined threshold are selected for removal.). Regarding claim 22, Dettmers in view of Boo further in view of Sun teaches all the elements of claim 21, therefore is rejected for the same reasons as those presented for claim 21. Dettmers in view of Boo further in view of Sun teaches: wherein the defined threshold metric is a magnitude of weight (Sun, [page 2] “magnitude pruning removes individual weights based on their magnitudes, where weights with magnitudes below a certain threshold are removed.” – As previously mapped in claim 21, Sun selects values for pruning according to a defined threshold metric. Sun further teaches that the metric used for the threshold comparison is the magnitude of the weight, with weights having magnitudes below the threshold being selected for removal.). Regarding claim 23, Dettmers in view of Boo further in view of Sun teaches all the elements of claim 21, therefore is rejected for the same reasons as those presented for claim 21. Dettmers in view of Boo further in view of Sun teaches: wherein the defined threshold metric is an error after pruning (Sun, [page 5, section 4] “SparseGPT [20] is a second-order pruning method based on solving a layer-wise reconstruction problem. To scale existing second-order based approaches [21] to LLMs, an efficient weight update procedure was proposed that iterates between weight removal and weight update at each layer.” [page 6] “our pruning metric shares an implicit connection to the OBS reconstruction error in Equation 3” – As previously mapped in claim 21, Sun teaches selecting values for pruning according to a defined pruned metric. Sun further teaches a pruning approach that iterates between removal of a weight and updating the remaining weights, and identifies the pruning metric with an OBS reconstruction error. Thus, the metric evaluates the reconstruction error associated with the model after a weight has been removed, corresponding to the claimed error after pruning.). Regarding claim 24, Dettmers in view of Boo further in view of Sun teaches all the elements of claim 21, therefore is rejected for the same reasons as those presented for claim 21. Dettmers in view of Boo further in view of Sun teaches: wherein the defined threshold metric is a product of a corresponding weight and activation obtained with training or validation data (Sun, [page 2] “Specifically, we introduce a novel pruning metric, where each weight is evaluated by the product of its magnitude and the norm of the corresponding input activations, estimated using a small set of calibration data.” [page 5] “Both Wanda and SparseGPT require calibration data to estimate certain input statistics (see Table 1). To control this variable factor, we use the exact same set of calibration data as SparseGPT, which consists of 128 sequences (2048 tokens each) sampled from the first shard of the C4 training data [53].” – As previously mapped in claim 21, Sun teaches selecting values for pruning according to a defined threshold metric. Sun further teaches a defined pruning metric form from the product of a corresponding weight magnitude and its corresponding input activation, where the input-activation statistics are estimated using calibration samples obtained from C4 training data. Thus, the defined threshold metric is a product of a corresponding weight and activation obtained with training data.). Conclusion The prior art of record and not relied upon is considered pertinent to Applicant’s disclosure: 1. Aytekin, C., Cricri, F., & Aksu, E. (2019). Compressibility loss for neural network weights. arXiv preprint arXiv:1905.01044. – is considered pertinent to the claimed subject matter as it teaches compression of neural network weights through sparsification and quantization. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Daravanh Phakousonh whose telephone number is (571)272-6324. The examiner can normally be reached Mon - Thurs 7 AM - 5 PM, Every other Friday 7 AM - 4PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Li B Zhen can be reached at 571-272-3768. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Daravanh Phakousonh/Examiner, Art Unit 2121 /Li B. Zhen/Supervisory Patent Examiner, Art Unit 2121
Read full office action

Prosecution Timeline

Mar 12, 2024
Application Filed
Aug 24, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12572821
ACCURACY PRIOR AND DIVERSITY PRIOR BASED FUTURE PREDICTION
4y 0m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
25%
Grant Probability
99%
With Interview (+100.0%)
3y 3m (~8m remaining)
Median Time to Grant
Low
PTA Risk
Based on 4 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month