Prosecution Insights
Last updated: October 02, 2026
Application No. 17/650,440

EFFICIENT OPTIMIZATION OF MACHINE LEARNING MODELS

Final Rejection §101§103
Filed
Feb 09, 2022
Examiner
KIM, SEHWAN
Art Unit
2129
Tech Center
2100 — Computer Architecture & Software
Assignee
International Business Machines Corporation
OA Round
4 (Final)
61%
Grant Probability
Moderate
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 61% of resolved cases
61%
Career Allowance Rate
95 granted / 156 resolved
+5.9% vs TC avg
Strong +67% interview lift
Without
With
+67.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 0m
Avg Prosecution
32 currently pending
Career history
188
Total Applications
across all art units

Statute-Specific Performance

§101
20.3%
-19.7% vs TC avg
§103
46.5%
+6.5% vs TC avg
§102
7.7%
-32.3% vs TC avg
§112
23.3%
-16.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 156 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Examiner’s Note The Examiner encourages Applicant to schedule an interview (via Automated Interview Request (AIR)) to discuss issues related to, for example, the rejections noted below under 35 U.S.C § 101 and § 103, for moving toward allowance. (e.g., integrating dependent claims, elaborating claim languages with specificities regarding training) Providing supporting paragraph(s) for each limitation of amended/new claim(s) in Remarks is strongly requested for clear and definite claim interpretations by Examiner (e.g., to avoid rejections under 35 U.S.C § 112(a) “Lack of written description”) Applicant can schedule interviews (via Automated Interview Request (AIR)) at any stage of the prosecution (e.g., Non-Final, Final, and After-Final) to discuss any issues related to, for example, rejections under 35 U.S.C § 101 and § 102/103, for moving toward allowance. If a limitation has bold brackets (i.e. [·]) around claim languages, the bracketed claim languages indicate that they have not been taught yet by the current prior art reference but they will be taught by another prior art reference afterwards. If a limitation has one or more bold underlines, the one or more bold underlined claim languages indicate that they are taught by the current prior art reference, while the one or more non-underlined claim languages indicate that they have been taught already by one or more previous art references. Priority Acknowledgment is made of applicant's claim for the present application filed on 02/09/2022. Response to Arguments Applicant's arguments filed on 07/09/2026 have been fully considered but they are not persuasive. In Remarks, regarding 35 U.S.C. 101, Applicant contends: Applicant respectfully submits that one or more aspects of the above-quoted language of amended claim I do not recite an abstract idea that is a mental process. … One or more of these aspects cannot be performed by the human mind, and therefore, are not an abstract idea. … Applicant respectfully submits that the above-noted processes support a practical application, that is, improving the computer-implemented deployment and runtime use of a machine learning model by reducing model size, reducing memory/storage/bandwidth demands, and improving scoring/inference performance. The specification expressly links the recited optimizations to improved storage efficiency, transmission, reduced computational expense and reduced inference time. Accordingly, Applicant respectfully submits that amended claim 1 is patentable eligible under 35 U.S.C. §101, and respectfully requests withdrawal of the rejection. … Applicant respectfully submits that these aspects of amended claim 1 are a specific technological process that transforms a machine learning model into a smaller and more efficient deployable model. Examiner’s response: The examiner understands the applicant’s assertion. However, it appears that each processing step is just applying the abstract idea to a general field of endeavor with additional elements. In addition, improvements to technology or technical field are not necessarily reflected in the claims. Thus, the claim does not integrate the judicial exception into a practical application, and the claim does not amount to significantly more than the judicial exception. The examiner understands the applicant’s assertion regarding abstract ideas. Generating a second machine learning model may be interpreted as an additional element. However, as rejected under Claim Rejections - 35 USC § 101, it appears that the claim 1 still has mental processes except additional elements. Thus, it still appears that each processing step is just applying the abstract idea to a general field of endeavor with additional elements, without improvements to technology or technical field. The examiner understands the applicant’s assertion “Applicant respectfully submits that the above-noted processes support a practical application, that is, improving the computer-implemented deployment and runtime use of a machine learning model by reducing model size, reducing memory/storage/bandwidth demands, and improving scoring/inference performance. The specification expressly links the recited optimizations to improved storage efficiency, transmission, reduced computational expense and reduced inference time. Accordingly, Applicant respectfully submits that amended claim 1 is patentable eligible under 35 U.S.C. §101, and respectfully requests withdrawal of the rejection” and “Applicant respectfully submits that these aspects of amended claim 1 are a specific technological process that transforms a machine learning model into a smaller and more efficient deployable model.” However, it is not clear how the claims can reduce model size, reduce memory/storage/bandwidth demands, and improve scoring/inference performance. In addition, it is not clear how the claims reflect the asserted improvements of storage efficiency, transmission, reduced computational expense and reduced inference time, and a smaller and more efficient deployable model. The claims recite parsing, removing elements, encoding, generating a binary model, etc. However, it appears that they are kind of known techniques and it is not clear how/why the claims of the present invention provide the improvements. Currently, the claims do not clearly show e.g., improvements in computer technology and improvements to other technical fields. It doesn’t appear that the claims clearly show how the inventive concept of the claims enables the asserted improvements and how they are tied together. The applicant may need to amend the claims to show how the claim languages and improvements are tied together. To find a valid improvement to a technology, MPEP 2106.04(d)(1) says the specification must explain the improvement and that the claim must reflect the disclosed improvement. Furthermore, the improvement should not be merely a consequence of the abstract idea. See MPEP 2106.05(a). An improvement in the abstract idea itself is not an improvement to technology. For at least these reasons, Applicant's arguments are not convincing. The Examiner encourages Applicant to schedule an interview to discuss issues related to, for example, the rejections noted below under 35 U.S.C § 101. Applicant’s arguments regarding 35 USC § 103 with respect to the independent claims have been considered but are moot because the arguments are directed to amended limitation(s) that has/have not been previously examined. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The limitations of “… method, comprising: …; … comprises: parsing the first machine learning model to identify elements of the first machine learning model that are used during scoring of model input; removing one or more non-scoring elements from the first machine learning model, …; removing, from the first machine learning model, one or more of redundant elements and unused elements identified based on the parsing; encoding one or more string values of the first machine learning model using encoded values that are used in place of the one or more string values during computation; generating a converted model by converting the second machine learning model into a common intermediate format; generating a binary model by marshaling the converted model from the common intermediate format into a specified binary format representing the converted model; and facilitating use of the second machine learning model for runtime inferencing by …”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper). If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). In particular, the claim recites an additional element(s) (“A computer-implemented”) – using a device and/or a model to process data. The device and the model in each step are recited at a high-level of generality (i.e., as a generic computer performing a generic computer function of processing data) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. In particular, the claim recites an additional element(s) (“receiving a first machine learning model in a first format”, “transmitting the binary model to a remote system for deployment”) – the act of receiving/sending data. The claim is adding an insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). The act of receiving/sending data is recited at a high-level of generality (i.e., as a generic act of performing a generic act function of receiving/sending data) such that it amounts no more than a mere act to apply the exception using a generic act of receiving/sending. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. In particular, the claim recites an additional element(s) (“generating a second machine learning model by applying both a common optimization technique and a model-specific optimization technique to the first machine learning model, …, and wherein the common optimization technique applies to multiple model types, wherein the model specific optimization technique applies to a particular model type, and wherein generating the second machine learning model comprises”). The additional element is recited at such a high level without any details as to how a model is generated such that it amounts to only the idea of a solution or outcome because it fails to recite details of how a solution to a problem is accomplished, and, therefore, represents no more than mere instructions to apply the judicial exception on a computer (see MPEP 2106.05(f)). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. In particular, the claim recites an additional element (“wherein the second machine learning model is an optimized version of the first machine learning model”, “wherein the one or more non-scoring elements are not used when generating output of the first machine learning model”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, with respect to integration of the abstract idea into a practical application, the additional elements of using a generic computer component to perform each step amount to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible. MPEP 2106.05(f). As discussed above, the claim recites the additional element(s) of receiving/sending data at a high-level of generality and is adding an insignificant extra-solution activity – see MPEP 2106.05(g). However, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood, routine, and conventional. See MPEP 2106.05(d)(II) – “Receiving or transmitting data over a network” or “Storing and retrieving information in memory”. Accordingly, this additional element does not provide an inventive concept and significantly more than the abstract idea. Thus, the claim is not patent eligible. The additional elements regarding training are recited at such a high level without any details as to how a model is generated such that it amounts to only the idea of a solution or outcome because it fails to recite details of how a solution to a problem is accomplished, and, therefore, represents no more than mere instructions to apply the judicial exception on a computer (see MPEP 2106.05(f)). Accordingly, this additional element does not amount to significantly more than the abstract idea. The claim is directed to an abstract idea. This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h). Regarding claim 2 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The claim recites the abstract idea identified above regarding claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim recites an additional element (“wherein the first machine learning model is an ensemble model, and wherein the common optimization technique and the model-specific optimization technique further include cross-model optimization techniques, wherein the cross-model optimization techniques apply to ensemble model types”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h). Regarding claim 3 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The limitations of “wherein the common optimization technique comprises at least one of: removing redundant elements; removing unused elements; and encoding string values”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper). If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible. Regarding claim 4 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The limitations of “wherein removing the redundant elements comprises: parsing the first machine learning model; determining, based on the parsing, that a first element corresponds to a first input value; and upon determining, based on the parsing, that a second element also corresponds to the first input value, removing either the first element or the second element from the first machine learning model”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper). If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible. Regarding claim 5 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The limitations of “wherein a first machine learning model specific optimization technique applies to trees, and comprises at least one of: pruning useless nodes; or combining duplicated nodes”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper). If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible. Regarding claim 6 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The limitations of “wherein a first machine learning model specific optimization technique applies to regression models, and comprises removing nodes with a coefficient of zero”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper). If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible. Regarding claim 7 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The limitations of “wherein the cross-model optimization techniques comprise: identifying a plurality of base models included in the first machine learning model; clustering the plurality of base models based on a set of model features; and for each respective cluster, extracting a model structure that matches output of each base model in the respective cluster”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper). If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim does not recite additional elements. Thus, the claim is directed to an abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claim is not patent eligible. Regarding claim 8 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The claim recites the abstract idea identified above regarding claim 1. Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim recites an additional element (“wherein the converted model comprises metadata information and model content, and”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h) In particular, the claim recites an additional element(s) (“wherein the metadata information of the converted model is stored as tabular data”) – the act of storing data. The claim is adding an insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). The act of storing data is recited at a high-level of generality (i.e., as a generic act of storing performing a generic act function of storing data) such that it amounts no more than a mere act to apply the exception using a generic act of storing. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h). As discussed above, the claim recites the additional element(s) of storing data at a high-level of generality and is adding an insignificant extra-solution activity – see MPEP 2106.05(g) – storing data. However, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood, routine, and conventional. See MPEP 2106.05(d)(II) – “Receiving or transmitting data over a network” or “Storing and retrieving information in memory”. Accordingly, this additional element does not provide an inventive concept and significantly more than the abstract idea. Thus, the claim is not patent eligible. Regarding claim 9 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The limitations of “… by: converting the binary model to a third model in the common intermediate format; …”, as drafted, are a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. That is, nothing in the claim element precludes the step from practically being performed in the mind. For example, the limitations in the context of this claim encompass the user mentally thinking with a physical aid (e.g., pencil and paper). If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea. Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites additional elements that are mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea. See MPEP 2106.05(f). In particular, the claim recites an additional element(s) (“the remote system”) – using a device to process data. The device in each step is recited at a high-level of generality (i.e., as a generic computer performing a generic computer function of processing data) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. In particular, the claim recites an additional element(s) (“instantiates an optimized version of the first machine learning model”, “converting the third model to a fourth machine learning model in the first format; and instantiating the optimized version of the first machine learning model based on the fourth machine learning model in the first format”). The additional element is recited at such a high level without any details as to how a model is generated such that it amounts to only the idea of a solution or outcome because it fails to recite details of how a solution to a problem is accomplished, and, therefore, represents no more than mere instructions to apply the judicial exception on a computer (see MPEP 2106.05(f)). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above, with respect to integration of the abstract idea into a practical application, the additional elements of using a generic computer component to perform each step amount to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claim is not patent eligible. MPEP 2106.05(f). The additional elements regarding training are recited at such a high level without any details as to how a model is generated such that it amounts to only the idea of a solution or outcome because it fails to recite details of how a solution to a problem is accomplished, and, therefore, represents no more than mere instructions to apply the judicial exception on a computer (see MPEP 2106.05(f)). Accordingly, this additional element does not amount to significantly more than the abstract idea. The claim is directed to an abstract idea. Regarding claim 10 The claim recites “A system, comprising: one or more computer processors; and a memory containing a program which when executed by the one or more computer processors performs an operation, the operation comprising” to perform precisely the method of Claim 1. As performance of an abstract idea on generic computer components (see MPEP 2106.05(f)) cannot integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself, the claim is rejected for reasons set forth in the rejection of Claim 1. Regarding claim 11 The claim is rejected for the reasons set forth in the rejection of Claim 2 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception. Regarding claim 12 The claim is rejected for the reasons set forth in the rejection of Claim 3 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception. Regarding claim 13 The claim is rejected for the reasons set forth in the rejection of Claim 7 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception. Regarding claim 14 The claim is rejected for the reasons set forth in the rejection of Claim 8 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception. Regarding claim 15 The claim is rejected for the reasons set forth in the rejection of Claim 9 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception. Regarding claim 16 The claim recites “A computer program product comprising: a computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by one or more computer processors to:” to perform precisely the method of Claim 1. As performance of an abstract idea on generic computer components (see MPEP 2106.05(f)) cannot integrate the abstract idea into a practical application nor provide significantly more than the abstract idea itself, the claim is rejected for reasons set forth in the rejection of Claim 1. Regarding claim 17 The claim is rejected for the reasons set forth in the rejection of Claim 2 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception. Regarding claim 18 The claim is rejected for the reasons set forth in the rejection of Claim 3 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception. Regarding claim 19 The claim is rejected for the reasons set forth in the rejection of Claim 7 under 35 U.S.C. 101, mutatis mutandis, as reciting an abstract idea without integrating the judicial exception into a practical application nor providing significantly more than the judicial exception. Regarding claim 20 The claim is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: The claim recites a method; therefore, it falls into the statutory category of processes. Step 2A Prong 1: The claim recites the abstract idea identified above regarding claim 16. Step 2A Prong 2: This judicial exception is not integrated into a practical application. In particular, the claim recites an additional element (“wherein the converted model in the common intermediate format comprises metadata information and model content, and”). This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not integrate the abstract idea into a practical application. See MPEP 2106.05(h) In particular, the claim recites an additional element(s) (“wherein the metadata information of the converted model in the common intermediate format is stored by tabular data”) – the act of storing data. The claim is adding an insignificant extra-solution activity to the judicial exception – see MPEP 2106.05(g). The act of storing data is recited at a high-level of generality (i.e., as a generic act of storing performing a generic act function of storing data) such that it amounts no more than a mere act to apply the exception using a generic act of storing. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. This is a recitation of a particular type or source of model/data to be used in performing the abstract idea. Limiting the abstract idea to a particular type or source of model/data is an attempt to limit the abstract idea to a particular field of use or technological environment, which does not amount to significantly more than the abstract idea. See MPEP 2106.05(h). As discussed above, the claim recites the additional element(s) of storing data at a high-level of generality and is adding an insignificant extra-solution activity – see MPEP 2106.05(g) – storing data. However, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood, routine, and conventional. See MPEP 2106.05(d)(II) – “Receiving or transmitting data over a network” or “Storing and retrieving information in memory”. Accordingly, this additional element does not provide an inventive concept and significantly more than the abstract idea. Thus, the claim is not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 3, 10, 12, 16, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 11,301,762 B1) in view of Cummings et al. (US 20220036123 A1) Regarding claim 1 Chen teaches A computer-implemented method, comprising: receiving a first machine learning model in a first format; (Chen [col 1 ln 66– col 3 ln 34] “a unified inference framework for heterogenous edge devices is provided that can accept machine learning (ML) models from a user in any of multiple formats ( as generated by multiple different frameworks), convert and optimize these ML models for use by heterogeneous "edge" devices having heterogeneous computing resources, and deploy these ML models for use in one or more edge devices 10 of one or more different types.” [col 3 ln 43– col 7 ln 55] “FIG. 1 is a diagram illustrating an exemplary environment including a unified inference framework (“UIF”) server module 112 according to some embodiments. The UIF server module 112, in some embodiments, is a portion of software allowing users 118 to deploy and manage high performance machine learning models 130 running on connected devices 122 in production. Users 118 (e.g., individuals, organizations, even OEMs) can import to—or train machine learning models in—a provider network 100 (“the cloud”), and reliably deploy these models to large numbers of devices 122 at the edge. Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices.”;) generating a second machine learning model by applying both a common optimization technique and a model-specific optimization technique to the first machine learning model, wherein the second machine learning model is an optimized version of the first machine learning model, and (Chen [fig(s) 2] “ML Frameworks 202”, “Framework-specific models 204”, “Model Optimizer 114” [fig(s) 3] “OPTIONALLY TRANS., OPTIONALLY PARTIALLY OPTIMIZED MODEL 129” [col 3 ln 43– col 7 ln 55] “FIG. 1 is a diagram illustrating an exemplary environment including a unified inference framework (“UIF”) server module 112 according to some embodiments. The UIF server module 112, in some embodiments, is a portion of software allowing users 118 to deploy and manage high performance machine learning models 130 running on connected devices 122 in production. Users 118 (e.g., individuals, organizations, even OEMs) can import to—or train machine learning models in—a provider network 100 (“the cloud”), and reliably deploy these models to large numbers of devices 122 at the edge. Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices.” [col 6 ln 47– col 6 ln 52] “In some embodiments, the (possibly translated) model can be provided to the one or more edge devices 122A-122N (e.g., as indicated by the deployment request) directly (via 50 circle (3A)), or by storing the (at least partially) optimized model 129 ( e.g., a translated model, a translated and partially optimized model, an optimized model, etc.) in a storage service 116 at circle (3B), where it can be obtained by the one or more edge devices 122A-122N as shown at 55 circle (3C), e.g., via the one or more edge devices 122A122N sending requests ( e.g., web service requests) to download the optimized model 129 files.”;) and wherein the common optimization technique applies to multiple model types, wherein the model specific optimization technique applies to a particular model type, and wherein generating the second machine learning model comprises: (Chen [fig(s) 2] “ML Frameworks 202”, “Framework-specific models 204”, “Model Optimizer 114” [col 1 ln 66– col 3 ln 34] “a unified inference framework for heterogenous edge devices is provided that can accept machine 5 learning (ML) models from a user in any of multiple formats ( as generated by multiple different frameworks), convert and optimize these ML models for use by heterogeneous "edge" devices having heterogeneous computing resources, and deploy these ML models for use in one or more edge devices 10 of one or more different types.” [col 3 ln 43– col 7 ln 55] “FIG. 1 is a diagram illustrating an exemplary environment including a unified inference framework (“UIF”) server module 112 according to some embodiments. The UIF server module 112, in some embodiments, is a portion of software allowing users 118 to deploy and manage high performance machine learning models 130 running on connected devices 122 in production. Users 118 (e.g., individuals, organizations, even OEMs) can import to—or train machine learning models in—a provider network 100 (“the cloud”), and reliably deploy these models to large numbers of devices 122 at the edge. Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices.”; Examiner notes that paragraph 33 of the Instant Specification describes “At block 215, the type of the received model is determined. In some embodiments, determining the model type includes identifying the architecture of the model and/or the format used to represent the model (which can be used to identify which other optimization techniques can be applied). For example, some optimizations may be applicable to neural networks, while others are uniquely applicable to decision trees.”) encoding one or more string values of the first machine learning model using encoded values that are used in place of the one or more string values during computation; (Cummings [fig(s) 1-2] [par(s) 82-88] “Each neuron 410 has one or more inputs and produces an output, which can be sent to one or more other neurons 410 (the inputs and outputs may be referred to as “signals”). Inputs to the neurons 410 of the input layer Lx can be feature values of a sample of external data (e.g., input variables xi). The input variables xi can be set as a vector containing relevant data (e.g., observations, ML features, etc.). The inputs to hidden units 410 of the hidden layers La, Lb, and Lc may be based on the outputs of other neurons 410. The outputs of the final output neurons 410 of the output layer Ly (e.g., output variables yj) include predictions, inferences, and/or accomplish a desired/configured task. The output variables yj may be in the form of determinations, inferences, predictions, and/or assessments. Additionally or alternatively, the output variables yj can be set as a vector containing the relevant data (e.g., determinations, inferences, predictions, assessments, and/or the like). In the context of ML, an “ML feature” (or simply “feature”) is an individual measureable property or characteristic of a phenomenon being observed. Features are usually represented using numbers/numerals (e.g., integers), strings, variables, ordinals, real-values, categories, and/or the like. Additionally or alternatively, ML features are individual variables, which may be independent variables, based on observable phenomenon that can be quantified and recorded. ML models use one or more features to make predictions or inferences. In some implementations, new features can be derived from old features.”; e.g., “The input variables xi can be set as a vector containing relevant data (e.g., observations, ML features, etc.)” along with “Features are usually represented using … strings” read(s) on “encoding one or more string values” since the input layer does embedding for input features.) generating a converted model by converting the second machine learning model into a common intermediate format; (Chen [col 3 ln 43– col 7 ln 55] “Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices.” and “a translator 113A module (e.g., library, function, binary, etc.) of the model optimizer 114A translates the ML model 108 from a first format (as generated by a particular ML framework) into a “common” second format that is “unified” in that it can be run by all inference engines 132 of all edge devices 122. …Thus, by translating the ML model 108 into a common format, the model optimizer 114A can make the model “portable” in that it can be run at any/all of the one or more edge devices 122A-122N using a common inference engine 132 (e.g., via use of an inference library 134). This translation can be performed using tools and techniques known to those of skill in the art, which may include using a conversion library/module that identifies certain values (e.g., weights) in model files generated by a first framework and inserts them in a different format (or location) within files adherent to a different framework or format (e.g., a different framework's format, a standardized “generic” format such as the Open Neural Network eXchange format “ONNX”, etc.). As another example, the translator may create low-level machine code that is executable for multiple types of hardware backends.”;) generating a binary model by marshaling the converted model from the common intermediate format into a specified binary format representing the converted model; and (Chen [col 3 ln 43– col 7 ln 55] “Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices.” and “a translator 113A module (e.g., library, function, binary, etc.) of the model optimizer 114A translates the ML model 108 from a first format (as generated by a particular ML framework) into a “common” second format that is “unified” in that it can be run by all inference engines 132 of all edge devices 122. … Thus, by translating the ML model 108 into a common format, the model optimizer 114A can make the model “portable” in that it can be run at any/all of the one or more edge devices 122A-122N using a common inference engine 132 (e.g., via use of an inference library 134). This translation can be performed using tools and techniques known to those of skill in the art, which may include using a conversion library/module that identifies certain values (e.g., weights) in model files generated by a first framework and inserts them in a different format (or location) within files adherent to a different framework or format (e.g., a different framework's format, a standardized “generic” format such as the Open Neural Network eXchange format “ONNX”, etc.). As another example, the translator may create low-level machine code that is executable for multiple types of hardware backends.”; e.g., “translator 113A module (e.g., library, function, binary, etc.) of the model optimizer 114A translates the ML model 108 from a first format (as generated by a particular ML framework) into a “common” second format that is “unified” in that it can be run by all inference engines 132 of all edge devices 122” along with “the translator may create low-level machine code that is executable for multiple types of hardware backends” read(s) on “generating a binary model by converting the converted model from the common intermediate format into a specified binary format”.) facilitating use of the second machine learning model for runtime inferencing by transmitting the binary model to a remote system for deployment. (Chen [col 3 ln 43– col 7 ln 55] “Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices” and “a translator 113A module (e.g., library, function, binary, etc.) of the model optimizer 114A translates the ML model 108 from a first format (as generated by a particular ML framework) into a “common” second format that is “unified” in that it can be run by all inference engines 132 of all edge devices 122. … Thus, by translating the ML model 108 into a common format, the model optimizer 114A can make the model “portable” in that it can be run at any/all of the one or more edge devices 122A-122N using a common inference engine 132 (e.g., via use of an inference library 134). This translation can be performed using tools and techniques known to those of skill in the art, which may include using a conversion library/module that identifies certain values (e.g., weights) in model files generated by a first framework and inserts them in a different format (or location) within files adherent to a different framework or format (e.g., a different framework's format, a standardized “generic” format such as the Open Neural Network eXchange format “ONNX”, etc.). As another example, the translator may create low-level machine code that is executable for multiple types of hardware backends. … the (possibly translated) model can be provided to the one or more edge devices 122A-122N (e.g., as indicated by the deployment request) directly (via 50 circle (3A)), or by storing the (at least partially) optimized model 129 ( e.g., a translated model, a translated and partially optimized model, an optimized model, etc.) in a storage service 116 at circle (3B), where it can be obtained by the one or more edge devices 122A-122N as shown at 55 circle (3C), e.g., via the one or more edge devices 122A122N sending requests ( e.g., web service requests) to download the optimized model 129 files.” [col 5 ln 47– col 5 ln 63] “In this scenario, the user 118 may cause the client device 120 to send, at circle (2), a request to deploy the ML model(s) 108 to one or more edge devices 122A-122N.”;) However, Chen does not appear to explicitly teach: parsing the first machine learning model to identify elements of the first machine learning model that are used during scoring of model input; removing one or more non-scoring elements from the first machine learning model, wherein the one or more non-scoring elements are not used when generating output of the first machine learning model; removing, from the first machine learning model, one or more of redundant elements and unused elements identified based on the parsing; Cummings teaches parsing the first machine learning model to identify elements of the first machine learning model that are used during scoring of model input; (Cummings [fig(s) 1-2] [par(s) 53-65] “Each depth component in vector 200 represents a block in the CNN, where the value in each component represents a number of convolutional layers in each block, For example, a first block (B0) includes 2 layers, a second block (B1) includes 3 layers, a third block (B2) includes 2 layers, a fourth block (B3) includes 4 layers, and a fifth block (B4) includes 4 layers. Therefore, the set of values for the depth component of vector 200 includes the values 2, 3, or 4 (e.g., indicated by “Depth={2,3,4}” in FIG. 2). Each kernel size component indicates a kernel size of each layer. For example, a first layer in the first block (index 0) has a kernel size of 3, a second layer in the third block (index 9) has a kernel size of 5, a third layer in the fifth block (index 18) has a kernel size of 7, and so forth. Therefore, the set of values for the kernel size of vector 200 includes the values 3, 5, or 7 (e.g., indicated by “Kernel Size={3,5,7}” in FIG. 2). Additionally, each expansion ratio (width) component indicates an expansion ratio of each layer. For example, a first layer in the first block (index 0) has an expansion ratio (width) of 4, a second layer in the third block (index 9) has an expansion ratio (width) of 3, a fourth layer in the fifth block (index 18) has an expansion ratio (width) of 6, and so forth. Therefore, the set of values for the expansion ratio component of vector 200 includes the values 3, 4, or 6 (e.g., indicated by “Expansion Ratio (Width)={3,4,6}” in FIG. 2). In the example subnet configuration of FIG. 2, unused parameters (e.g., indicated by the value “X” in FIG. 2) are obtained when there are less than the maximum of 4 layers (e.g., 2 or 3 layers) in a block. These unused parameters do not significantly contribute to the prediction or inference determination and may be removed to reduce the overall size of the ML model.”;) removing one or more non-scoring elements from the first machine learning model, wherein the one or more non-scoring elements are not used when generating output of the first machine learning model; (Cummings [fig(s) 1-2] [par(s) 53-65] “Each depth component in vector 200 represents a block in the CNN, where the value in each component represents a number of convolutional layers in each block, For example, a first block (B0) includes 2 layers, a second block (B1) includes 3 layers, a third block (B2) includes 2 layers, a fourth block (B3) includes 4 layers, and a fifth block (B4) includes 4 layers. Therefore, the set of values for the depth component of vector 200 includes the values 2, 3, or 4 (e.g., indicated by “Depth={2,3,4}” in FIG. 2). Each kernel size component indicates a kernel size of each layer. For example, a first layer in the first block (index 0) has a kernel size of 3, a second layer in the third block (index 9) has a kernel size of 5, a third layer in the fifth block (index 18) has a kernel size of 7, and so forth. Therefore, the set of values for the kernel size of vector 200 includes the values 3, 5, or 7 (e.g., indicated by “Kernel Size={3,5,7}” in FIG. 2). Additionally, each expansion ratio (width) component indicates an expansion ratio of each layer. For example, a first layer in the first block (index 0) has an expansion ratio (width) of 4, a second layer in the third block (index 9) has an expansion ratio (width) of 3, a fourth layer in the fifth block (index 18) has an expansion ratio (width) of 6, and so forth. Therefore, the set of values for the expansion ratio component of vector 200 includes the values 3, 4, or 6 (e.g., indicated by “Expansion Ratio (Width)={3,4,6}” in FIG. 2). In the example subnet configuration of FIG. 2, unused parameters (e.g., indicated by the value “X” in FIG. 2) are obtained when there are less than the maximum of 4 layers (e.g., 2 or 3 layers) in a block. These unused parameters do not significantly contribute to the prediction or inference determination and may be removed to reduce the overall size of the ML model.”;) removing, from the first machine learning model, one or more of redundant elements and unused elements identified based on the parsing; (Cummings [fig(s) 1-2] [par(s) 53-65] “Each depth component in vector 200 represents a block in the CNN, where the value in each component represents a number of convolutional layers in each block, For example, a first block (B0) includes 2 layers, a second block (B1) includes 3 layers, a third block (B2) includes 2 layers, a fourth block (B3) includes 4 layers, and a fifth block (B4) includes 4 layers. Therefore, the set of values for the depth component of vector 200 includes the values 2, 3, or 4 (e.g., indicated by “Depth={2,3,4}” in FIG. 2). Each kernel size component indicates a kernel size of each layer. For example, a first layer in the first block (index 0) has a kernel size of 3, a second layer in the third block (index 9) has a kernel size of 5, a third layer in the fifth block (index 18) has a kernel size of 7, and so forth. Therefore, the set of values for the kernel size of vector 200 includes the values 3, 5, or 7 (e.g., indicated by “Kernel Size={3,5,7}” in FIG. 2). Additionally, each expansion ratio (width) component indicates an expansion ratio of each layer. For example, a first layer in the first block (index 0) has an expansion ratio (width) of 4, a second layer in the third block (index 9) has an expansion ratio (width) of 3, a fourth layer in the fifth block (index 18) has an expansion ratio (width) of 6, and so forth. Therefore, the set of values for the expansion ratio component of vector 200 includes the values 3, 4, or 6 (e.g., indicated by “Expansion Ratio (Width)={3,4,6}” in FIG. 2). In the example subnet configuration of FIG. 2, unused parameters (e.g., indicated by the value “X” in FIG. 2) are obtained when there are less than the maximum of 4 layers (e.g., 2 or 3 layers) in a block. These unused parameters do not significantly contribute to the prediction or inference determination and may be removed to reduce the overall size of the ML model.”;) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Chen with the element removal of Cummings. One of ordinary skill in the art would have been motived to combine in order to reduce resource consumption while improving AI/ML model performance for optimizing artificial intelligence (AI) and/or machine learning (ML) models. (Cummings [par(s) 8] “The present disclosure is related to techniques for optimizing artificial intelligence (AI) and/or machine learning (ML) models to reduce resource consumption while improving AI/ML model performance. In particular, the present disclosure provides a ML architecture search (MLAS) framework that involves generalized and/or hardware (HW)-aware ML architectures.”) Regarding claim 3 The combination of Chen, Cummings teaches claim 1. wherein the common optimization technique comprises at least one of: (See claim 1) Cummings further teaches removing redundant elements; (Cummings [fig(s) 1-2] [par(s) 53-65] “Each depth component in vector 200 represents a block in the CNN, where the value in each component represents a number of convolutional layers in each block, For example, a first block (B0) includes 2 layers, a second block (B1) includes 3 layers, a third block (B2) includes 2 layers, a fourth block (B3) includes 4 layers, and a fifth block (B4) includes 4 layers. Therefore, the set of values for the depth component of vector 200 includes the values 2, 3, or 4 (e.g., indicated by “Depth={2,3,4}” in FIG. 2). … In the example subnet configuration of FIG. 2, unused parameters (e.g., indicated by the value “X” in FIG. 2) are obtained when there are less than the maximum of 4 layers (e.g., 2 or 3 layers) in a block. These unused parameters do not significantly contribute to the prediction or inference determination and may be removed to reduce the overall size of the ML model.”;) removing unused elements; and (Cummings [fig(s) 1-2] [par(s) 53-65] “Each depth component in vector 200 represents a block in the CNN, where the value in each component represents a number of convolutional layers in each block, For example, a first block (B0) includes 2 layers, a second block (B1) includes 3 layers, a third block (B2) includes 2 layers, a fourth block (B3) includes 4 layers, and a fifth block (B4) includes 4 layers. Therefore, the set of values for the depth component of vector 200 includes the values 2, 3, or 4 (e.g., indicated by “Depth={2,3,4}” in FIG. 2). … In the example subnet configuration of FIG. 2, unused parameters (e.g., indicated by the value “X” in FIG. 2) are obtained when there are less than the maximum of 4 layers (e.g., 2 or 3 layers) in a block. These unused parameters do not significantly contribute to the prediction or inference determination and may be removed to reduce the overall size of the ML model.”;) encoding string values. (Cummings [fig(s) 1-2] [par(s) 82-88] “Each neuron 410 has one or more inputs and produces an output, which can be sent to one or more other neurons 410 (the inputs and outputs may be referred to as “signals”). Inputs to the neurons 410 of the input layer Lx can be feature values of a sample of external data (e.g., input variables xi). The input variables xi can be set as a vector containing relevant data (e.g., observations, ML features, etc.). The inputs to hidden units 410 of the hidden layers La, Lb, and Lc may be based on the outputs of other neurons 410. The outputs of the final output neurons 410 of the output layer Ly (e.g., output variables yj) include predictions, inferences, and/or accomplish a desired/configured task. The output variables yj may be in the form of determinations, inferences, predictions, and/or assessments. Additionally or alternatively, the output variables yj can be set as a vector containing the relevant data (e.g., determinations, inferences, predictions, assessments, and/or the like). In the context of ML, an “ML feature” (or simply “feature”) is an individual measureable property or characteristic of a phenomenon being observed. Features are usually represented using numbers/numerals (e.g., integers), strings, variables, ordinals, real-values, categories, and/or the like. Additionally or alternatively, ML features are individual variables, which may be independent variables, based on observable phenomenon that can be quantified and recorded. ML models use one or more features to make predictions or inferences. In some implementations, new features can be derived from old features.”; e.g., “The input variables xi can be set as a vector containing relevant data (e.g., observations, ML features, etc.)” along with “Features are usually represented using … strings” read(s) on “encoding string values” since the input layer does embedding for input features.) The combination of Chen, Cummings is combinable with Cummings for the same rationale as set forth above with respect to claim 1. Regarding claim 10 The claim is a system claim corresponding to the method claim 1, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Regarding claim 12 The claim is a system claim corresponding to the method claim 3, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Regarding claim 16 The claim is a computer program product claim corresponding to the method claim 1, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Regarding claim 18 The claim is a computer program product claim corresponding to the method claim 3, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Claim(s) 2, 7, 11, 13, 17, 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 11,301,762 B1) in view of Cummings et al. (US 20220036123 A1) in view of Bakker et al. (Clustering ensembles of neural network models) Regarding claim 2 The combination of Chen, Cummings teaches claim 1. Chen further teaches wherein the first machine learning model is an ensemble model, and wherein the common optimization technique and the model-specific optimization technique further include [cross]-model optimization techniques, wherein the [cross]-model optimization techniques apply to ensemble model types. (Chen [col 1 ln 66– col 3 ln 34] “a unified inference framework for heterogenous edge devices is provided that can accept machine 5 learning (ML) models from a user in any of multiple formats ( as generated by multiple different frameworks), convert and optimize these ML models for use by heterogeneous "edge" devices having heterogeneous computing resources, and deploy these ML models for use in one or more edge devices 10 of one or more different types.” [col 3 ln 43– col 7 ln 55] “FIG. 1 is a diagram illustrating an exemplary environment including a unified inference framework (“UIF”) server module 112 according to some embodiments. The UIF server module 112, in some embodiments, is a portion of software allowing users 118 to deploy and manage high performance machine learning models 130 running on connected devices 122 in production. Users 118 (e.g., individuals, organizations, even OEMs) can import to—or train machine learning models in—a provider network 100 (“the cloud”), and reliably deploy these models to large numbers of devices 122 at the edge. Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices.” [col 21 ln 42– col 24 ln 39] “In some embodiments, the operating environment sup ports many different types of machine learning models, such as multi arm bandit models, reinforcement learning models, ensemble machine learning models, deep learning models, and/or the like.”;) However, the combination of Chen, Cummings does not appear to explicitly teach: [cross]-model optimization techniques, wherein the [cross]-model optimization techniques apply to ensemble model types. Bakker teaches cross-model optimization techniques, wherein the cross-model optimization techniques apply to ensemble model types. (Bakker [sec(s) 7] “we have presented a method to summarize large ensembles of models to a small number of representative models. We have shown that predictions based on a weighted average of these representatives can be as good as, and sometimes even better than, predictions based on the full ensemble. We believe that this method provides an extremely useful addition to any method featuring an ensemble of models, such as bootstrapping, sampling of Bayesian posterior distributions or multitask learning. The method is not only valuable in terms of predictive quality, but also on a more abstract level. This improvement was apparent on the newspaper data, where different clusters of models brought out different aspects of the data.” [sec(s) 1] “Since W and M may contain any kind of elements, the method of clustering may well be used to find a workable representation of any oversized ensemble of (neural network) models. Taking the elements of W to be the models in the original (large) ensemble, represented in the case of neural networks by their weights, biases and overall structure (number of hidden layers, transfer functions, etc.) M can be a smaller set of networks best representing the features contained in W.” [sec(s) Abs] “We show that large ensembles of (neural network) models, obtained e.g. in bootstrapping or sampling from (Bayesian) probability distributions, can be effectively summarized by a relatively small number of representative models. In some cases this summary may even yield better function estimates. We present a method to find representative models through clustering based on the models’ outputs on a data set. We apply the method on an ensemble of neural network models obtained from bootstrapping on the Boston housing data, and use the results to discuss bootstrapping in terms of bias and variance. … a small number of representative models generally matches, or even surpasses, the performance of the full ensemble”) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Chen, Cummings with the cross-model optimization of Bakker. One of ordinary skill in the art would have been motived to combine in order to show that predictions based on a weighted average of the representatives can be as good as, and sometimes even better than, predictions based on the full ensemble, and provide an extremely useful addition to any method featuring an ensemble of models. (Bakker [sec(s) 7] “In the present article (which elaborates on our earlier work in Bakker and Heskes (1999)) we have presented a method to summarize large ensembles of models to a small number of representative models. We have shown that predictions based on a weighted average of these representatives can be as good as, and sometimes even better than, predictions based on the full ensemble. We believe that this method provides an extremely useful addition to any method featuring an ensemble of models, such as bootstrapping, sampling of Bayesian posterior distributions or multitask learning. The method is not only valuable in terms of predictive quality, but also on a more abstract level. This improvement was apparent on the newspaper data, where different clusters of models brought out different aspects of the data.”) Regarding claim 7 The combination of Chen, Cummings, Bakker teaches claim 2. Bakker further teaches wherein the cross-model optimization techniques comprise: (Bakker [sec(s) 7] “we have presented a method to summarize large ensembles of models to a small number of representative models. We have shown that predictions based on a weighted average of these representatives can be as good as, and sometimes even better than, predictions based on the full ensemble. We believe that this method provides an extremely useful addition to any method featuring an ensemble of models, such as bootstrapping, sampling of Bayesian posterior distributions or multitask learning. The method is not only valuable in terms of predictive quality, but also on a more abstract level. This improvement was apparent on the newspaper data, where different clusters of models brought out different aspects of the data.” [sec(s) 1] “Since W and M may contain any kind of elements, the method of clustering may well be used to find a workable representation of any oversized ensemble of (neural network) models. Taking the elements of W to be the models in the original (large) ensemble, represented in the case of neural networks by their weights, biases and overall structure (number of hidden layers, transfer functions, etc.) M can be a smaller set of networks best representing the features contained in W.”;) identifying a plurality of base models included in the first machine learning model; (Bakker [sec(s) 6] “The clustering algorithm was performed for each ensemble. We predicted the outputs in the smaller part of the database (the test set) through Eqs. (9) and (10). The predictions on the Boston housing data were compared to prediction through bagging (taking the average over the outputs of all models in the ensemble on a new input). For the newspaper data, for each split of the data we also trained one network similar to those described previously, with two hidden units, but with one output for each task (outlet). This means that all tasks shared the same input-to-hidden weights of the network, but had independent hidden-to-output weights. See Caruana (1997) for similar work. All prediction results are rated through their sum-squared error on the test set. Boston housing. We modeled the Boston housing problem by bootstrapping with multilayered perceptrons with their numbers of hidden units varying from 1 to 14. The models within one bootstrap ensemble always had the same numbers of hidden units. Each ensemble was clustered up to the point where no more improvement in sum-squared error was gained by adding more cluster centers”;) clustering the plurality of base models based on a set of model features; and (Bakker [sec(s) 6] “The clustering algorithm was performed for each ensemble. We predicted the outputs in the smaller part of the database (the test set) through Eqs. (9) and (10). The predictions on the Boston housing data were compared to prediction through bagging (taking the average over the outputs of all models in the ensemble on a new input). For the newspaper data, for each split of the data we also trained one network similar to those described previously, with two hidden units, but with one output for each task (outlet). This means that all tasks shared the same input-to-hidden weights of the network, but had independent hidden-to-output weights. See Caruana (1997) for similar work. All prediction results are rated through their sum-squared error on the test set. Boston housing. We modeled the Boston housing problem by bootstrapping with multilayered perceptrons with their numbers of hidden units varying from 1 to 14. The models within one bootstrap ensemble always had the same numbers of hidden units. Each ensemble was clustered up to the point where no more improvement in sum-squared error was gained by adding more cluster centers”;) for each respective cluster, extracting a model structure that matches output of each base model in the respective cluster. (Bakker [sec(s) Abs] “We show that large ensembles of (neural network) models, obtained e.g. in bootstrapping or sampling from (Bayesian) probability distributions, can be effectively summarized by a relatively small number of representative models. In some cases this summary may even yield better function estimates. We present a method to find representative models through clustering based on the models’ outputs on a data set. We apply the method on an ensemble of neural network models obtained from bootstrapping on the Boston housing data, and use the results to discuss bootstrapping in terms of bias and variance. … a small number of representative models generally matches, or even surpasses, the performance of the full ensemble” [sec(s) 1] “In our approach the distance function D(W, M) is based on model outputs instead of model parameters. We feel this approach is more intuitive, since model outputs on a known database provide a more direct representation of the models’ characteristics than the more abstract model parameters themselves” [sec(s) 6] “The clustering algorithm was performed for each ensemble. We predicted the outputs in the smaller part of the database (the test set) through Eqs. (9) and (10). The predictions on the Boston housing data were compared to prediction through bagging (taking the average over the outputs of all models in the ensemble on a new input). For the newspaper data, for each split of the data we also trained one network similar to those described previously, with two hidden units, but with one output for each task (outlet). This means that all tasks shared the same input-to-hidden weights of the network, but had independent hidden-to-output weights. See Caruana (1997) for similar work. All prediction results are rated through their sum-squared error on the test set. Boston housing. We modeled the Boston housing problem by bootstrapping with multilayered perceptrons with their numbers of hidden units varying from 1 to 14. The models within one bootstrap ensemble always had the same numbers of hidden units. Each ensemble was clustered up to the point where no more improvement in sum-squared error was gained by adding more cluster centers”; e.g., “representative models” read(s) on “model structure”.) The combination of Chen, Cummings, Bakker is combinable with Bakker for the same rationale as set forth above with respect to claim 2. Regarding claim 11 The claim is a system claim corresponding to the method claim 2, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Regarding claim 13 The claim is a system claim corresponding to the method claim 7, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Regarding claim 17 The claim is a computer program product claim corresponding to the method claim 2, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Regarding claim 19 The claim is a computer program product claim corresponding to the method claim 7, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 11,301,762 B1) in view of Cummings et al. (US 20220036123 A1) in view of Molchanov et al. (PRUNING CONVOLUTIONAL NEURAL NETWORKS FOR RESOURCE EFFICIENT INFERENCE) Regarding claim 4 The combination of Chen, Cummings teaches claim 3. wherein removing the redundant elements comprises: (See claim 3) However, the combination of Chen, Cummings does not appear to explicitly teach: parsing the first machine learning model; determining, based on the parsing, that a first element corresponds to a first input value; and upon determining, based on the parsing, that a second element also corresponds to the first input value, removing either the first element or the second element from the first machine learning model. Molchanov teaches parsing the first machine learning model; (Molchanov [fig(s) 1] [sec(s) 1] “While modern deep CNNs are composed of a variety of layer types, runtime during prediction is dominated by the evaluation of convolutional layers. With the goal of speeding up inference, we prune entire feature maps so the resulting networks may be run efficiently even on embedded devices. We interleave greedy criteria-based pruning with fine-tuning by backpropagation, a computationally efficient procedure that maintains good generalization in the pruned network.” [sec(s) 2] “Finding a good subset of parameters while maintaining a cost value as close as possible to the original is a combinatorial problem. It will require 2|W| evaluations of the cost function for a selected subset of data. For current networks it would be impossible to compute: for example, VGG-16 has |W| = 4224 convolutional feature maps. While it is impossible to solve this optimization exactly for networks of any reasonable size, in this work we investigate a class of greedy methods. Starting with a full set of parameters W, we iteratively identify and remove the least important parameters, as illustrated in Figure 1. By removing parameters at each iteration, we ensure the eventual satisfaction of the l0 bound on W’.”;) determining, based on the parsing, that a first element corresponds to a first input value; and (Molchanov [fig(s) 1] [sec(s) 1] “While modern deep CNNs are composed of a variety of layer types, runtime during prediction is dominated by the evaluation of convolutional layers. With the goal of speeding up inference, we prune entire feature maps so the resulting networks may be run efficiently even on embedded devices. We interleave greedy criteria-based pruning with fine-tuning by backpropagation, a computationally efficient procedure that maintains good generalization in the pruned network.” [sec(s) 2] “Consider a set of training examples D = {X = {x0, x1, ..., xN }, Y = {y0, y1, ..., yN}}, where x and y represent an input and a target output, respectively. The network’s parameters1 PNG media_image1.png 74 946 media_image1.png Greyscale are optimized to minimize a cost value C(D|W). … In the case of transfer learning, we adapt a large network initialized with parameters W0 pretrained on a related but distinct dataset. During pruning, we refine a subset of parameters which preserves the accuracy of the adapted network, C(D|W’) ≈ C(D|W). … Finding a good subset of parameters while maintaining a cost value as close as possible to the original is a combinatorial problem. It will require 2|W| evaluations of the cost function for a selected subset of data. For current networks it would be impossible to compute: for example, VGG-16 has |W| = 4224 convolutional feature maps. While it is impossible to solve this optimization exactly for networks of any reasonable size, in this work we investigate a class of greedy methods. Starting with a full set of parameters W, we iteratively identify and remove the least important parameters, as illustrated in Figure 1. By removing parameters at each iteration, we ensure the eventual satisfaction of the l0 bound on W’.”; e.g., “Starting with a full set of parameters W, we iteratively identify and remove the least important parameters, as illustrated in Figure 1” read(s) on “determining, based on the parsing, that a first element corresponds to a first input value”.) upon determining, based on the parsing, that a second element also corresponds to the first input value, removing either the first element or the second element from the first machine learning model. (Molchanov [fig(s) 1] [sec(s) 1] “While modern deep CNNs are composed of a variety of layer types, runtime during prediction is dominated by the evaluation of convolutional layers. With the goal of speeding up inference, we prune entire feature maps so the resulting networks may be run efficiently even on embedded devices. We interleave greedy criteria-based pruning with fine-tuning by backpropagation, a computationally efficient procedure that maintains good generalization in the pruned network.” [sec(s) 2] “Consider a set of training examples D = {X = {x0, x1, ..., xN }, Y = {y0, y1, ..., yN}}, where x and y represent an input and a target output, respectively. The network’s parameters1 PNG media_image1.png 74 946 media_image1.png Greyscale are optimized to minimize a cost value C(D|W). … In the case of transfer learning, we adapt a large network initialized with parameters W0 pretrained on a related but distinct dataset. During pruning, we refine a subset of parameters which preserves the accuracy of the adapted network, C(D|W’) ≈ C(D|W). … Finding a good subset of parameters while maintaining a cost value as close as possible to the original is a combinatorial problem. It will require 2|W| evaluations of the cost function for a selected subset of data. For current networks it would be impossible to compute: for example, VGG-16 has |W| = 4224 convolutional feature maps. While it is impossible to solve this optimization exactly for networks of any reasonable size, in this work we investigate a class of greedy methods. Starting with a full set of parameters W, we iteratively identify and remove the least important parameters, as illustrated in Figure 1. By removing parameters at each iteration, we ensure the eventual satisfaction of the l0 bound on W’.”; e.g., “Starting with a full set of parameters W, we iteratively identify and remove the least important parameters, as illustrated in Figure 1” read(s) on “upon determining, based on the parsing, that a second element also corresponds to the first input value, removing either the first element or the second element from the first machine learning model”.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Chen, Cummings with the redundant element removal of Molchanov. One of ordinary skill in the art would have been motived to combine in order to demonstrate superior performance compared to other criteria-based pruning approaches and reduce the number of convolutional feature maps and the total estimated floating-point operations. (Molchanov [sec(s) Abs] “The proposed criterion demonstrates superior performance compared to other criteria, e.g. the norm of kernel weights or feature map activation, for pruning large CNNs after adaptation to fine-grained classification tasks (Birds-200 and Flowers-102) relaying only on the first order gradient information. We also show that pruning can lead to more than 10× theoretical reduction in adapted 3D-convolutional filters with a small drop in accuracy in a recurrent gesture classifier. Finally, we show results for the largescale ImageNet dataset to emphasize the flexibility of our approach.” [sec(s) 3.3] “We now evaluate the full iterative pruning procedure on two transfer learning problems. We focus on reducing the number of convolutional feature maps and the total estimated floating point operations (FLOPs).”) Claim(s) 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 11,301,762 B1) in view of Cummings et al. (US 20220036123 A1) in view of Polleri et al. (US 20220366298 A1) Regarding claim 5 The combination of Chen, Cummings teaches claim 1. However, the combination of Chen, Cummings does not appear to explicitly teach: wherein a first machine learning model specific optimization technique applies to trees, and comprises at least one of: pruning useless nodes; or combining duplicated nodes. Polleri teaches wherein a first machine learning model specific optimization technique applies to trees, and comprises at least one of: pruning useless nodes; or combining duplicated nodes. (Polleri [par(s) 21] “If the number of data items, assigned to a particular category represented by a leaf node, does not meet a threshold level of data items, the system prunes the particular leaf node from the hierarchical tree. Furthermore, the system reassigns the data items, assigned to the particular category, to a parent category of the particular category. Subsequent to reassignment of the data items, the data items are used to train a machine learning model with the assigned category serving as a label in a supervised learning algorithm. The system applies the trained machine learning model to classify new data items.” [par(s) 91] “Correspondingly, the system also removes (“prunes”) the leaf node from the hierarchical tree of nodes that is the graphical representation of the classification (operation 228). Removing the category associated with too few data items, and removing the corresponding node, may improve the distinguishability of nodes from one another by improving the statistics associated with the remaining leaf and/or non-leaf nodes. The operation 228 may also improve the computational efficiency of the hierarchical classification when applied to target data.” [par(s) 86] “If the total number of data items assigned to direct and/or indirect children of a particular node does not meet a threshold, then those direct and/or indirect children of the particular node are pruned, and the corresponding data items are reassigned to the particular node.”;) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Chen, Cummings with the tree pruning of Polleri. One of ordinary skill in the art would have been motived to combine in order to improve the distinguishability of nodes from one another by improving the statistics associated with the remaining leaf and/or non-leaf nodes, and improve the computational efficiency of the hierarchical classification when applied to target data. (Polleri [par(s) 91] “Removing the category associated with too few data items, and removing the corresponding node, may improve the distinguishability of nodes from one another by improving the statistics associated with the remaining leaf and/or non-leaf nodes. The operation 228 may also improve the computational efficiency of the hierarchical classification when applied to target data.”) Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 11,301,762 B1) in view of Cummings et al. (US 20220036123 A1) in view of Sürer et al. (Coefficient tree regression for generalized linear models) Regarding claim 6 The combination of Chen, Cummings teaches claim 1. However, the combination of Chen, Cummings does not appear to explicitly teach: wherein a first machine learning model specific optimization technique applies to regression models, and comprises removing nodes with a coefficient of zero. Sürer teaches wherein a first machine learning model specific optimization technique applies to regression models, and comprises removing nodes with a coefficient of zero. (Sürer [fig(s) 1] [sec(s) 2 COEFFICIENT TREE REGRESSION (CTR) FOR GENERALIZED LINEAR MODELS (GLM)] “As mentioned above, CTRGLM iteratively splits the predictors into groups, which can be depicted by a tree structure as illustrated by Figure 1. After each iteration k − 1, CTRGLM splits the predictors into k mutually exclusive groups, and each level of the tree corresponds to the group structure produced after iteration k − 1. In level k − 1, we have k nodes, each of which corresponds to a group Gk−1,l for l = 1, … , k, and each label on the edge above the node corresponds to the estimated coefficient of the corresponding derived predictor at that iteration. In other words, after iteration k − 1, we have k mutually exclusive groups such that Gk−1,1 ∪ Gk−1,2 ∪···∪ Gk−1,k = {1, … , p} and Gk−1,i ∩ Gk−1,𝑗 =∅ ∀i, 𝑗 ∈ {1, … , k} and i ≠ 𝑗. The group Gk−1,k denotes the excluded group of predictors, i.e., having a coefficient of zero. We initialize the root node with G0,1 = {1, … , p}, since none of the predictors are included in the model initially, and thus, all predictors have a coefficient of zero. … Consequently, the total number of predictors decreased from 10 in the standard logistic regression model to four in the final CTRGLM model. In this sense, CTRGLM represents the data in a lower dimensional feature space by exploring the group structure. Moreover, the zero-coefficient predictors in group G4,5 = {5,7,9} are correctly excluded from the final CTRGLM model. Thus, our approach performs both feature extraction and feature selection.”;) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Chen, Cummings with the regression models of Sürer. One of ordinary skill in the art would have been motived to combine in order to provide a clear and concise interpretation of the data based on simple and highly interpretable models, and perform better than existing competitors in terms of computing time and predictive accuracy. (Sürer [sec(s) Abs] “In this paper we develop a coefficient tree regression algorithm for generalized linear models to discover the group structure from the data. The approach results in simple and highly interpretable models, and we demonstrated with real examples that it can provide a clear and concise interpretation of the data. Via simulation studies under different scenarios we showed that our approach performs better than existing competitors in terms of computing time and predictive accuracy.”) Claim(s) 8, 14, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 11,301,762 B1) in view of Cummings et al. (US 20220036123 A1) in view of Turco et al. (US 2019/0073570 A1) Regarding claim 8 The combination of Chen, Cummings teaches claim 1. Chen further teaches wherein the converted model comprises metadata information and model content, and wherein the metadata information of the converted model is stored as [tabular] data. (Chen [col 3 ln 43– col 7 ln 55] “Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices.” and “a translator 113A module (e.g., library, function, binary, etc.) of the model optimizer 114A translates the ML model 108 from a first format (as generated by a particular ML framework) into a “common” second format that is “unified” in that it can be run by all inference engines 132 of all edge devices 122. …Thus, by translating the ML model 108 into a common format, the model optimizer 114A can make the model “portable” in that it can be run at any/all of the one or more edge devices 122A-122N using a common inference engine 132 (e.g., via use of an inference library 134). This translation can be performed using tools and techniques known to those of skill in the art, which may include using a conversion library/module that identifies certain values (e.g., weights) in model files generated by a first framework and inserts them in a different format (or location) within files adherent to a different framework or format (e.g., a different framework's format, a standardized “generic” format such as the Open Neural Network eXchange format “ONNX”, etc.). As another example, the translator may create low-level machine code that is executable for multiple types of hardware backends. … an optimized model 129-whether it is simply translated, optimized, or both translated at optimized for execution-is provided to the identified one or more edge devices 122A-122N. … the optimized model 130 can include a first file carrying a graph-based representation of the model ( e.g., in JSON/XML) indicating the structure of a neural network, and include another file carrying the model weights. As another example, the optimized model 130 may include three Intermediate Representation (IR) files: a JSON file describing the optimized graph, a "params" file saving the values of model parameters, and a "so" file for the inference engine to run the model inference.”;) However, the combination of Chen, Cummings does not appear to explicitly teach: wherein the metadata information of the converted model is stored as [tabular] data. Turco teaches wherein the metadata information of the converted model is stored as tabular data. (Turco [fig(s) 1] [par(s) 32] “the model repository may be a file directory comprising models stored as individual XML or JSON files. According to other embodiments, the model repository consists of two database tables: a metadata table for storing model metadata such as model-ID, model name, author, model type, etc., and a content table which comprises a specification of the model in a generic data format, e.g. in JSON format, as a list of property-value pairs, in a binary format, etc. The best model for solving a computational task that is stored in the model repository has assigned a best-model-label or flag.” [par(s) 96] “The server computer system comprises a DBMS 132 which comprises one or more databases 134. A plurality of predictive models M1, M2, . . . , are stored in the model repository 136. In the depicted embodiment, the model repository consists of two database tables 138, 140. In the first repository table 138, a model identifier and metadata of the model such as model name, model type, creator, creation date, number of hyperparameters or the like are stored. In the second model repository table 140, the specification (“body”) of each model is stored in one or more respective table rows. For example, the specification can be stored as a string in XML or JSON format. The format used for storing models and the model repository typically requires a complex parsing operation in order to extract individual model parameters or parameter values.” [par(s) 113] “FIG. 3 is a flow chart of performing a model-based computational task that can be performed e.g. by components of the system depicted in FIGS. 1 and 2. In step 302, upon instantiation of the analytical program 156, the analytical program analyzes a configuration 162 in which the best-model is specified and accesses the model repository, e.g. for reading model metadata that is used for creating a table for storing the best-model.”;) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Chen, Cummings with the tabular data of Turco. One of ordinary skill in the art would have been motived to combine in order to enable to use a particular one of many different models within the context of a relational database without creating bottlenecks and with a very small storage and memory footprint. (Turco [par(s) 4] “Often, several thousand or even more predictive models need to be defined, trained and evaluated, whereby the structure and parameters of the created models vary greatly to increase the probability that a good predictive model can be identified. The storing and management of large number of models therefore is a bottleneck in many model-based prediction use case scenarios like drug screening.” [par(s) 48] “Some embodiments may allow a server to enable any client to use a particular one of 10,000 or 100,000 different models, e.g. different SVM and/or neural network models, within the context of a relational database without creating bottlenecks and with a very small storage and memory footprint.”) Regarding claim 14 The claim is a system claim corresponding to the method claim 8, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Regarding claim 20 The combination of Chen, Cummings teaches claim 1. (Note: Hereinafter, if a limitation has bold brackets (i.e. [·]) around claim languages, the bracketed claim languages indicate that they have not been taught yet by the current prior art reference but they will be taught by another prior art reference afterwards.) Chen further teaches wherein the converted model in the common intermediate format comprises metadata information and model content, and wherein the metadata information of the converted model in the common intermediate format is stored by [tabular] data. (Chen [col 3 ln 43– col 7 ln 55] “Embodiments can—via a model optimizer 114 of a UIF server module 112 and/or UIF client module 124—automatically tune ML models for optimal performance across multiple underlying hardware platforms, resulting in improved prediction/inference speeds that allow sophisticated computer vision, audio, and anomaly detection models to run efficiently, even on low-power devices.” and “a translator 113A module (e.g., library, function, binary, etc.) of the model optimizer 114A translates the ML model 108 from a first format (as generated by a particular ML framework) into a “common” second format that is “unified” in that it can be run by all inference engines 132 of all edge devices 122. …Thus, by translating the ML model 108 into a common format, the model optimizer 114A can make the model “portable” in that it can be run at any/all of the one or more edge devices 122A-122N using a common inference engine 132 (e.g., via use of an inference library 134). This translation can be performed using tools and techniques known to those of skill in the art, which may include using a conversion library/module that identifies certain values (e.g., weights) in model files generated by a first framework and inserts them in a different format (or location) within files adherent to a different framework or format (e.g., a different framework's format, a standardized “generic” format such as the Open Neural Network eXchange format “ONNX”, etc.). As another example, the translator may create low-level machine code that is executable for multiple types of hardware backends. … an optimized model 129-whether it is simply translated, optimized, or both translated at optimized for execution-is provided to the identified one or more edge devices 122A-122N. … the optimized model 130 can include a first file carrying a graph-based representation of the model ( e.g., in JSON/XML) indicating the structure of a neural network, and include another file carrying the model weights. As another example, the optimized model 130 may include three Intermediate Representation (IR) files: a JSON file describing the optimized graph, a "params" file saving the values of model parameters, and a "so" file for the inference engine to run the model inference.”;) However, the combination of Chen, Cummings does not appear to explicitly teach: wherein the metadata information of the converted model in the common intermediate format is stored by [tabular] data. (Note: Hereinafter, if a limitation has one or more bold underlines, the one or more underlined claim languages indicate that they are taught by the current prior art reference, while the one or more non-underlined claim languages indicate that they have been taught already by one or more previous art references.) Turco teaches wherein the metadata information of the converted model in the common intermediate format is stored by tabular data. (Turco [fig(s) 1] [par(s) 32] “the model repository may be a file directory comprising models stored as individual XML or JSON files. According to other embodiments, the model repository consists of two database tables: a metadata table for storing model metadata such as model-ID, model name, author, model type, etc., and a content table which comprises a specification of the model in a generic data format, e.g. in JSON format, as a list of property-value pairs, in a binary format, etc. The best model for solving a computational task that is stored in the model repository has assigned a best-model-label or flag.” [par(s) 96] “The server computer system comprises a DBMS 132 which comprises one or more databases 134. A plurality of predictive models M1, M2, . . . , are stored in the model repository 136. In the depicted embodiment, the model repository consists of two database tables 138, 140. In the first repository table 138, a model identifier and metadata of the model such as model name, model type, creator, creation date, number of hyperparameters or the like are stored. In the second model repository table 140, the specification (“body”) of each model is stored in one or more respective table rows. For example, the specification can be stored as a string in XML or JSON format. The format used for storing models and the model repository typically requires a complex parsing operation in order to extract individual model parameters or parameter values.” [par(s) 113] “FIG. 3 is a flow chart of performing a model-based computational task that can be performed e.g. by components of the system depicted in FIGS. 1 and 2. In step 302, upon instantiation of the analytical program 156, the analytical program analyzes a configuration 162 in which the best-model is specified and accesses the model repository, e.g. for reading model metadata that is used for creating a table for storing the best-model.”;) The combination of Chen, Cummings is combinable with Turco for the same rationale as set forth above with respect to claim 8. Claim(s) 9, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 11,301,762 B1) in view of Cummings et al. (US 20220036123 A1) in view of Shridhar et al. (Interoperating Deep Learning models with ONNX.jl) Regarding claim 9 The combination of Chen, Cummings teaches claim 1. Chen further teaches the remote system instantiates an optimized version of the first machine learning model by: (Chen [fig(s) 1] “Model Optimizer 114B” and “Optimized Model 130” [col 3 ln 43– col 7 ln 55] “the translator may create low-level machine code that is executable for multiple types of hardware backends. … Thus, an optimized model 129-whether it is simply translated, optimized, or both translated at optimized for execution-is provided to the identified one or more edge devices 122A-122N. … Thus, the (possibly optimized) model provided to each edge device 122 at (3A) or (3C) may be further optimized (at (4A)) or not further optimized (circle (4B)), and then provided to an inference engine 132 as optimized model 130. The optimized model 130 may be in a variety of different formats based on the particular implementation. As one example, the optimized model 130 can include a first file carrying a graph-based representation of the model ( e.g., in JSON/XML) indicating the structure of a neural network, and include another file carrying the model weights. As another example, the optimized model 130 may include three Intermediate Representation (IR) files: a JSON file describing the optimized graph, a "params" file saving the values of model parameters, and a "so" file for the inference engine to run the model inference.” [col 5 ln 47– col 5 ln 63] “In this scenario, the user 118 may cause the client device 120 to send, at circle (2), a request to deploy the ML model(s) 108 to one or more edge devices 122A-122N.”;) instantiating the optimized version of the first machine learning model based on [the fourth machine learning model in the first format]. (Chen [fig(s) 1] “Model Optimizer 114B” and “Optimized Model 130” [col 3 ln 43– col 7 ln 55] “the translator may create low-level machine code that is executable for multiple types of hardware backends. … Thus, an optimized model 129-whether it is simply translated, optimized, or both translated at optimized for execution-is provided to the identified one or more edge devices 122A-122N. … Thus, the (possibly optimized) model provided to each edge device 122 at (3A) or (3C) may be further optimized (at (4A)) or not further optimized (circle (4B)), and then provided to an inference engine 132 as optimized model 130. The optimized model 130 may be in a variety of different formats based on the particular implementation. As one example, the optimized model 130 can include a first file carrying a graph-based representation of the model ( e.g., in JSON/XML) indicating the structure of a neural network, and include another file carrying the model weights. As another example, the optimized model 130 may include three Intermediate Representation (IR) files: a JSON file describing the optimized graph, a "params" file saving the values of model parameters, and a "so" file for the inference engine to run the model inference.” [col 5 ln 47– col 5 ln 63] “In this scenario, the user 118 may cause the client device 120 to send, at circle (2), a request to deploy the ML model(s) 108 to one or more edge devices 122A-122N.”;) However, the combination of Chen, Cummings does not appear to explicitly teach: converting the binary model to a third model in the common intermediate format; converting the third model to a fourth machine learning model in the first format; and instantiating the optimized version of the first model based on [the fourth machine learning model in the first format]. Shridhar teaches converting the binary model to a third model in the common intermediate format; (Shridhar [fig(s) 1] “ONNX Serialized model” and “ModelProto object” [sec(s) 3-3.1] “—ModelProto: A very high-level struct that holds all the information. ONNX models are read directly into this structure. —GraphProto: This structure captures the entire computation graph of the model. —NodeProto and TensorProto: Information regarding individual nodes in the graph (inputs, outputs and finer attributes) and weights associated with the nodes.” [sec(s) 3.2] “ModelProto structure is the structure that holds all the information needed to load a model. Internally, it holds data such as the version information, model version, docstring, producer details and most importantly: the computation graph. An ONNX model, once read using ProtoBuf.jl is loaded into this ModelProto object before extracting the graph details. Naturally, at the heart of this is the graph::GraphProto attribute that stores the computation graph of the model.”; e.g., “ONNX Serialized model” read(s) on “binary model”. In addition, e.g., “ModelProto object” along with “ONNX models are read directly into this structure” read(s) on “third model in the common intermediate format”.) converting the third model to a fourth machine learning model in the first format; and (Shridhar [fig(s) 1] “ONNX Serialized model” and “ModelProto object” [sec(s) 3-3.1] “—ModelProto: A very high-level struct that holds all the information. ONNX models are read directly into this structure. —GraphProto: This structure captures the entire computation graph of the model. —NodeProto and TensorProto: Information regarding individual nodes in the graph (inputs, outputs and finer attributes) and weights associated with the nodes.” [sec(s) 3.2] “ModelProto structure is the structure that holds all the information needed to load a model. Internally, it holds data such as the version information, model version, docstring, producer details and most importantly: the computation graph. An ONNX model, once read using ProtoBuf.jl is loaded into this ModelProto object before extracting the graph details. Naturally, at the heart of this is the graph::GraphProto attribute that stores the computation graph of the model.” [sec(s) 3.3] “The GraphProto structure stores information about particular nodes in the graph. This includes the node metadata, name, input, output and the pre-trained parameters in the initializer attribute.”; e.g., “GraphProto: This structure captures the entire computation graph of the model” read(s) on “fourth machine learning model in the first format”.) instantiating the optimized version of the first machine learning model based on the fourth machine learning model in the first format. (Shridhar [fig(s) 1] “ONNX Serialized model” and “ModelProto object” [sec(s) 2] “ONNX defines the computation graph for a deep learning model along with various operators used in the model. It provides a set of specifications to convert a model to a basic ONNX format, and another set of specifications to get the model back from this ONNX form. … Machine learning models can be converted to a serialized ONNX format which can then be run on a number devices.” [sec(s) 2.1] “ONNX is is usable anywhere from small mobile devices to large server farms, across chipsets and vendors, and with extensive runtimes and tools support. ONNX reduces the friction of moving trained AI models among your favorite tools and frameworks and platforms. A simple example of how ONNX is ideal for ML is the case when large deep learning models need to be deployed. … By connecting the common dots from different frameworks, ONNX makes it possible to express a model of type A to type B, thus saving time and the need to train the model again.” [sec(s) 3-3.1] “—ModelProto: A very high-level struct that holds all the information. ONNX models are read directly into this structure. —GraphProto: This structure captures the entire computation graph of the model. —NodeProto and TensorProto: Information regarding individual nodes in the graph (inputs, outputs and finer attributes) and weights associated with the nodes.” [sec(s) 3.2] “ModelProto structure is the structure that holds all the information needed to load a model. Internally, it holds data such as the version information, model version, docstring, producer details and most importantly: the computation graph. An ONNX model, once read using ProtoBuf.jl is loaded into this ModelProto object before extracting the graph details. Naturally, at the heart of this is the graph::GraphProto attribute that stores the computation graph of the model.”; e.g., “GraphProto: This structure captures the entire computation graph of the model” read(s) on “fourth machine learning model in the first format”.) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Chen, Cummings with the machine learning model instantiation of Shridhar. One of ordinary skill in the art would have been motived to combine in order to make it very easy and straight-forward to use trained models while allowing framework interoperability for many excellent machine learning libraries in various languages. (Shridhar [sec(s) 2] “At a high level, ONNX is designed to allow framework interoperability. There are many excellent machine learning libraries in various languages: PyTorch [2], TensorFlow [1], MXNet [5], and Caffe [20] are just a few that have become very popular in recent years, but there are many others as well.” [sec(s) 8] “Developing ONNX.jl has been tremendous learning experience. From studying about Intermediate Representation formats for deep learning models with millions of parameters to loading them in just a couple of lines of code, ONNX.jl has made it very easy and straight-forward to use a high quality trained model as a starting point for many projects. Once such example I’d like to point out is DenseNet-121 model. This is deep convolutional network that has multiple Convolutional, Dense and Pooling blocks. Naturally, implementing this in any framework is going to be a challenging task. However, thanks to ONNX, we can now use an earlier implementation to import this model into any other framework.”) Regarding claim 15 The claim is a system claim corresponding to the method claim 9, and is directed to largely the same subject matter. Thus, it is rejected for the same reasons as given in the rejections of the method claim. Prior Art The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Rio et al. (Decoupling Application Intelligence and its Orchestration on IoT Devices) teaches generating a binary data based on serialization for ONNX. Jin et al. (Compiling ONNX Neural Network Models Using MLIR) teaches generating executables for ONNX. Szabo et al. (Distributed Machine Learning Using Data Parallelism on Mobile Platform) teaches storing model data in a relational database. Shekar et al. (US 20210357196 A1) teaches storing metadata in a relational database. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEHWAN KIM whose telephone number is (571)270-7409. The examiner can normally be reached Mon - Thu 7:00 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached on (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SEHWAN KIM/Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Show 8 earlier events
Apr 06, 2026
Request for Continued Examination
Apr 10, 2026
Response after Non-Final Action
Apr 20, 2026
Non-Final Rejection mailed — §101, §103
Jul 09, 2026
Response Filed
Jul 28, 2026
Examiner Interview Summary
Jul 28, 2026
Applicant Interview (Telephonic)
Jul 31, 2026
Final Rejection mailed — §101, §103
Sep 18, 2026
Interview Requested

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12619853
DECISION-MAKING DEVICE, UNMANNED SYSTEM, DECISION-MAKING METHOD, AND PROGRAM
5y 6m to grant Granted May 05, 2026
Patent 12619921
PREDICTIVE FOG DATA CENTER MIGRATION
3y 8m to grant Granted May 05, 2026
Patent 12608592
AUTOMATED ELECTRIC SUBMERSIBLE PUMP (ESP) FAILURE ANALYSIS
3y 4m to grant Granted Apr 21, 2026
Patent 12602595
SYSTEM AND METHOD OF USING A KNOWLEDGE REPRESENTATION FOR FEATURES IN A MACHINE LEARNING CLASSIFIER
9y 4m to grant Granted Apr 14, 2026
Patent 12602580
Dataset Dependent Low Rank Decomposition Of Neural Networks
6y 9m to grant Granted Apr 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
61%
Grant Probability
99%
With Interview (+67.3%)
4y 0m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 156 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month