Prosecution Insights
Last updated: October 02, 2026
Application No. 18/125,062

MACHINE LEARNING (ML) BASED SOFTWARE KERNEL SELECTION

Final Rejection §103§112
Filed
Mar 22, 2023
Examiner
STANDKE, ADAM C
Art Unit
2100
Tech Center
2100 — Computer Architecture & Software
Assignee
Qualcomm Incorporated
OA Round
2 (Final)
53%
Grant Probability
Moderate
3-4
OA Rounds
9m
Est. Remaining
79%
With Interview

Examiner Intelligence

Grants 53% of resolved cases
53%
Career Allowance Rate
77 granted / 146 resolved
-2.3% vs TC avg
Strong +27% interview lift
Without
With
+26.6%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
15 currently pending
Career history
174
Total Applications
across all art units

Statute-Specific Performance

§101
18.4%
-21.6% vs TC avg
§103
56.5%
+16.5% vs TC avg
§102
9.0%
-31.0% vs TC avg
§112
14.4%
-25.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 146 resolved cases

Office Action

§103 §112
DETAILED ACTION Examiner Remarks The Examiner notes that the instant application was previously examined by a different examiner. As such, the Examiner proceeds with prosecution giving full faith and credit to the search and action of the previous examiner per MPEP § 704.01:When an examiner is assigned to act on an application which has received one or more actions by some other examiner, full faith and credit should be given to the search and action of the previous examiner unless there is a clear error in the previous action or knowledge of other prior art. In general, the second examiner should not take an entirely new approach to the application or attempt to reorient the point of view of the previous examiner, or make a new search in the mere hope of finding something. See MPEP § 719.05. In light of Applicant’s Remarks and amendments submitted on 02/03/2026, Examiner has withdrawn the objections to the Drawings, Specification, and Claims. Examiner has also withdrawn the claim interpretation and the claim rejections under l12(f), l12(a) and 112(b) for claims 1, 3, 5-10, 24-25, and 27-28. However, the claim interpretation under l12(f) for claim 4 has not been withdrawn and the 112(b)-rejection relating to relative terminology has also not been withdrawn. Response to Arguments Applicant’s arguments with respect to independent claims 1, 11, and 29 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged by Applicant in the arguments submitted in the Remarks of 02/03/2026. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: a training data repository configured to accept an optimal software kernel in claim 4. Regarding claim 4 and the above-noted three-prong test, the recited machine learning (ML) model selection engine is a generic placeholder, data repository is a generic placeholder, “accept an optimal software kernel”, is functional language and there is no recitation of sufficient structure in claim 4 to perform the acceptance of an optimal software kernel. Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. The term “valid software kernels” in claims 1 and 29 is a relative term which renders the claim indefinite. The term “valid” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. In particular, the specification describes, at a high level of generality, what makes a kernel “valid” within the described selection framework (see, e.g. paragraph [0052-0054]: “the validation rules engine 730 produces a plurality of valid software kernels 740 (labeled k 1, k2, k3 in FIG. 7)… In one example, the plurality of valid software kernels 740 is produced based on a plurality of validation rules which determines which software kernels of the plurality of software kernels 720 are capable of implementing the desired application. For example, the plurality of validation rules determines all software kernels which can support a desired mathematical operation” and paragraph [0068]: “In one example, the machine learning process may employ a validation rules engine which produces a plurality of valid software kernels based on a plurality of validation rules and determines valid software kernels of the plurality of software kernels which are capable of implementing the desired application. In one example, the plurality of validation rules determines all software kernels which may support a desired mathematical operation. For example, the machine learning process may select the ML-selected software kernel based on the trained ML model instead of the static optimization parameters (e.g., static bid values)”). The described “mathematical operation” and “validation rules” that “determine valid software kernels” are vague and do not provide any concrete statistical measure or quantitative standard that one of ordinary skill could use to identify or verify kernel validity independently of the described system. Thus, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. For examination purposes, the term valid software kernels, is being interpreted as any plurality of software kernels that are functionally viable/practically executable for a given task or purpose. Additionally, claims 2-10, 12-28 and 30 which depend either directly or indirectly from independent claims 1, 11, and 29 respectively, are also rejected under 35 U.S.C. 112(b) as being indefinite under the same rationales as independent claims 1, 11, and 29. Claims 9-10 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 9 recites the limitation "the first processor.” There is insufficient antecedent basis for this limitation in the claim. For examination purposes claim 9 is dependent upon claim 8. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-4, 8-16, 24-26, and 29 are rejected under 35 U.S.C. 103 as being unpatentable over Barker (US 20210192334 A1; hereinafter Barker) in view of Liu et al., "SEER: A time prediction model for CNNs from GPU kernel's view." 2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT). IEEE, 2021(“Liu”). Regarding Independent Claim 1, Barker discloses an apparatus for kernel selection, the apparatus comprising: a non-transitory memory configured to store one or more input parameters(Barker, paras., 0073-0077, see also fig., 8, “[I]nference and/or training logic 815 may include, without limitation, code and/or data storage 801 to store forward and/or output weight and/or input/output data, and/or other parameters... code and/or code and/or data storage 801 may be cache memory, dynamic randomly addressable memory ("DRAM"), static randomly addressable memory ("SRAM"), non-volatile memory (e.g., Flash memory), or other storage.”); a first processor coupled to the non-transitory memory, the first processor configured to accept the one or more input parameters and a plurality of software kernels, and configured to generate a plurality of valid software kernels1 based on the input parameters and the plurality of software kernels (see, e.g., paragraph [0016]: “FIG. 1 is an illustration of an example environment… The environment includes a kernel selection system 100” [i.e., the example environment in FIG. 1 is an apparatus for kernel selection] and paragraph [0019]: “The candidate kernel generator 106 determines one or more candidate kernels that may be utilized by the kernel processor 104 to determine a result computation. The candidate kernel generator 106 identifies the kernels in database 102 that are available to execute the operation that has been received in a request as well as characteristics of the input data. For example, one or more kernels may be constrained in the size of the matrices that may be used as input to calculate a result and/or one or more of the kernels may be specific to a particular computation. Thus, candidate kernel generator 106 can determine a list of the kernels stored in the database 102 that may be utilized to provide the application 130 with a result” [i.e., the kernel selection system functions as a validation rules engine that accepts both input data (input parameters) and software kernels from the database]); a second processor coupled to the non-transitory memory(Barker, para. 0079, “ALUs 910 may be included within a processor's execution units or otherwise within a bank of ALUs accessible by a processor's execution units either within same processor or distributed between different processors of different types ( e.g., central processing units, graphics processing units, fixed function units, etc.)... any portion of activation storage 920 may be included with other on-chip or off-chip data storage, including a processor's Ll, L2, or L3 cache or system memory.”). While Barker discloses the second processor, Backer does not disclose: configured to generate a machine learning (ML) selected software kernel based on the plurality of valid software kernels and with a shorter timing performance than a software kernel, wherein the software kernel includes at least one static optimization parameter. However, Liu teaches configured to generate a machine learning (ML) selected software kernel based on the plurality of valid software kernels and with a shorter timing performance than a software kernel, wherein the software kernel includes at least one static optimization parameter(Liu, pgs., 174-184, see also figs. 2, 9, 10, and 13,“ There are three classes of efficient convolution algorithms, based on GEMM, Winograd and FFT[based on the plurality of valid software kernels]... [t]o clearly show the building process of our model, we illustrated it as a 4-step workflow, as shown in Fig.2. The 4 steps are: 1. Kernel classification. We train a classification tree to predict the resource bound type of a target kernel[]...[w]e propose polynomial models for Static metrics which are used in our model... in group(a) in Table I... we try to find the relationship between convolution hyper-parameters and Static metrics[wherein the software kernel includes at least one static optimization parameter]... [w]e use our proposed performance model to predict the best convolution algorithm[configured to generate a machine learning (ML) selected software kernel]... [b]y using SEER as algorithm picker, we can save 44.58% time on this operator, and 11.37% time on the overall forward computation time[and with a shorter timing performance than a software kernel].”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Barker with the teachings of Liu the motivation to do so would be optimize the execution of a deep learning model by using a machine learning based model to select the best kernel for executing operations for the deep learning model(Liu, pg., 174, “We propose SEER, a hybrid time prediction model for CNNs, which combines both analytical models and learning-based models... [w]e show that SEER can achieve low prediction error rate of 14.71% at op-level and 1.79% at network-level. When used for choosing best convolution algorithm, SEER outperforms cuDNN official algorithm picker, with 7.14% lower error rate.”). Regarding claim 2, as discussed above, Barker in view of Liu discloses the apparatus of claim 1. Barker further discloses wherein the one or more input parameters include one of the following: a plurality of tensors, an attribute of a mathematical operation or function, a data attribute or an attribute of a tensor descriptor (see, e.g., paragraph [0017]: “Each kernel may be associated with a particular computation and may be utilized with provided input data to produce a result. For example, a kernel stored in database 105 may be associated with a general matrix multiplication (GeMM) computation. The GeMM kernel may then be utilized by an execution component, such as kernel processor 104, which can perform the computation utilizing the GeMM kernel” [i.e., the input parameters used in the candidate kernel generator includes a mathematical operation /function (GeMM)] and paragraph [0069]: In at least one embodiment, an untrained neural network is trained using a training dataset. In at least one embodiment, a training framework is a PyTorch framework, Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit/CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework” [i.e., the training frameworks inherently include tensor operations, mathematical functions and tensor attributes (such as shape/dimension, data type, and/or memory layout)]). Regarding claim 3, as discussed above, Barker in view of Liu discloses the apparatus of claim 1. Barker further discloses wherein the second processor is further configured to generate the ML-selected software kernel based on the plurality of valid software kernels by using a first trained ML model (see, e.g., paragraph [0038]: “The neural network is trained to generate a ranked list of candidate kernels that may be utilized to perform a matrix computation. In some embodiments, the neural network may be trained using one or more embodiments described herein. For example, the neural network can be trained using a training system 300, as illustrated in FIG. 3” and paragraph [0039]: “At step 415, the neural network generates a ranked list of kernels to perform the requested computation. For example, the neural network can determine a relevancy score for each of the kernels in a list of kernels. A relevancy score can be indicative of how likely a kernel is an optimal kernel for performing the computation. At step 420, the optimal kernel is selected” [i.e., a ranked list candidate kernels are (ML-selected software kernels) based on a relevancy score (validation) of software kernels, is produced by the neural network generated by the training system (trained ML model engine)]). Regarding claim 4, as discussed above, Barker in view of Liu discloses the apparatus of claim 3. Barker further discloses further comprising a training data repository2 configured to accept an optimal software kernel (see, e.g., paragraph [0028]: “Kernel database 308 includes a plurality of kernels, each of which can perform a computation on a set of input matrices. For each training input, each of the plurality of kernels stored in kernel database 308 may be provided to the kernel processor 302 to perform a computation… The result set may include, for example, for each kernel executed on a processor, processor information, run-time for the kernel, and/or a class for the training input” [i.e., kernel database 308 depicted in Fig. 3 with bi-directional arrows meaning it accepts/receives and sends/outputs software kernels to/from training system 300], paragraph [0030]: “Relevancy calculator 310 determines a relevancy score for each of the kernels in the result set based on the run-time of the kernel while executing the matrix computation… The result set may then be ranked according to relevancy scores for further processing” and paragraph [0031]: “The result set and the training input can be provided to the neural network as training data. The neural network may then process the input to determine a predicted relevancy score for each of the kernels, which then can be ranked into a predicted ranking of the kernels” [i.e., the kernels in the training system communicated with database 308 are ranked for relevancy which is a form of optimization]). Regarding claim 8, as discussed above, Barker in view of Liu discloses the apparatus of claim 1. Barker further discloses further comprising a third processor coupled to the non-transitory memory, the third processor configured to receive the plurality of valid software kernels and further configured to generate a plurality of performance metrics based on the plurality of valid software kernels (see, e.g., paragraph [0020]: “Filter 108 removes any kernels from the list of candidate kernels that are not practical and/or impossible to execute given particular restraints of the system” [i.e., the kernels are validated by the Filter 108] and paragraph [0028]: “For each training input, each of the plurality of kernels stored in kernel database 308 may be provided to the kernel processor 302 to perform a computation. For example, kernel database may include kernels K1 . . . Kn, each of which can be provided to kernel processor 302 (or provided to a plurality of kernel processors) to generate a result set. The result set may include information regarding the execution of the training input on the kernel processor 302. The result set may include, for example, for each kernel executed on a processor, processor information, run-time for the kernel” [i.e., the kernel processer component in the training system produces a result set of performance metrics for the valid input kernels, functioning as a performance evaluation engine]). Regarding claim 9, as discussed above, Barker in view of Liu discloses the apparatus of claim 8. Barker further discloses wherein the third processor and the first processor are the same(see, e.g., Barker paragraph [0030]: “Relevancy calculator 310 determines a relevancy score for each of the kernels in the result set based on the run-time of the kernel while executing the matrix computation… Thus, a kernel with a relevancy score of 0.25 may have performed worse (e.g., taken longer to run) than a kernel that is assigned a relevancy score of 0.75. The result set may then be ranked according to relevancy scores for further processing” [i.e., the relevancy calculator component in the training system receives the result set of performance metrics from the kernel processor component for scoring and ranking the kernels]). Regarding claim 10, as discussed above, Barker in view of Liu discloses the apparatus of claim 9. Barker further discloses wherein the third processor is further configured to implement a selection function for each of the plurality of valid software kernels to determine an optimal software kernel (see, e.g., paragraph [0030]: “Relevancy calculator 310 determines a relevancy score for each of the kernels in the result set based on the run-time of the kernel while executing the matrix computation… The result set may then be ranked according to relevancy scores for further processing” [i.e., the relevancy calculator determines a relevancy score for the kernels and ranks the kernels (determining optimal software kernels), which is functionally a selection function performed on each of the valid kernels]). Regarding Independent claim 11, Barker discloses A method for kernel selection, the method comprising: inputting a plurality of valid software kernels3 to a trained machine learning (ML) model engine (see, e.g., paragraph [0020]: “Filter 108 removes any kernels from the list of candidate kernels that are not practical and/or impossible to execute given particular restraints of the system” [i.e., the kernels are validated by the Filter 108], paragraph [0026]: “FIG. 3 is an example environment that may be utilized to train a neural network to rank kernels to perform a computation. The training system 300 can be the same system as illustrated in FIG. 1 as training system 120. The training system includes a kernel processor 302, which may be the same, or share characteristics with kernel processor 104 of FIG. 1” and paragraph [0028]: “For each training input, each of the plurality of kernels stored in kernel database 308 may be provided to the kernel processor 302 to perform a computation. For example, kernel database may include kernels K1 . . . Kn, each of which can be provided to kernel processor 302 (or provided to a plurality of kernel processors) to generate a result set.” [i.e., the training system (trained ML model engine) includes a kernel processor 302, which receives a plurality of valid software kernels]); configuring the trained ML model engine to generate a first trained machine learning (ML) model based on the plurality of valid software kernels (see, e.g., paragraph [0021]: “Once a candidate list of kernels has been determined, the neural network 110 is provided the list along with the input data. The neural network is trained to determine, based on the list of candidate kernels, a relevancy score for each of the kernels that is predictive of how a given kernel will perform a computation on the input data. The neural network may be trained by a training system 120” [i.e., the trained neural network (trained ML model) is generated via training by the training system 120 (trained ML model engine)]); and using the first trained ML model to generate a machine learning (ML)-selected software kernel based on the plurality of valid software kernels (see, e.g., paragraph [0021]: “Once a candidate list of kernels has been determined, the neural network 110 is provided the list along with the input data. The neural network is trained to determine, based on the list of candidate kernels, a relevancy score for each of the kernels that is predictive of how a given kernel will perform a computation on the input data” and paragraph [0022]: “Sorter 112 sorts the candidate kernels by relevancy based on the output of the neural network. For example, for a given list of kernels with relevancy scores, sorter 112 can sort the list so that the first kernel in the list is the most relevant kernel for performing a computation on the input data. Selection engine 114 then chooses a kernel from the sorted list, such as the kernel with the highest relevancy score” [i.e., under the broadest reasonable interpretation, the neural network (trained ML Model) generates a (ML)-selected software kernel by determining ML relevancy score data that results in a selected kernel from the coupled selection engine]). Barker does not disclose: and with a shorter timing performance than a software kernel, wherein the software kernel includes at least one static optimization parameter. However, Liu teaches and with a shorter timing performance than a software kernel, wherein the software kernel includes at least one static optimization parameter(Liu, pgs., 174-184, see also figs. 2, 9, 10, and 13,“To clearly show the building process of our model, we illustrated it as a 4-step workflow, as shown in Fig.2. The 4 steps are: 1. Kernel classification. We train a classification tree to predict the resource bound type of a target kernel...[w]e propose polynomial models for Static metrics which are used in our model... in group(a) in Table I... we try to find the relationship between convolution hyper-parameters and Static metrics[wherein the software kernel includes at least one static optimization parameter]... [b]y using SEER as algorithm picker, we can save 44.58% time on this operator, and 11.37% time on the overall forward computation time[and with a shorter timing performance than a software kernel].”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Barker with the teachings of Liu the motivation to do so would be optimize the execution of a deep learning model by using a machine learning based model to select the best kernel for executing operations for the deep learning model(Liu, pg., 174, “We propose SEER, a hybrid time prediction model for CNNs, which combines both analytical models and learning-based models... [w]e show that SEER can achieve low prediction error rate of 14.71% at op-level and 1.79% at network-level. When used for choosing best convolution algorithm, SEER outperforms cuDNN official algorithm picker, with 7.14% lower error rate.”). Regarding claim 12, as discussed above, Barker in view of Liu discloses the method of claim 11. Barker further discloses further comprising inputting one or more input parameters to a validation rules engine (see, e.g., paragraph [0019]: “The candidate kernel generator 106 determines one or more candidate kernels that may be utilized by the kernel processor 104 to determine a result computation. The candidate kernel generator 106 identifies the kernels in database 102 that are available to execute the operation that has been received in a request as well as characteristics of the input data.” [i.e., the components in the kernel selection system receive both input data (input parameters) and software kernels from the database]). Regarding claim 13, as discussed above, Barker in view of Liu discloses the method of claim 12. Barker further discloses further comprising inputting a plurality of software kernels to the validation rules engine (see, e.g., paragraph [0019]: “The candidate kernel generator 106 determines one or more candidate kernels that may be utilized by the kernel processor 104 to determine a result computation. The candidate kernel generator 106 identifies the kernels in database 102 that are available to execute the operation that has been received in a request as well as characteristics of the input data.” [i.e., the components in the kernel selection system receive both input data (input parameters) and software kernels from the database] ). Regarding claim 14, as discussed above, Barker in view of Liu discloses the method of claim 13. Barker further discloses further comprising generating the plurality of valid software kernels based on the one or more input parameters and the plurality of software kernels (see, e.g., paragraph [0019]: “The candidate kernel generator 106 identifies the kernels in database 102 that are available to execute the operation that has been received in a request as well as characteristics of the input data. For example, one or more kernels may be constrained in the size of the matrices that may be used as input to calculate a result and/or one or more of the kernels may be specific to a particular computation. Thus, candidate kernel generator 106 can determine a list of the kernels stored in the database 102 that may be utilized to provide the application 130 with a result” and paragraph [0020]: “Filter 108 removes any kernels from the list of candidate kernels that are not practical and/or impossible to execute given particular restraints of the system. For example, filter 108 may identify the hardware constraints of kernel processor 104 and determine that, for a given kernel, the kernel processor 104 does not have (or is unlikely to have) the resources to process the given inputs with that kernel. Thus, filter 108 can remove the kernel from the candidate kernel list so that the neural network does not process that kernel as a potentially optimal kernel” [i.e., a filter generates a plurality of valid software kernels based on the list of kernels and input data from the candidate kernel generator]). Regarding claim 15, as discussed above, Barker in view of Liu discloses the method of claim 14. Barker further discloses wherein the generating the plurality of valid software kernels is implemented by the validation rules engine (see, e.g., paragraph [0019]: “The candidate kernel generator 106 determines one or more candidate kernels that may be utilized by the kernel processor 104 to determine a result computation. The candidate kernel generator 106 identifies the kernels in database 102 that are available to execute the operation that has been received in a request as well as characteristics of the input data. For example, one or more kernels may be constrained in the size of the matrices that may be used as input to calculate a result and/or one or more of the kernels may be specific to a particular computation. Thus, candidate kernel generator 106 can determine a list of the kernels stored in the database 102 that may be utilized to provide the application 130 with a result” and paragraph [0020]: “Thus, filter 108 can remove the kernel from the candidate kernel list so that the neural network does not process that kernel as a potentially optimal kernel” [i.e., the kernel selection system functions as a validation rules engine that uses a filter 108 to determine one or more candidate kernels that may be utilized by the kernel processor 104 to determine a result computation (valid software kernels)]). Regarding claim 16, as discussed above, Barker in view of Liu discloses the method of claim 14. Barker further discloses wherein the one or more input parameters include one of the following: a plurality of tensors, an attribute of a mathematical operation or function, a data attribute or an attribute of a tensor descriptor (see, e.g., paragraph [0017]: “Each kernel may be associated with a particular computation and may be utilized with provided input data to produce a result. For example, a kernel stored in database 105 may be associated with a general matrix multiplication (GeMM) computation. The GeMM kernel may then be utilized by an execution component, such as kernel processor 104, which can perform the computation utilizing the GeMM kernel” [i.e., the input parameters used in the candidate kernel generator includes a mathematical operation /function (GeMM)] and paragraph [0069]: In at least one embodiment, an untrained neural network is trained using a training dataset. In at least one embodiment, a training framework is a PyTorch framework, Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit/CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework” [i.e., the training frameworks inherently include tensor operations, mathematical functions and tensor attributes (such as shape/dimension, data type, and/or memory layout)]). Regarding Claim 24, as discussed above, Barker in view of Liu discloses the apparatus of claim 1. Barker further discloses wherein the first processor and the second processor are the same processor(see, e.g. paragraph [0020]: “Filter 108 removes any kernels from the list of candidate kernels that are not practical and/or impossible to execute given particular restraints of the system” [i.e., the kernels are validated by the Filter 108], paragraph [0026]: “FIG. 3 is an example environment that may be utilized to train a neural network to rank kernels to perform a computation. The training system 300 can be the same system as illustrated in FIG. 1 as training system 120. The training system includes a kernel processor 302, which may be the same, or share characteristics with kernel processor 104 of FIG. 1”, paragraph [0028]: “For each training input, each of the plurality of kernels stored in kernel database 308 may be provided to the kernel processor 302 to perform a computation. For example, kernel database may include kernels K1 . . . Kn, each of which can be provided to kernel processor 302 (or provided to a plurality of kernel processors) to generate a result set” and paragraph [0039]: “A relevancy score can be indicative of how likely a kernel is an optimal kernel for performing the computation. At step 420, the optimal kernel is selected” [i.e., the training system (trained ML model engine) includes a kernel processor 302, which receives a plurality of valid software kernels in a kernel selection process]). Regarding claim 25, as discussed above, Barker in view of Liu discloses the non-transitory computer-readable medium of claim 29. Barker further discloses further comprising instructions for causing the computer to generate the plurality of valid software kernels based on one or more input parameters and a plurality of software kernels (see, e.g. paragraph [0019]: “The candidate kernel generator 106 determines one or more candidate kernels that may be utilized by the kernel processor 104 to determine a result computation. The candidate kernel generator 106 identifies the kernels in database 102 that are available to execute the operation that has been received in a request as well as characteristics of the input data.” [i.e., the components in the kernel selection system receive both input data (input parameters) and software kernels from the database]). Regarding claim 26, as discussed above, Barker in view of Liu discloses the non-transitory computer-readable medium of claim 25. Barker further discloses wherein the one or more input parameters include one of the following: a plurality of tensors, an attribute of a mathematical operation or function, a data attribute or an attribute of a tensor descriptor (see, e.g., paragraph [0017]: “Each kernel may be associated with a particular computation and may be utilized with provided input data to produce a result. For example, a kernel stored in database 105 may be associated with a general matrix multiplication (GeMM) computation. The GeMM kernel may then be utilized by an execution component, such as kernel processor 104, which can perform the computation utilizing the GeMM kernel” [i.e., the input parameters used in the candidate kernel generator includes a mathematical operation /function (GeMM)] and paragraph [0069]: In at least one embodiment, an untrained neural network is trained using a training dataset. In at least one embodiment, a training framework is a PyTorch framework, Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit/CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework” [i.e., the training frameworks inherently include tensor operations, mathematical functions and tensor attributes (such as shape/dimension, data type, and/or memory layout)]). Regarding Independent Claim 29, Barker discloses A non-transitory computer-readable medium storing computer executable code, operable on a device comprising at least one processor and at least one memory coupled to the at least one processor, wherein the at least one processor is configured to implement kernel selection (see, e.g., paragraph [0086]: “In at least one embodiment, code is stored on a computer-readable storage medium, for example, in form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transitory signals (e.g., a propagating transient electric or electromagnetic transmission) but includes non-transitory data storage circuitry (e.g., buffers, cache, and queues) within transceivers of transitory signals”, paragraph [0026]: “FIG. 3 is an example environment that may be utilized to train a neural network to rank kernels to perform a computation. The training system 300 can be the same system as illustrated in FIG. 1 as training system 120. The training system includes a kernel processor 302, which may be the same, or share characteristics with kernel processor 104 of FIG. 1”, paragraph [0028]: “For each training input, each of the plurality of kernels stored in kernel database 308 may be provided to the kernel processor 302 to perform a computation. For example, kernel database may include kernels K1 . . . Kn, each of which can be provided to kernel processor 302 (or provided to a plurality of kernel processors) to generate a result set” and paragraph [0039]: “A relevancy score can be indicative of how likely a kernel is an optimal kernel for performing the computation. At step 420, the optimal kernel is selected” [i.e., the training system (trained ML model engine) includes a kernel processor 302, which receives a plurality of valid software kernels in a kernel selection process]), the computer executable code comprising: instructions for causing a computer to input a plurality of valid software kernels to a trained machine learning (ML) model engine (see, e.g., paragraph [0021]: “Once a candidate list of kernels has been determined, the neural network 110 is provided the list along with the input data. The neural network is trained to determine, based on the list of candidate kernels, a relevancy score for each of the kernels that is predictive of how a given kernel will perform a computation on the input data. The neural network may be trained by a training system 120” [i.e., the trained neural network (trained ML model) is generated via training by the training system 120 (trained ML model engine)]); and instructions for causing the computer to use the first trained ML model to generate a machine learning (ML)-selected software kernel based on the plurality of valid software kernels (see, e.g., paragraph [0021]: “Once a candidate list of kernels has been determined, the neural network 110 is provided the list along with the input data. The neural network is trained to determine, based on the list of candidate kernels, a relevancy score for each of the kernels that is predictive of how a given kernel will perform a computation on the input data” and paragraph [0022]: “Sorter 112 sorts the candidate kernels by relevancy based on the output of the neural network. For example, for a given list of kernels with relevancy scores, sorter 112 can sort the list so that the first kernel in the list is the most relevant kernel for performing a computation on the input data. Selection engine 114 then chooses a kernel from the sorted list, such as the kernel with the highest relevancy score” [i.e., under the broadest reasonable interpretation, the neural network (trained ML Model) generates a (ML)-selected software kernel by determining ML relevancy score data that results in a selected kernel from the coupled selection engine]). Barker does not disclose: and with a shorter timing performance than a software kernel, wherein the software kernel includes at least one static optimization parameter. However, Liu teaches and with a shorter timing performance than a software kernel, wherein the software kernel includes at least one static optimization parameter(Liu, pgs., 174-184, see also figs. 2, 9, 10, and 13,“To clearly show the building process of our model, we illustrated it as a 4-step workflow, as shown in Fig.2. The 4 steps are: 1. Kernel classification. We train a classification tree to predict the resource bound type of a target kernel...[w]e propose polynomial models for Static metrics which are used in our model... in group(a) in Table I... we try to find the relationship between convolution hyper-parameters and Static metrics[wherein the software kernel includes at least one static optimization parameter]... [b]y using SEER as algorithm picker, we can save 44.58% time on this operator, and 11.37% time on the overall forward computation time[and with a shorter timing performance than a software kernel].”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Barker with the teachings of Liu the motivation to do so would be optimize the execution of a deep learning model by using a machine learning based model to select the best kernel for executing operations for the deep learning model(Liu, pg., 174, “We propose SEER, a hybrid time prediction model for CNNs, which combines both analytical models and learning-based models... [w]e show that SEER can achieve low prediction error rate of 14.71% at op-level and 1.79% at network-level. When used for choosing best convolution algorithm, SEER outperforms cuDNN official algorithm picker, with 7.14% lower error rate.”). Claims 5-7, 17-23, 27-28 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Barker in view of Liu et al., "SEER: A time prediction model for CNNs from GPU kernel's view." 2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT). IEEE, 2021(“Liu”) and in view of Jakubiuk (US 20210295158 A1; hereinafter Jakubiuk). Regarding claim 5, as discussed above, Barker in view of Liu discloses the apparatus of claim 4. However, Barker in view of Liu fails to explicitly teach wherein the second processor is further configured to accept the one or more input parameters and the optimal software kernel from the training data repository Nevertheless, in the same field, analogous art Jakubiuk teaches wherein the second processor is further configured to accept the one or more input parameters (see, e.g., Jakubiuk paragraph [0033]: “As shown in FIG. 2, In various embodiments, the Optimization Engine receives as input a plurality of characteristics 202 of target AI network (such as a neural network) and generates output 210… The Optimization Engine executes according to a training phase and an optimization phase. During the training phase, the Optimization Engine trains a heuristic AI network according to various types of heuristic training data, as shown in FIG. 6. In the optimization phase, the Optimization Engine receives the plurality of characteristics 202 of the target AI network”, paragraph [0035]: “Upon selection of the heuristic artificial intelligence network model during the optimization phase…” and paragraph [0051]: “The heuristic training data 124 may also include various types of operation hyper-parameters 604-6” [i.e., the optimization engine functions as an ML model selection engine that accepts the input parameters from heuristic training data]) and the optimal software kernel from the training data repository (see, e.g., paragraph [0035]: “Upon selection of the heuristic artificial intelligence network model during the optimization phase, the heuristic AI network model module 206 may utilize the dynamic auto-tune module 208 to identify the run times of various operation combinations and kernel implementations” [i.e., the optimization engine accepts the optimized kernel implementations from the dynamic auto-tune module 208], paragraph [0051]: “The heuristic training data 124 may also include various types of operation hyper-parameters 604-6” and paragraph [0052]: “A non-limiting exemplary list of operation hyper-parameters includes: kernel width/height, input depth, one or more input/output features” [i.e., the heuristic training data 124 functions as the training data repository that includes kernel run times of various kernel implementations (optimal software kernel)]). Barker in view of Liu and Jakubiuk are analogous art because they are both directed to kernel optimization using machine learning techniques. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to utilize a machine learning model selection engine that accepts one or more input parameters and an optimal software kernel from a training data repository. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “create a new and useful system and method for providing optimized software implementations of artificial intelligence networks”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraph [0006]). Regarding claim 6, as discussed above, Barker in view of Liu and Jakubiuk teaches the apparatus of claim 5. Jakubiuk teaches wherein the second processor is further configured to generate a second trained machine learning (ML) model (see, e.g., Jakubiuk paragraph [0051]: “During a training phase, the heuristic AI network training module 110 trains a heuristic AI network 130 according to various types of heuristic training data 124, as shown in FIG. 6” [i.e., the heuristic AI network training module 110 (first ML model), which is a component of the Optimization engine (model selection engine), generates (via training) a heuristic AI network 130 (second ML model)]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to generate a second trained machine learning model. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “result in efficient AI networks which are deployable across many different hardware components and platforms”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraph [0006]). Regarding claim 7, as discussed above, Barker in view of Liu and Jakubiuk teaches the apparatus of claim 6. Jakubiuk teaches wherein the ML model selection engine is further configured to tune the second trained ML model to generate a tuned machine learning (ML) model (see, e.g., Jakubiuk paragraphs [0051-0052]: “During a training phase, the heuristic AI network training module 110 trains a heuristic AI network 130 according to various types of heuristic training data 124, as shown in FIG. 6. The heuristic training data 124 may include… various types of operation hyper-parameters 604-6… A non-limiting exemplary list of operation hyper-parameters includes: kernel width/height, input depth, one or more input/output features, padding type, padding dimensions, strides, dilations, epsilons and thresholds… A non-limiting exemplary list of workload types, i.e., varieties of optimization functions, includes: latency-optimized inference, throughput-optimized inference, latency-bound throughput” [i.e., the heuristic AI network 130 (second ML model) is generated using hyperparameters and optimized inference algorithms which is tuning the model to produce a tuned ML model]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to generate a second trained machine learning model. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “result in efficient AI networks which are deployable across many different hardware components and platforms”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraph [0006]). Regarding claim 17, as discussed above, Barker in view of Liu discloses the method of claim 11. However, Barker in view of Liu fails to explicitly teach further comprising providing one or more input parameters and an optimal software kernel to a machine learning (ML) model selection engine from a training data repository. Nevertheless, in the same field, analogous art Jakubiuk teaches further comprising providing one or more input parameters and an optimal software kernel to a machine learning (ML) model selection engine from a training data repository (see, e.g., Jakubiuk paragraph [0033]: “As shown in FIG. 2, In various embodiments, the Optimization Engine receives as input a plurality of characteristics 202 of target AI network (such as a neural network) and generates output 210… The Optimization Engine executes according to a training phase and an optimization phase. During the training phase, the Optimization Engine trains a heuristic AI network according to various types of heuristic training data, as shown in FIG. 6. In the optimization phase, the Optimization Engine receives the plurality of characteristics 202 of the target AI network”, paragraph [0035]: “Upon selection of the heuristic artificial intelligence network model during the optimization phase…”, paragraph [0051]: “The heuristic training data 124 may also include various types of operation hyper-parameters 604-6” [i.e., the optimization engine functions as an ML model selection engine that accepts the input parameters from heuristic training data] and paragraph [0052]: “A non-limiting exemplary list of operation hyper-parameters includes: kernel width/height, input depth, one or more input/output features” [i.e., the heuristic training data 124 functions as the training data repository that includes kernel run times of various kernel implementations (optimal software kernel)]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to generate a second trained machine learning model. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “result in efficient AI networks which are deployable across many different hardware components and platforms”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraph [0006]). Regarding claim 18, as discussed above, Barker in view of Liu and Jakubiuk teaches the method of claim 17. Jakubiuk teaches further comprising configuring the machine learning (ML) model selection engine to generate a second trained machine learning (ML) model (see, e.g., Jakubiuk paragraph [0051]: “During a training phase, the heuristic AI network training module 110 trains a heuristic AI network 130 according to various types of heuristic training data 124, as shown in FIG. 6” [i.e., the heuristic AI network training module 110 (first ML model), which is a component of the Optimization engine (model selection engine), generates (via training) a heuristic AI network 130 (second ML model)]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to generate a second trained machine learning model. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “result in efficient AI networks which are deployable across many different hardware components and platforms”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraph [0006]). Regarding claim 19, as discussed above, Barker in view of Liu and Jakubiuk teaches the method of claim 18. Jakubiuk teaches further comprising tuning the second trained ML model by using a training data from the training data repository to generate a tuned machine learning (ML) model (see, e.g., Jakubiuk paragraphs [0051-0052]: “During a training phase, the heuristic AI network training module 110 trains a heuristic AI network 130 according to various types of heuristic training data 124, as shown in FIG. 6. The heuristic training data 124 may include… various types of operation hyper-parameters 604-6… A non-limiting exemplary list of operation hyper-parameters includes: kernel width/height, input depth, one or more input/output features, padding type, padding dimensions, strides, dilations, epsilons and thresholds… A non-limiting exemplary list of workload types, i.e., varieties of optimization functions, includes: latency-optimized inference, throughput-optimized inference, latency-bound throughput” [i.e., the heuristic AI network 130 (second ML model) is generated using hyperparameters and optimized inference algorithms from training data 124 (training data repository), which is tuning the model to produce a tuned ML model]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to generate a second trained machine learning model. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “result in efficient AI networks which are deployable across many different hardware components and platforms”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraph [0006]). Regarding claim 20, as discussed above, Barker in view of Liu and Jakubiuk teaches the method of claim 19. Barker further teaches further comprising using the tuned ML model in a kernel selection process in a kernel selection engine based on machine learning (ML) (see, e.g., Barker paragraph [0030]: “Relevancy calculator 310 determines a relevancy score for each of the kernels in the result set based on the run-time of the kernel while executing the matrix computation… The result set may then be ranked according to relevancy scores for further processing” and paragraph [0031]: “The result set and the training input can be provided to the neural network as training data. The neural network may then process the input to determine a predicted relevancy score for each of the kernels” [i.e., the relevancy calculator and neural network determine a relevancy score for the kernels and ranks the kernels (outputting optimal software kernels), which functions as a kernel selection engine]). Regarding claim 21, as discussed above, Barker in view of Liu and Jakubiuk teaches the method of claim 20. Barker further teaches further comprising supplying a plurality of performance metrics to the kernel selection engine (see, e.g., Barker paragraph [0030]: “Relevancy calculator 310 determines a relevancy score for each of the kernels in the result set based on the run-time of the kernel while executing the matrix computation… Thus, a kernel with a relevancy score of 0.25 may have performed worse (e.g., taken longer to run) than a kernel that is assigned a relevancy score of 0.75. The result set may then be ranked according to relevancy scores for further processing” [i.e., the relevancy calculator component in the training system receives the result set of performance metrics from the kernel processor component for scoring and ranking the kernels]). Regarding claim 22, as discussed above, Barker in view of Liu and Jakubiuk teaches the method of claim 21. Barker further teaches further comprising configuring the kernel selection engine to implement a selection function for each of the plurality of valid software kernels to determine the optimal software kernel (see, e.g., Barker paragraph [0020]: “Filter 108 removes any kernels from the list of candidate kernels that are not practical and/or impossible to execute given particular restraints of the system” [i.e., valid kernels are produced by the Filter 108 and transmitted to adjacent components for further computation] and paragraph [0030]: “Relevancy calculator 310 determines a relevancy score for each of the kernels in the result set based on the run-time of the kernel while executing the matrix computation… The result set may then be ranked according to relevancy scores for further processing” [i.e., the relevancy calculator determines a relevancy score for the kernels and ranks the kernels (determining optimal software kernels), which is functionally a selection function performed on each of the valid kernels by the combined neural network and relevancy calculator components (kernel selection engine)]). Regarding claim 23, as discussed above, Barker in view of Liu and Jakubiuk teaches the method of claim 21. Barker further teaches further comprising generating the plurality of performance metrics based on the plurality of valid software kernels (see, e.g., Barker paragraph [0043]: “At step 510, a result set is generated for the training input. The result set includes one or more kernels that can be utilized to perform a computation on the result set… At step 515, a relevancy score is assigned to each kernel in the result set that is indicative of the quality of performance of that kernel in performing the computation. Thus, kernels with faster run-times may be assigned a higher score than kernels that performed slower. At step 520, the kernels can be ranked based on run-times and/or assigned relevancy scores” [i.e., kernels that can be utilized to perform a computation on a result set are valid software kernels, the relevancy scores (performance metrics) are then generated based on the plurality of valid software kernels]). Regarding claim 27, as discussed above, Barker in view of Liu discloses the non-transitory computer-readable medium of claim 29. However, Barker in view of Liu fails to explicitly teach further comprising instruction for causing the computer to generate a second trained machine learning (ML) model. Nevertheless, in the same field, analogous art Jakubiuk teaches instruction for causing the computer to generate a second trained machine learning (ML) model(see, e.g., Jakubiuk paragraph [0051]: “During a training phase, the heuristic AI network training module 110 trains a heuristic AI network 130 according to various types of heuristic training data 124, as shown in FIG. 6” [i.e., the heuristic AI network training module 110 (first ML model), which is a component of the Optimization engine (model selection engine), generates (via training) a heuristic AI network 130 (second ML model)]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to generate a second trained machine learning model. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “result in efficient AI networks which are deployable across many different hardware components and platforms”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraph [0006]). Regarding claim 28, as discussed above, Barker in view of Liu teaches the non-transitory computer-readable medium of claim 29. However, Barker in view of Liu fails to explicitly teach further comprising instructions for causing the computer to tune a second trained ML model by using a training data from a training data repository to generate a tuned machine learning (ML) model. Nevertheless, in the same field, analogous art Jakubiuk teaches further comprising means for tuning the second trained ML model by using a training data from the training data repository to generate a tuned machine learning (ML) model (see, e.g., Jakubiuk paragraphs [0051-0052]: “During a training phase, the heuristic AI network training module 110 trains a heuristic AI network 130 according to various types of heuristic training data 124, as shown in FIG. 6. The heuristic training data 124 may include… various types of operation hyper-parameters 604-6… A non-limiting exemplary list of operation hyper-parameters includes: kernel width/height, input depth, one or more input/output features, padding type, padding dimensions, strides, dilations, epsilons and thresholds… A non-limiting exemplary list of workload types, i.e., varieties of optimization functions, includes: latency-optimized inference, throughput-optimized inference, latency-bound throughput” [i.e., the heuristic AI network 130 (second ML model) is generated using hyperparameters and optimized inference algorithms from training data 124 (training data repository), which is tuning the model to produce a tuned ML model]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to tune a second trained ML model using training data from a training data repository to generate a tuned machine learning model using a model selection engine. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “simulate[s] the execution of various combinations of operations and kernels and other algorithms and components over the entire structure of the artificial intelligence network to identify operation, kernel, algorithm and component combinations that optimize the performance of the entire artificial intelligence network”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraph [0007]). Regarding claim 30, as discussed above, Barker in view of Liu discloses the non-transitory computer-readable medium of claim 29. However, Barker in view of Liu fails to explicitly teach further comprising instructions for causing the computer to generate a second trained machine learning (ML) model and to tune the second trained ML model by using a training data to generate a tuned machine learning (ML) model. Nevertheless, in the same field, analogous art Jakubiuk teaches further comprising instructions for causing the computer to generate a second trained machine learning (ML) model (see, e.g., Jakubiuk paragraph [0051]: “During a training phase, the heuristic AI network training module 110 trains a heuristic AI network 130 according to various types of heuristic training data 124, as shown in FIG. 6” [i.e., the heuristic AI network training module 110 (first ML model), which is a component of the Optimization engine (model selection engine), generates (via training) a heuristic AI network 130 (second ML model)]). and to tune the second trained ML model by using a training data to generate a tuned machine learning (ML) model (see, e.g. Jakubiuk paragraphs [0051-0052]: “During a training phase, the heuristic AI network training module 110 trains a heuristic AI network 130 according to various types of heuristic training data 124, as shown in FIG. 6. The heuristic training data 124 may include… various types of operation hyper-parameters 604-6… A non-limiting exemplary list of operation hyper-parameters includes: kernel width/height, input depth, one or more input/output features, padding type, padding dimensions, strides, dilations, epsilons and thresholds… A non-limiting exemplary list of workload types, i.e., varieties of optimization functions, includes: latency-optimized inference, throughput-optimized inference, latency-bound throughput” [i.e., the heuristic AI network 130 (second ML model) is generated using hyperparameters and optimized inference algorithms from training data 124 (training data repository), which is corresponding to tuning the model to produce a tuned ML model]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Barker in view of Liu to incorporate the teachings of Jakubiuk to generate a second trained machine learning model and tune the second trained ML model using data from a training data repository to generate a tuned machine learning model using a model selection engine. Doing so would have allowed Barker in view of Liu to use Jakubiuk’s method in order to “result in efficient AI networks which are deployable across many different hardware components and platforms” and “simulate[s] the execution of various combinations of operations and kernels and other algorithms and components over the entire structure of the artificial intelligence network to identify operation, kernel, algorithm and component combinations that optimize the performance of the entire artificial intelligence network”, as suggested by Jakubiuk (see, e.g., Jakubiuk, paragraphs [0006-0007]). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADAM C STANDKE whose telephone number is (571)270-1806. The examiner can normally be reached Gen. M-F 9-9PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Adam C Standke/ Primary Examiner Art Unit 2129 1 As discussed above in the section 112(b) rejection of this claim, for examination purposes the term “valid software kernels” is interpreted as any plurality of software kernels that are functionally viable/practically executable for a given task or purpose. 2 As discussed above in the 112(f) interpretation of this claim, the term “data repository” has been interpreted as covering the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. 3 As discussed above in the section 112(b) rejection of this claim, for examination purposes the term “valid software kernels” is interpreted as any plurality of software kernels that are functionally viable/practically executable for a given task or purpose.
Read full office action

Prosecution Timeline

Mar 22, 2023
Application Filed
Nov 19, 2025
Non-Final Rejection mailed — §103, §112
Feb 03, 2026
Response Filed
Aug 10, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743614
METHOD FOR TRAINING ASYMMETRIC GENERATIVE ADVERSARIAL NETWORK TO GENERATE IMAGE AND ELECTRIC APPARATUS USING THE SAME
5y 0m to grant Granted Sep 22, 2026
Patent 12725015
RECONFIGURABLE PREDICTION ENGINE FOR GENERAL PROCESSOR COUNTING
8y 5m to grant Granted Sep 01, 2026
Patent 12725026
RUNTIME OPTIMIZATION OF COMPUTATIONS OF AN ARTIFICIAL NEURAL NETWORK COMPILED FOR EXECUTION ON A DEEP LEARNING ACCELERATOR
5y 10m to grant Granted Sep 01, 2026
Patent 12717817
SYSTEMS AND METHODS FOR TIME-BASED ABNORMALITY IDENTIFICATION WITHIN UNIFORM DATASET
7y 7m to grant Granted Aug 25, 2026
Patent 12718145
Method for correcting bias introduced by weighted training in machine learning
3y 12m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
53%
Grant Probability
79%
With Interview (+26.6%)
4y 4m (~9m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 146 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month