Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Detailed Action
2. This Non-Final Office Action is responsive to Applicants’ arguments as received 5/8/26. Claims 1-3, 5-14, and 16-22 remain pending, of which claims 1 and 12 are independent.
Claim Rejections - 35 USC § 112
3. The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
4. Claims 1-3, 5-14, and 16-22 are rejected under 35 U.S.C. 112(a) as failing to comply with the written description requirement. The claims contain subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Taking independent claim 1 as representative, the claim recites, in part, a limitation for “utilizing a value of the hyperparameter from the previously trained machine learning model to configure an initial probability distribution of a search model defined over a search space.” Upon close review of Applicants specification and varying keyword searches, the Examiner cannot find any concrete teaching, and hence support, of the following terms that are recited in the aforementioned limitation: “initial probability distribution”, “search model”, “probability distribution”, and “distribution.” The Examiner believes Applicants are attempting to capture some of the subject matter disclosed in [0029]-[0035] in the published version of the specification but are using terminology via the claim that are not consistent with the teachings found in the specification. Applicants are advised to point out where the specification support is for this limitation, particularly with regards to the recited feature of “an initial probability distribution of a search model defined over a search space”, and further if the same terms are not used in the specification, then to at least offer a clarifying explanation as to what the equivalent teachings are in the specification that correspond with these same terms per the claim limitation.
Independent claim 12 features the same limitation and is hence rejected under the same rationale for lacking support in the specification. The dependent claims 2-3, 5-11, 13-14, and 16-22 each depend from one of the two aforementioned independent claims, and therefore inherit the deficiency explained here without otherwise curing it. Hence, they too are rejected under the same rationale.
Upon a proper showing of support by Applicants, the Examiner will reconsider the rejection and withdraw it if deemed appropriate.
5. The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
6. Claim 7 is rejected under 35 U.S.C. 112(b) as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor regards as the invention. In particular, claim 7 features a limitation for “wherein providing the performance of each of the trained unique machine learning models to the user device comprises providing, to the user device, an indication indicating which trained unique machine learning model has the best performance based on the training data.” The claim appears to provide a further clarification, in italics provided here, for a broader limitation as bolded here. However, upon close review of the claims, there is no prior mention or recitation of the would-be clarified bolded language either in the present claim 7 of in independent claim 1 from which it depends. Hence, on this basis, it is not clear whether any of these limitations as detailed in this present claim are actually required. For example, one way of understanding this claim is that a clarification of a claim limitation that is not actively recited or even previously presented would then not appear to be a clarification of subject matter that is actively/even required. The ambiguity explained here renders the claim vague and indefinite.
The Examiner contrasts this with the recitation of similar language found in claim 18, and notes that in claim 18 the clarification actively requires the further limitation.
Claim Rejections - 35 USC § 103
7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office Action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
8. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
9. Claims 1-2, 5-7, 9-13, 16-18, and 20-22 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Patent Application Publication No. 2018/0240041 (“Koch”) in view of Non-Patent Literature “Hyperparameter Importance Across Datasets” (“van Rijn”) and further in view of Non-Patent Literature “Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads” (“Narayanan”).
Regarding claim 1, KOCH teaches A computer-implemented method (FIG. 3 teaching a block diagram for a selection manager device, which features modules 314-322 that are used to perform data processing relating to the automatic selection of hyperparameters to train a predictive model (as summarized per the Abstract and as shown in more detail per FIGs. 5 and especially 6A-6C for example), where the selection manager device of FIG. 3 includes a processor (FIG. 3’s element 310)) comprising:
receiving, by data processing hardware (FIG. 3 element 310) and from a user device (a requesting user and corresponding device for that user, per FIG. 2 and [0033]), a hyperparameter optimization request requesting optimization of one or more hyperparameters of a machine learning model (the aforementioned requesting user and corresponding device issues the request of FIG. 5 step 528, which the Examiner equates with the recited “hyperparameter optimization request” which is understood to relate to a model to be tuned accordingly (per FIG. 5 and [0067] for example));
obtaining, by the data processing hardware (FIG. 3 element 310), training data for training the machine learning model (FIG. 5 step 506, where a user provides indication of an “input dataset”, which the Examiner understands to be a basis for training data selection ([0061], [0160]) in accordance with training and tuning a model as previously discussed just above and to be further discussed just below);
determining, by the data processing hardware (FIG. 3 element 310), a set of hyperparameter permutations of the one or more hyperparameters of the machine learning model ... and for each respective hyperparameter permutation in the set of hyperparameter permutations:
training, by the data processing hardware (FIG. 3 element 310) ... , a unique machine learning model using the training data and the respective hyperparameter permutation; and determining, by the data processing hardware (FIG. 3 element 310), a performance of the trained unique machine learning model ([0152] discussing the determination of a configuration list of hyperparameter configurations to be evaluated (i.e., akin to the recited “set of hyperparameter permutations”), which are selected for a particular model, where the hyperparameter configurations are iteratively selected and assigned to a session per FIG. 6B and [0169] (i.e., corresponding to the recitations for “training” and “determining a performance of ... model” that is required “for each respective hyperparameter permutation”), such that the model is trained and scored with respect to each hyperparameter configuration ([0171]), where at the end of the iterative evaluation as described there is a result (FIG. 6C step 672, which clarifies FIG. 5 steps 530-532 and serves as a basis for a hyperparameter selection for further model training/evaluation per FIG. 5 step 534));
selecting, by the data processing hardware (FIG. 3 element 310) and based on the performance of each of the trained unique machine learning models, one of the trained unique machine learning models and generating, by the data processing hardware (FIG. 3 element 310), one or more predictions using the selected one of the trained unique machine learning models (FIG. 5’s step 532 is a hyperparameter selection, e.g. based on its performance for training a particular model with a particular training dataset, and then is used for a new dataset per step 534 (which the Examiner equates with the prediction generation as recited)).
The Examiner notes that the limitation discussed above for determining a set of hyperparameter permutations of the one or more hyperparameters of the machine learning model is further clarified by way of further limitations to be performed by at least:
identifying at least one previously trained machine learning model that shares a hyperparameter with the machine learning model and
utilizing a value of the hyperparameter from the previously trained machine learning model to configure an initial probability distribution of a search model defined over a search space, and
searching the defined search space.
Regarding the identifying limitation, see Koch: [0098], clarifying FIG. 5’ step 518, such that a hyperparameter configuration to be evaluated / used for training may be permitted or subject to discarding or revision (i.e., an identifying step) based on a similarity tolerance to a prior configuration, where this comparison is performed on a per-hyperparameter basis, and hence from this the Examiner understands that across the configurations to be evaluated, the hyperparameters included therein may be same/similar and hence subject to a reuse condition (which the Examiner equates to the recited shared concept)). In the Examiner’s view, this at least encompasses that prior configurations exist and serve as a basis for comparison with the present model’s hyperparameters, which necessarily involves an identification such that the identified element can serve as the basis for this comparison as taught.
Moreover, it would seem that the finding of these other configurations and hyperparameters, which serves as a basis for identifying and comparing, is a searching per the limitation for “searching.” Koch, for example, teaches at [0167]: “For illustration, the LHS, the Random, and/or the Grid search methods may be used in a first iteration to define the first set of hyperparameter configurations that sample the search space.”
However, the Examiner notes that the limitation for utilizing a value of the hyperparameter from the previously trained machine learning model to configure an initial probability distribution of a search model defined over a search space is not so clearly taught by Koch. Rather, the Examiner turns to VAN RIJN to teach what Koch otherwise lacks, see e.g., van Rijn’s comparable framework for hyperparameter optimization, in which the consideration of “priors” (other models with similar features) and their performances are utilized to determine similarities between models, the hyperparameters that are same/similar to both, and then model performance, such that when there is an appropriate fit, values from those other similar and well performing models can be used to facilitate warm start and more efficient search space evaluation (section 2’s subheadings for Hyperparameter Importance and Priors, specifically, as found on page 2 of the reference, with further clarification found in sections 4.1 and 4.2). Section 4.3 is also noted here for its teaching of the framework’s consideration of other models, known hyperparameter configurations, and their performance data, which serves as the informational basis to perform its hyperparameter importance and priors evaluation steps. The Examiner understands this framework, based on these cited sections, to essentially look to the repository mentioned in section 4.3 for performance information for previously trained models in terms of their known hyperparameter configurations, and to then use that information to select an important hyperparameter and determine values from it (sections 4.1-4.2) as applied to whatever problem/subject model is being optimized in terms of its training. Based on this understanding, van Rijn reads on claim’s identifying..., utilizing... , and searching... limitations as being addressed here, particularly in view of the possibility of this improved manner of hypermeter optimization through improved search, per van Rijn, being used to improve the existing search aspect more generally taught per Koch.
Like Koch, van Rijn relates to hyperparameter optimization for a particular model/task, in part based on a hyperparameter search step. Hence, they are similarly directed and therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the particular search aspect per van Rijn, including its importance and priors considerations as taught, into Koch’s more generalized user-facing framework, with a reasonable expectation of success, such that van Rijn’s importance and priors considerations as discussed above could be understood to provide for a more efficient search.
Koch etc. does not teach the further limitation of performing the training as discussed above based on a priority order of the machine learning model. Rather, the Examiner relies upon NARAYANAN to teach what Koch otherwise lacks, see e.g., Narayanan’s framework having a scheduling policy for deep learning (Abstract and Introduction sections, per page 481), as depicted generally via Figure 2 on page 483, where the system manages the training of different jobs (section 3 on page 483), e.g. using a scheduling mechanism as discussed per section 3.2 on page 485 that involves an explicit “priority score for every job” to facilitate the scheduling in accordance with a priority order for the jobs.
Both Koch and Narayanan relate to deep learning frameworks and the resource management thereof. Hence, they are similarly directed and therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate a job priority for a deep learning task, per Narayanan, into a multi-task/job framework as contemplated by either/both reference, such that the available compute resources can be strategically applied to satisfy the needs of not just one deep learning task/job but many, thereby improving efficiency and throughput for a compute paradigm that is often resource constrained and benefits from being smartly resource aware.
Regarding claim 2, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references further teach the additional limitation wherein determining the set of hyperparameter permutations comprises performing a search on a hyperparameter search space of the one or more hyperparameters of the machine learning model (Koch’s [0167] teaches its framework’s capability to perform a hyperparameter search over a defined search space, and the Examiner emphasizes what is believed to be an improved and more efficient search as taught by van Rijn, which leverages steps for determining hyperparameter importance and priors information (as discussed per claim 1), to essentially better confine the search and evaluation steps to less than all Koch’s less finely targeted methods). The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 5, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references further teach the additional limitation wherein the at least previously trained machine learning model is trained for a user of the user device (the prior iterations of training and hyperparameter model evaluation are all associated with a common user involved with Koch’s FIG. 5 teachings). The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 6, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references further teach the additional limitation wherein training the unique machine learning model comprises training two or more unique machine learning models in parallel (Koch: [0067], [0070], [0157], and [0231] teaching concurrency and parallel execution advantages in facilitating the steps of FIGs. 5 and 6A-6C, such that [0150]’s workers and sessions can be understood to be working in parallel). The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 7, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references further teach the additional limitation wherein providing the performance of each of the trained unique machine learning models to the user device comprises providing, to the user device, an indication indicating which trained unique machine learning model has the best performance based on the training data (Koch: [0073]: “one or more of the output tables may be selected by the user for presentation on display”, as referring to the different output tables discussed therein same paragraph, and also [0147] providing further clarification of tuning evaluation results as subject to a display/presentation, e.g. per FIG. 5’s step 530). The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 9, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references further teach the additional limitation wherein the hyperparameter optimization request comprises a budget and a size of the set of hyperparameter permutations of the one or more hyperparameters of the machine learning model is based on the budget (Koch: [0156] discussing a user and/or administrator’s capability to adjust a size constrain for evaluation per FIG. 6 step 604). The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 10, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references further teach the additional limitation wherein the data processing hardware is part of a distributed computing database system (Koch: [0066]: “... the input dataset may be stored in a cube distributed across the computing devices of each session that is a grid of computers as understood by a person of skill in the art ...”, e.g. to facilitate the management and use of the many worker computers shown per FIG. 1 and described per [0030] and [0036]). The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 11, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references further teach the additional limitation wherein selecting the one of the trained unique machine learning models comprises: transmitting the performance of each of the trained unique machine learning models to the user device, and receiving, from the user device, a trained unique machine learning model selection selecting the one of the trained unique machine learning models (Koch: FIG. 5’s step 530 and 532 respectively). The motivation for combining the references is as discussed above in relation to claim 1.
Regarding claim 12, the claim includes the same or similar limitations as claim 1 discussed above and is therefore rejected under the same rationale.
Regarding claim 13, the claim includes the same or similar limitations as claim 2 discussed above and is therefore rejected under the same rationale.
Regarding claim 16, the claim includes the same or similar limitations as claim 5 discussed above and is therefore rejected under the same rationale.
Regarding claim 17, the claim includes the same or similar limitations as claim 6 discussed above and is therefore rejected under the same rationale.
Regarding claim 18, the claim includes the same or similar limitations as claim 7 discussed above and is therefore rejected under the same rationale.
Regarding claim 20, the claim includes the same or similar limitations as claim 9 discussed above and is therefore rejected under the same rationale.
Regarding claim 21, the claim includes the same or similar limitations as claim 10 discussed above and is therefore rejected under the same rationale.
Regarding claim 22, the claim includes the same or similar limitations as claim 11 discussed above and is therefore rejected under the same rationale.
10. Claims 3 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Koch in view of van Rijn and further in view of Narayanan and further yet in view of Non-Patent Literature “Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization” (“Li”).
Regarding claim 3, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references teach a hyperparameter search and corresponding space, as discussed per claim 2 (citing Koch’s [0167]). Further, the search may be Gaussian as discussed per Koch’s [0139] (“For illustration, the Bayesian search method is based on creating and exploring a kriging surrogate model to search for improved solutions. A Kriging model is a type of interpolation algorithm for which the interpolated values are modeled by a Gaussian process governed by prior covariance values.”). Van Rijin, similarly, teaches a hyperparameter search space that is more targeted using importance and priors considerations, as discussed per claim 1 and citing to its section. Hence, Koch and van Rijn may already teach the entirety of the further limitation wherein performing the search on the hyperparameter search space comprises performing the search using a batched Gaussian process bandit optimization. However, to the extent that they do not sufficiently teach “a batched Gaussian process bandit optimization” for the search, as recited, the Examiner then relies upon LI to teach what Koch etc. otherwise lacks, see e.g., Li’s section 2.1 beginning on page 3, discussing Gaussian processes to model and sample hyperparameters in pursuit of hyperparameter optimization (citing to Spearmint) and more explicitly (on page 4’s second full paragraph) that “Gaussian processes have also been studied in the bandit setting using confidence bound acquisition functions”, where Li’s Hyperband framework itself expands upon the sampling aspect mentioned just above per Spearmint with a batched approach (section 3.3, beginning on page 9, but see also the bulleted Data Set Subsampling discussion on page 10).
Like Koch, Li is directed to hyperparameter optimization, and is therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Li’s sampling contributions to hyperparameter search and optimization aspects into a framework such as Koch’s, with a reasonable expectation of success, such as to improve speed and efficiency as discussed per Li’s Abstract and page 2’s first full paragraph.
Regarding claim 14, the claim includes the same or similar limitations as claim 3 discussed above and is therefore rejected under the same rationale.
11. Claims 8 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Koch in view of van Rijn and further in view of Narayanan and further yet in view of U.S. Patent Application Publication No. 2019/0073570 (“Turco”).
Regarding claim 8, Koch in view of van Rijn and further in view of Narayanan teach the method of claim 1, as discussed above. The aforementioned references, e.g. per Koch’s [0066], teach the use and management of the input dataset, which might be stored using “a structured query language database.” Hence, Koch etc. very clearly teach a SQL database as linked to the input dataset, where the input dataset is part of the user’s request formulation per FIG. 5 step 506. That said, it is not entirely clear whether the request itself (Koch’s FIG. 5 step 528) includes or comprises a query to the database, e.g. per the further limitation wherein the hyperparameter optimization request comprises a SQL query. Rather, the Examiner relies upon TURCO to teach what Koch etc. otherwise lack, see e.g. Turco’s [0101] (“... Performing the predictions in the described manner in a database contexts may provide a huge performance gain compared to other systems, because the creation of tables for thousands of models which may never actually be used by any client is avoided and because in some embodiments the received input data is stored in a structured manner in temporary input tables such that fast, specially adapted analytical SQL routines 174 can be applied on the input data without having to export the data to a higher-level application program.”) and [0111] (“... the model manager module forwards the model M14 to the predictor module 160 which performs a prediction on the input data 124, thereby using the model M14 and optionally a stored SQL procedure ... ”).
Like Koch, Turco is directed to efficient and optimal practices relating to machine learning models, and is therefore analogous. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Turco’s SQL procedure approach, e.g. to facilitate access and management of input data, into a framework such as Koch’s, with a reasonable expectation of success, such as to simplify and improve memory/data management for the input data as Turco contemplates.
Regarding claim 19, the claim includes the same or similar limitations as claim 8 discussed above and is therefore rejected under the same rationale.
Conclusion
12. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHOURJO DASGUPTA whose telephone number is (571)272-7207. The examiner can normally be reached M-F 8am-5pm CST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at 571 272 4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHOURJO DASGUPTA/Primary Examiner, Art Unit 2144