Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
The objection to the Drawings is withdrawn in view of the amendment filed on 07/29/2026.
The objections to the specification are withdrawn in view of the amendment filed on 07/29/2026.
Applicant’s arguments in view of the claim amendments filed on 07/29/2026, with respect to the 35 U.S.C. §112 rejections have been fully considered and are persuasive. The 35 U.S.C. §112 rejections have been withdrawn.
Applicant’s arguments with respect to prior art rejections have been considered but are moot because the new ground of rejection does not rely on any of the same combination of references applied in the prior rejections of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 2-7, 10-15 and 18-21 are rejected under 35 U.S.C. 103 as being unpatentable over US Pat. Pub. No. 2015/0379424 to Dirac et al. (previously cited, hereinafter Dirac) in view of US Pat. Pub. No. 2018/0114114 to Molchanov et al. (hereinafter Molchanov).
Per claim 2, Dirac discloses A computer-implemented method comprising: under control of one or more processors, (Dirac: ¶[0032], fig. 1:185…Dirac teaches that the machine learning service is carried out by a control plane and a data plane made up of pools of servers and their associated storage, which constitutes performance of the recited method under control of one or more processors, "The data plane of the MLS may include, for example, at least a subset of the servers of pool(s) 185, storage devices that are used to store input data sets, intermediate results or final results (some of which may be part of the MLS artifact repository), and the network pathways used for transferring client input data and results").
receiving, from an electronic communication device, a request for a machine learning model, wherein the request identifies: (i) an input type to be received by the machine learning model, (ii) an output type to be provided by the machine learning model, and (iii) training data for the machine learning model (Dirac: ¶[0066], fig. 8:812, 861…Dirac teaches that a client device submits a model request over a programmatic interface and that the request itself names the input data the model is to consume and the type of output the model is to produce, which constitutes a request identifying an input type and an output type, "In the depicted embodiment, a client 164 of the MLS may submit a model execution request 812 to the MLS control plane 180 via a programmatic interface 861. The model execution request may specify the execution mode (batch, online or local), the input data to be used for the model run (which may be produced using a specified data source or recipe in some cases), the type of output (e.g., a prediction or an evaluation) that is desired, and/or optional parameters"; ¶[0029]…Dirac further teaches that the same programmatic interfaces let a client name the data source supplying the training data used to build the model, which constitutes the third recited element of the request, "The APIs implemented by the MLS may in some embodiments allow clients to submit requests to create, query the attributes of, read, update/modify, search, or delete an instance of at least some of the various entity types supported. For example, for the entity type ‘DataSource’, respective APIs similar to ‘createDataSource’, ‘describeDataSource’ (to obtain the values of attributes of the data source), ‘updateDataSource’, ‘searchForDataSource’, and ‘deleteDataSource’ may be supported by the MLS. A similar set of APIs may be supported for recipes, models, and so on").
identifying, from a library of machine learning models, a trained machine learning model, wherein identifying the trained machine learning model is based at least partly on at least one of the input type or the output type identified in the request (Dirac: ¶[0030]…Dirac teaches an artifact repository holding already-trained models that are published for reuse and reached through an alias pointer, which constitutes the recited library of machine learning models and the trained machine learning model identified from it, "an alias may comprise an immutable name … and a pointer to a model that has already been created and stored in an MLS artifact repository … an internal identifier generated for the model by the MLS"; ¶[0091], fig. 15:1501, 1507, 1509, 1511, 1513, 1515…Dirac teaches that the stored artifacts are indexed by problem domain, and each of Dirac’s enumerated domains is defined by the kind of data it consumes and the kind of result it produces, so selecting among them constitutes a selection made on the basis of input type or output type, "In the depicted example, a MLS customer can use a check-box to select from among the problem domains fraud detection 1507, sentiment analysis 1509, image analysis 1511, genome analysis 1513, or voice recognition 1515").
generating the machine learning model using the trained machine learning model and the training data, wherein generating the machine learning model comprises (Dirac: ¶[0030]…Dirac teaches that an already-trained model in the repository is carried forward and reworked with further input data to yield an improved successor model, which constitutes generating a machine learning model using the trained machine learning model and the training data, "The model developers may continue to experiment with various algorithms, parameters and/or input data sets to obtain improved versions of the underlying model, and may be able to change the pointer to point to an enhanced version to improve the quality of predictions obtained by the business analysts"; ¶[0030]…Dirac further teaches that the successor model put in the place of a stored model must deliver the input and prediction types the requester expects of it, which constitutes the requirement that the generated model provide an output of the output type identified in the request or receive the input type identified in the request, "when an alias pointer is changed, both the original model and the new model (i.e., the respective models being pointed to by the old pointer and the new pointer) consume the same type of input and provide the same type of prediction (e.g., binary classification, multi-class classification or regression)"): …
Dirac does not expressly disclose, but Molchanov does teach:
modifying a layer of the trained machine learning model to provide an output of the output type identified in the request or receive the input type identified in the request to form the machine learning model (Molchanov: ¶[0021], fig. 1A:140…Molchanov teaches that the operation is performed upon a network that has already been trained and alters the composition of that network’s layers to yield a different network, which constitutes modifying a layer of the trained machine learning model to form the machine learning model, "At step 140, the at least one neuron is removed from the trained neural network to produce a pruned neural network"; ¶[0030], fig. 1D…Molchanov further teaches that the alteration is made to the neurons of a given layer, which constitutes the modification being made to a layer of that model rather than to the network at large, "In coarse pruning, entire neurons (or feature maps) are removed. As shown in FIG. 1D, the patterned neuron is removed during coarse pruning. When a neuron is removed, all connections to and from the neuron are removed").
wherein modifying the layer of the trained machine learning model comprises removing nodes from the layer of the trained machine learning model (Molchanov: ¶[0030], fig. 1D…Molchanov teaches that the removal operates on whole neurons of a layer together with every connection they carry, which constitutes removing nodes from the layer of the trained model rather than merely zeroing individual weights, "In coarse pruning, entire neurons (or feature maps) are removed. As shown in FIG. 1D, the patterned neuron is removed during coarse pruning. When a neuron is removed, all connections to and from the neuron are removed").
training the machine learning model using the training data (Molchanov: ¶[0045], fig. 2C:210…Molchanov teaches that the network from which neurons have been removed is thereafter optimized against a supplied dataset, which constitutes training the model formed by the modification using the training data, "At step 210, the pruned neural network is fine-tuned using conventional techniques. Fine-tuning involves optimizing parameters of the network to minimize a cost function on a given dataset").
Dirac and Molchanov are analogous art because both references are from the same field of endeavor, specifically the construction, adaptation and reuse of trained machine learning models. Dirac addresses that field at the service level, maintaining a repository of already-trained models that are republished and improved for new client requests (Dirac: ¶[0030]). Molchanov addresses it at the layer level, removing neurons from a trained network and refitting the result (Molchanov: ¶[0021], ¶[0045]). Each reference is additionally reasonably pertinent to the particular problem with which the inventor was involved, namely producing a model that satisfies a requester’s input and output requirements by reworking an existing trained model instead of building one from scratch.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to carry out the layer modification by which Dirac’s republished trained model is fitted to the requester’s output type through Molchanov’s removal of neurons from that layer. This is the application of a known technique to a known device ready for improvement to yield predictable results (MPEP 2143(I)(D)).
The suggestion/motivation for doing so is provided by Molchanov itself, which identifies the very deficiency that arises when a service such as Dirac’s fits a large pre-trained model to a requester’s narrower task and teaches neuron removal as the remedy, "In these cases, accuracy may be improved by fine-tuning an existing deep network previously trained on a much larger labeled vision dataset. While transfer learning of this form supports state of the art accuracy, inference is expensive due to the time, power, and memory demanded by the heavyweight architecture of the fine-tuned network. Thus, there is a need for addressing these issues and/or other issues associated with the prior art" (Molchanov: ¶[0003]). Dirac supplies the matching incentive, teaching that a published model is expected to be executed repeatedly on behalf of a wider audience of users than its creators (Dirac: ¶[0030]), so a PHOSITA would have had concrete reason to reduce the per-inference cost of each republished model by the pruning Molchanov describes.
Per claim 3, Dirac combined with Molchanov discloses claim 2. Dirac further teaches identifying a shape of an input of the input type for the machine learning model, wherein the shape indicates at least one of a number of input values or a data type for an input value to the machine learning model; and determining that the shape corresponds to a trained model input shape for the trained machine learning model (Dirac: ¶[0081], fig. 11:1110…Dirac teaches that before an artifact is run the service checks the submitted input against the format that artifact expects, which constitutes determining that the input’s shape corresponds to the trained model’s input shape, "perform a set of run-time validations (e.g., to ensure that the requester is permitted to execute the recipe, that the input data appears to be in the correct or expected format, and so on)"; ¶[0081]…Dirac further teaches that the transformation libraries are chosen from the data types of the input records, which constitutes the recited shape indicating a data type for an input value, "the specific libraries or functions to be used for the transformation may be selected based on the data types of the input records").
Per claim 4, Dirac combined with Molchanov discloses claim 2.
Dirac further teaches identifying a portion of the machine learning model based at least in part on annotation information for the machine learning model, wherein the portion includes a layer of the machine learning model (Dirac: ¶[0030]…Dirac teaches that each stored model artifact carries service-generated identifying metadata by which the artifact is addressed, "a pointer to a model that has already been created and stored in an MLS artifact repository (e.g., ‘samModel-23adf-2013-12-13-08-06-01’, an internal identifier generated for the model by the MLS)").
Molchanov further teaches determining an adjustment to a weight for a path from a node included in the layer of the machine learning model, wherein the machine learning model including the weight with the adjustment provides the output with an output value of a higher accuracy than the trained machine learning model including the weight without the adjustment (Molchanov: ¶[0045], fig. 2C:210…Molchanov teaches that after nodes are removed the remaining connection parameters are re-optimized against a cost function, which constitutes determining an adjustment to a weight for a path from a node of the affected layer, "At step 210, the pruned neural network is fine-tuned using conventional techniques. Fine-tuning involves optimizing parameters of the network to minimize a cost function on a given dataset"; ¶[0003]…Molchanov further teaches expressly that the adjustment so made raises the accuracy of the network relative to the network as previously trained, which constitutes the recited higher accuracy, "In these cases, accuracy may be improved by fine-tuning an existing deep network previously trained on a much larger labeled vision dataset").
The rationale to combine Molchanov with Dirac is the same as stated for the parent claim.
Per claim 5, Dirac combined with Molchanov discloses claim 2. Dirac further teaches receiving, from the electronic communication device, an image for processing by the machine learning model; retrieving the machine learning model; processing the image using the machine learning model to generate an image processing result, the image processing result including at least one of segmentation information or classification information for an object shown in the image; and transmitting the image processing result to the electronic communication device (Dirac: ¶[0035]…Dirac teaches an image processing data type among the input data types the service accepts, which constitutes receiving an image for processing by the model, "The input data may comprise data records that include variables of any of a variety of data types, such as, for example text, a numeric data type (e.g., real or integer), Boolean, a binary data type, a categorical data type, an image processing data type, an audio processing data type, a bioinformatics data type"; ¶[0091], fig. 15:1511…Dirac teaches image analysis as a supported problem domain of the service, "In the depicted example, a MLS customer can use a check-box to select from among the problem domains fraud detection 1507, sentiment analysis 1509, image analysis 1511, genome analysis 1513, or voice recognition 1515"; ¶[0030]…Dirac teaches that the result a served model provides is a classification, which, taken with the image input type and the image analysis domain, constitutes classification information for an object shown in the image, "consume the same type of input and provide the same type of prediction (e.g., binary classification, multi-class classification or regression)"; ¶[0066]…Dirac further teaches that the model is executed on the client-supplied input and the requested result returned to the requesting client, "The model execution request may specify the execution mode (batch, online or local), the input data to be used for the model run (which may be produced using a specified data source or recipe in some cases), the type of output (e.g., a prediction or an evaluation) that is desired").
Per claim 6, Dirac discloses A system comprising: one or more computing devices having a processor and a memory, wherein the one or more computing devices execute computer-readable instructions to at least (Dirac: ¶[0038], fig. 2:202, 255, 258…Dirac teaches that the machine learning service is deployed upon the compute, storage and database resources of a provider network, which constitutes the recited system of one or more computing devices having a processor and a memory, "In the embodiment shown in FIG. 2, the MLS utilizes storage service 202, computing service 258, and database service 255 of provider network 202").
receive, from an electronic communication device, a request for a machine learning model, wherein the request identifies: (i) an input type to be received by the machine learning model, (ii) an output type to be provided by the machine learning model, and (iii) training data for the machine learning model (Dirac: ¶[0066], fig. 8:812, 861…Dirac teaches that a client device submits a model request over a programmatic interface and that the request itself names the input data the model is to consume and the type of output the model is to produce, which constitutes a request identifying an input type and an output type, "In the depicted embodiment, a client 164 of the MLS may submit a model execution request 812 to the MLS control plane 180 via a programmatic interface 861. The model execution request may specify the execution mode (batch, online or local), the input data to be used for the model run (which may be produced using a specified data source or recipe in some cases), the type of output (e.g., a prediction or an evaluation) that is desired, and/or optional parameters"; ¶[0029]…Dirac further teaches that the same programmatic interfaces let a client name the data source supplying the training data used to build the model, which constitutes the third recited element of the request, "The APIs implemented by the MLS may in some embodiments allow clients to submit requests to create, query the attributes of, read, update/modify, search, or delete an instance of at least some of the various entity types supported. For example, for the entity type ‘DataSource’, respective APIs similar to ‘createDataSource’, ‘describeDataSource’ (to obtain the values of attributes of the data source), ‘updateDataSource’, ‘searchForDataSource’, and ‘deleteDataSource’ may be supported by the MLS. A similar set of APIs may be supported for recipes, models, and so on").
identify a first trained machine learning model based at least partly on the request (Dirac: ¶[0030]…Dirac teaches an artifact repository holding already-trained models that are published for reuse and reached through an alias pointer, which constitutes the recited library of machine learning models and the trained machine learning model identified from it, "an alias may comprise an immutable name … and a pointer to a model that has already been created and stored in an MLS artifact repository … an internal identifier generated for the model by the MLS"; ¶[0091], fig. 15:1501, 1507, 1509, 1511, 1513, 1515…Dirac teaches that the stored artifacts are indexed by problem domain, and each of Dirac’s enumerated domains is defined by the kind of data it consumes and the kind of result it produces, so selecting among them constitutes a selection made on the basis of input type or output type, "In the depicted example, a MLS customer can use a check-box to select from among the problem domains fraud detection 1507, sentiment analysis 1509, image analysis 1511, genome analysis 1513, or voice recognition 1515").
generate the machine learning model using the first trained machine learning model (Dirac: ¶[0030]…Dirac teaches that an already-trained model in the repository is carried forward and reworked with further input data to yield an improved successor model, which constitutes generating a machine learning model using the trained machine learning model and the training data, "The model developers may continue to experiment with various algorithms, parameters and/or input data sets to obtain improved versions of the underlying model, and may be able to change the pointer to point to an enhanced version to improve the quality of predictions obtained by the business analysts"; ¶[0030]…Dirac further teaches that the successor model put in the place of a stored model must deliver the input and prediction types the requester expects of it, which constitutes the requirement that the generated model provide an output of the output type identified in the request or receive the input type identified in the request, "when an alias pointer is changed, both the original model and the new model (i.e., the respective models being pointed to by the old pointer and the new pointer) consume the same type of input and provide the same type of prediction (e.g., binary classification, multi-class classification or regression)").
Dirac does not expressly disclose, but Molchanov does teach:
modify a layer of the first trained machine learning model to provide an output of the output type identified in the request or receive the input type identified in the request to form the machine learning model (Molchanov: ¶[0021], fig. 1A:140…Molchanov teaches that the operation is performed upon a network that has already been trained and alters the composition of that network’s layers to yield a different network, which constitutes modifying a layer of the first trained machine learning model to form the machine learning model, "At step 140, the at least one neuron is removed from the trained neural network to produce a pruned neural network"; ¶[0030], fig. 1D…Molchanov further teaches that the alteration is made to the neurons of a given layer, which constitutes the modification being made to a layer of that model rather than to the network at large, "In coarse pruning, entire neurons (or feature maps) are removed. As shown in FIG. 1D, the patterned neuron is removed during coarse pruning. When a neuron is removed, all connections to and from the neuron are removed").
wherein modifying the layer of the first trained machine learning model comprises removing nodes from the layer of the first trained machine learning model (Molchanov: ¶[0030], fig. 1D…Molchanov teaches that the removal operates on whole neurons of a layer together with every connection they carry, which constitutes removing nodes from the layer of the trained model rather than merely zeroing individual weights, "In coarse pruning, entire neurons (or feature maps) are removed. As shown in FIG. 1D, the patterned neuron is removed during coarse pruning. When a neuron is removed, all connections to and from the neuron are removed").
train at least a portion of the machine learning model using the training data (Molchanov: ¶[0045], fig. 2C:210…Molchanov teaches that the network from which neurons have been removed is thereafter optimized against a supplied dataset, which constitutes training the model formed by the modification using the training data, "At step 210, the pruned neural network is fine-tuned using conventional techniques. Fine-tuning involves optimizing parameters of the network to minimize a cost function on a given dataset").
Dirac and Molchanov are analogous art for the reasons stated above as to claim 2, both being from the same field of endeavor of constructing, adapting and reusing trained machine learning models and each being reasonably pertinent to the problem of fitting an existing trained model to a requester’s input and output requirements.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to build Dirac’s system so that the layer modification adapting a republished trained model to the requester’s output type is performed by removing nodes from that layer as Molchanov teaches, this being the application of a known technique to a known device ready for improvement to yield predictable results (MPEP 2143(I)(D)).
The suggestion/motivation for doing so is provided by Molchanov itself, which teaches that a model produced by fine-tuning a large pre-trained network is costly to run and that neuron removal is the answer, "While transfer learning of this form supports state of the art accuracy, inference is expensive due to the time, power, and memory demanded by the heavyweight architecture of the fine-tuned network" (Molchanov: ¶[0003]).
Per claim 7, Dirac combined with Molchanov discloses claim 6. Dirac further teaches identify a shape of the output of the output type for the machine learning model, wherein the shape indicates at least one of a number of output values or a data type for an output value to the machine learning model; and determine that the shape corresponds to a trained model output shape for the first trained machine learning model (Dirac: ¶[0030]…Dirac teaches that the service enforces a correspondence between the output type a stored model produces and the output type expected of the artifact standing in its place, which constitutes determining that the output shape corresponds to the trained model’s output shape, "when an alias pointer is changed, both the original model and the new model (i.e., the respective models being pointed to by the old pointer and the new pointer) consume the same type of input and provide the same type of prediction (e.g., binary classification, multi-class classification or regression)").
Per claim 10, this claim is substantially similar in scope and spirit to claim 4. Therefore the rejection of claim 4 is applied accordingly.
Per claim 11, Dirac combined with Molchanov discloses claim 6. Dirac further teaches the first trained machine learning model is associated with model metadata describing the first trained machine learning model, and wherein the one or more computing devices execute computer-readable instructions to at least identify the first trained machine learning model based at least in part on a comparison of the model metadata and metadata included in the request (Dirac: ¶[0029]…Dirac teaches describe and search APIs that return and match the stored attributes of an artifact, and teaches that the same API set is provided for models, so a model is located by matching request-supplied attribute values against the model’s stored attributes, which constitutes identifying the model by a comparison of model metadata with metadata in the request, "respective APIs similar to ‘createDataSource’, ‘describeDataSource’ (to obtain the values of attributes of the data source), ‘updateDataSource’, ‘searchForDataSource’, and ‘deleteDataSource’ may be supported by the MLS. A similar set of APIs may be supported for recipes, models, and so on").
Per claim 12, Dirac combined with Molchanov discloses claim 6. Dirac further teaches receive, from another computing device, audio data for processing by the machine learning model; process the audio data using the machine learning model to generate a language processing result, the language processing result including at least one of: (i) a transcription of an utterance encoded by the audio data, or (ii) an intent for the utterance encoded by the audio data; and transmit the language processing result to the other computing device (Dirac: ¶[0091], fig. 15:1515…Dirac teaches voice recognition as a supported problem domain of the service, "In the depicted example, a MLS customer can use a check-box to select from among the problem domains fraud detection 1507, sentiment analysis 1509, image analysis 1511, genome analysis 1513, or voice recognition 1515"; ¶[0035]…Dirac further teaches an audio processing data type among the input data types the service accepts, which, taken with the voice recognition domain, constitutes generating a language processing result from audio data, "an audio processing data type").
Per claim 13, Dirac combined with Molchanov discloses claim 6. Dirac further teaches determine that an accuracy of an output provided by the machine learning model corresponds to a target accuracy; and activate a network address to receive an input for processing via the machine learning model (Dirac: ¶[0066]…Dirac teaches that a request may carry a quality target against which the produced model is assessed and that an evaluation of the model is among the results the service returns, which constitutes determining that the model’s accuracy corresponds to a target accuracy, "optional parameters (such as desired model quality targets, minimum input record group sizes to be used for online predictions, and so on)"; ¶[0065]…Dirac further teaches that a network address is assigned as the destination to which input data records for the model are submitted, which constitutes activating a network address to receive an input for processing via the model, "In real-time mode, a network endpoint (e.g., an IP address) may be assigned as a destination to which input data records for a specified model are to be submitted, and model predictions may be generated on groups of streaming data records as the records are received"; ¶[0066]…Dirac further teaches that the model is mounted at that address, "For online mode 867, the model may be mounted (e.g., configured with a network address) to which data records may be streamed, and from which results including predictions 868 and/or evaluations 869 can be retrieved").
Per claim 14, Dirac discloses A computer-implemented method comprising: under control of one or more processors, (Dirac: ¶[0032], fig. 1:185…Dirac teaches that the machine learning service is carried out by a control plane and a data plane made up of pools of servers and their associated storage, which constitutes performance of the recited method under control of one or more processors, "The data plane of the MLS may include, for example, at least a subset of the servers of pool(s) 185, storage devices that are used to store input data sets, intermediate results or final results (some of which may be part of the MLS artifact repository), and the network pathways used for transferring client input data and results").
receiving, from an electronic communication device, a request for a machine learning model, wherein the request identifies: (i) an input type, the input type indicating a data format to be received by the machine learning model, (ii) an output type, the output type indicating a type of processing result to be provided by the machine learning model based on the input type, and (iii) training data for the machine learning model (Dirac: ¶[0066]…Dirac teaches that the request names both the data the model is to consume and the kind of result it is to produce, which constitutes an input type indicating a data format and an output type indicating a type of processing result, "The model execution request may specify the execution mode (batch, online or local), the input data to be used for the model run (which may be produced using a specified data source or recipe in some cases), the type of output (e.g., a prediction or an evaluation) that is desired"; ¶[0035]…Dirac teaches that the input data records carry variables of enumerated data formats, which constitutes the input type indicating a data format, "The input data may comprise data records that include variables of any of a variety of data types, such as, for example text, a numeric data type (e.g., real or integer), Boolean, a binary data type, a categorical data type, an image processing data type, an audio processing data type, a bioinformatics data type"; ¶[0030]…Dirac further teaches that the kind of prediction a served model provides is fixed together with the kind of input it consumes, which constitutes the output type indicating a type of processing result based on the input type, "consume the same type of input and provide the same type of prediction (e.g., binary classification, multi-class classification or regression)").
identifying, from a library of machine learning models, a first trained machine learning model, wherein identifying the first trained machine learning model is based at least partly on at least one of the input type or the output type identified in the request (Dirac: ¶[0030]…Dirac teaches an artifact repository holding already-trained models that are published for reuse and reached through an alias pointer, which constitutes the recited library of machine learning models and the trained machine learning model identified from it, "an alias may comprise an immutable name … and a pointer to a model that has already been created and stored in an MLS artifact repository … an internal identifier generated for the model by the MLS"; ¶[0091], fig. 15:1501, 1507, 1509, 1511, 1513, 1515…Dirac teaches that the stored artifacts are indexed by problem domain, and each of Dirac’s enumerated domains is defined by the kind of data it consumes and the kind of result it produces, so selecting among them constitutes a selection made on the basis of input type or output type, "In the depicted example, a MLS customer can use a check-box to select from among the problem domains fraud detection 1507, sentiment analysis 1509, image analysis 1511, genome analysis 1513, or voice recognition 1515").
generating the machine learning model using the first trained machine learning model and the training data, wherein generating the machine learning model comprises (Dirac: ¶[0030]…Dirac teaches that an already-trained model in the repository is carried forward and reworked with further input data to yield an improved successor model, which constitutes generating a machine learning model using the trained machine learning model and the training data, "The model developers may continue to experiment with various algorithms, parameters and/or input data sets to obtain improved versions of the underlying model, and may be able to change the pointer to point to an enhanced version to improve the quality of predictions obtained by the business analysts"; ¶[0030]…Dirac further teaches that the successor model put in the place of a stored model must deliver the input and prediction types the requester expects of it, which constitutes the requirement that the generated model provide an output of the output type identified in the request or receive the input type identified in the request, "when an alias pointer is changed, both the original model and the new model (i.e., the respective models being pointed to by the old pointer and the new pointer) consume the same type of input and provide the same type of prediction (e.g., binary classification, multi-class classification or regression)").
Dirac does not expressly disclose, but Molchanov does teach:
modifying a layer of the first trained machine learning model to form the machine learning model that provides a processing result of the output type identified in the request or receives the data format indicated by the input type identified in the request (Molchanov: ¶[0021], fig. 1A:140…Molchanov teaches that the operation is performed upon a network that has already been trained and alters the composition of that network’s layers to yield a different network, which constitutes modifying a layer of the first trained machine learning model to form the machine learning model, "At step 140, the at least one neuron is removed from the trained neural network to produce a pruned neural network"; ¶[0030], fig. 1D…Molchanov further teaches that the alteration is made to the neurons of a given layer, which constitutes the modification being made to a layer of that model rather than to the network at large, "In coarse pruning, entire neurons (or feature maps) are removed. As shown in FIG. 1D, the patterned neuron is removed during coarse pruning. When a neuron is removed, all connections to and from the neuron are removed").
wherein modifying the layer of the first trained machine learning model comprises removing nodes from the layer of the first trained machine learning model (Molchanov: ¶[0030], fig. 1D…Molchanov teaches that the removal operates on whole neurons of a layer together with every connection they carry, which constitutes removing nodes from the layer of the trained model rather than merely zeroing individual weights, "In coarse pruning, entire neurons (or feature maps) are removed. As shown in FIG. 1D, the patterned neuron is removed during coarse pruning. When a neuron is removed, all connections to and from the neuron are removed").
training at least a portion of the machine learning model using the training data (Molchanov: ¶[0045], fig. 2C:210…Molchanov teaches that the network from which neurons have been removed is thereafter optimized against a supplied dataset, which constitutes training the model formed by the modification using the training data, "At step 210, the pruned neural network is fine-tuned using conventional techniques. Fine-tuning involves optimizing parameters of the network to minimize a cost function on a given dataset").
Dirac and Molchanov are analogous art for the reasons stated above as to claim 2, both being from the same field of endeavor of constructing, adapting and reusing trained machine learning models and each being reasonably pertinent to the problem of fitting an existing trained model to a requester’s stated input format and output result type.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to carry out the layer modification of Dirac’s republished trained model, by which the model is fitted to the requester’s stated output type, through the removal of nodes from that layer as Molchanov teaches, this being the application of a known technique to a known device ready for improvement to yield predictable results (MPEP 2143(I)(D)).
The suggestion/motivation for doing so is provided by Molchanov itself, which teaches that the fine-tuned product of a large pre-trained network is expensive to run and identifies neuron removal as the remedy, "While transfer learning of this form supports state of the art accuracy, inference is expensive due to the time, power, and memory demanded by the heavyweight architecture of the fine-tuned network" (Molchanov: ¶[0003]).
Per claims 15, 18 and 19, these claims are substantially similar in scope and spirit to claims 7, 4 and 11, respectively. Therefore the rejections of claims 7, 4 and 11 are applied accordingly.
Per claim 20, Dirac combined with Molchanov discloses claim 14. Dirac further teaches receiving, from a client device, tabular data for processing by the machine learning model, the tabular data describing a user; processing the tabular data using the machine learning model to generate a prediction for the user; and transmitting the prediction to the client device (Dirac: ¶[0003]…Dirac teaches records of many variables per entity in a fraud detection setting, which constitutes tabular data describing a user, "transaction records, each representing dozens or even hundreds of variables"; ¶[0091], fig. 15:1507…Dirac teaches fraud detection as a supported problem domain, "In the depicted example, a MLS customer can use a check-box to select from among the problem domains fraud detection 1507, sentiment analysis 1509, image analysis 1511, genome analysis 1513, or voice recognition 1515"; ¶[0066]…Dirac further teaches that a prediction is the requested output returned to the requesting client, "the type of output (e.g., a prediction or an evaluation) that is desired").
Per claim 21, Dirac combined with Molchanov discloses claim 14. Dirac further teaches identifying a topical domain for a user associated with the request, wherein identifying the first trained machine learning model comprises determining that the topical domain relates to a domain associated with the first trained machine learning model (Dirac: ¶[0091], fig. 15:1501…Dirac teaches that the requester designates a problem domain and that the service returns the stored artifacts associated with that domain, which constitutes identifying a topical domain and determining that it relates to the domain associated with the identified trained model, "In the depicted example, a MLS customer can use a check-box to select from among the problem domains fraud detection 1507, sentiment analysis 1509, image analysis 1511, genome analysis 1513, or voice recognition 1515. A user may also search for recipes associated with other problem domains using search term text block 1517 in the depicted web page").
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Dirac in view of Molchanov, as applied in the rejection of claim 6 above, and further in view of Net2Net: Accelerating Learning via Knowledge Transfer to Chen et al. (hereinafter Chen).
Dirac combined with Molchanov discloses claim 6.
Dirac combined with Molchanov does not expressly disclose, but Chen does teach: add a node to the layer of the first trained machine learning model (Chen: pg. 3, §2.3…Chen teaches an operation that replaces a layer of an existing network with a layer holding a greater number of units, which constitutes adding a node to that layer, "Our first proposed transformation is Net2WiderNet transformation. This allows a layer to be replaced with a wider layer, meaning a layer that has more units"; fig. 2…Chen teaches that the network so enlarged is the already-trained teacher network and that the added unit is created by replicating an existing unit of that network, which constitutes the addition being made to the layer of the first trained machine learning model, "The student network is larger because we replicate the h[2] unit of the teacher. … To replicate the h[2] unit, we copy its weights c and d to the new h[3] unit"; p.5, §2.3…Chen further teaches that the enlargement may be applied to a single designated layer, which constitutes the addition being made to the layer, "This operator can be applied arbitrarily many times; we can expand only one layer of the network, or we can expand all non-output layers").
Dirac, Molchanov and Chen are analogous art because all three references are from the same field of endeavor, specifically the construction, adaptation and reuse of trained machine learning models, and each is reasonably pertinent to the particular problem of reworking an existing trained model to meet a new requirement rather than building one from scratch. Chen addresses that field by transferring the parameters of a trained network into a structurally altered successor network (Chen: p. 1), which is the same reuse of a trained artifact that Dirac's service performs over its repository (Dirac: ¶[0030]) and that Molchanov performs upon a trained network's layers (Molchanov: ¶[0021]).
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to have the layer modification of the combination of Dirac and Molchanov additionally add a node to that layer as Chen teaches, this being the application of a known technique to a known device ready for improvement to yield predictable results (MPEP 2143(I)(D)).
The suggestion/motivation for doing so is provided by Chen itself, which identifies the cost of rebuilding a model from scratch each time a new requirement arises and teaches the enlargement of an existing trained network as the remedy, "During real-world workflows, one often trains very many different neural networks during the experimentation and design process. This is a wasteful process in which each new model is trained from scratch. Our Net2Net technique accelerates the experimentation process by instantaneously transferring the knowledge from a previous network to each new deeper or wider network" (Chen: p.1, Abstract). Dirac supplies the matching incentive, teaching that stored models are carried forward and improved rather than rebuilt (Dirac: ¶[0030]), so a PHOSITA fitting a republished model to a requester whose output type demands greater capacity than the stored model provides would have had concrete reason to enlarge the layer by Chen's function-preserving addition of units.
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Dirac in view of Molchanov, as applied in the rejection of claim 14 above, and further in view of US Pat. Pub. No. 2016/0132787 to Drevo et al. (previously cited, hereinafter Drevo).
Dirac combined with Molchanov discloses claim 14.
Dirac combined with Molchanov does not expressly disclose, but Drevo does teach: wherein modifying the layer of the first trained machine learning model comprises updating a hyperparameter for the layer of the first trained machine learning model (Drevo: ¶[0003]…Drevo teaches that a layered model is configured by setting per-layer quantities including the count of hidden units in each layer together with continuous training parameters, and identifies these as the parameter and hyperparameter choices to be made for the model, which constitutes updating a hyperparameter for the layer, "Consider for example, a DBN model. In most cases, a data scientist needs to choose a number of layers and a transfer function for each layer. Then, the data scientist further needs to choose a number of hidden units for each layer and values for continuous parameters, such as learning rate, number of epochs, pre-training learning rate, and learning rate decay"; ¶[0087], fig. 3:300…Drevo teaches a structure whose nodes represent the parameter and hyperparameter choices spanning the model’s option space, "A CPT 300 expresses a modeling methodology’s option space, which includes combined discrete, categorical, and/or continuous parameters as well as any hyperparameters").
Dirac, Molchanov and Drevo are analogous art because all three references are from the same field of endeavor, specifically the construction, adaptation and reuse of trained machine learning models, and each is reasonably pertinent to the particular problem of configuring a model so that it satisfies a requester’s requirements without being rebuilt from scratch. Drevo addresses that field as a distributed platform that automates the selection and tuning of modeling methodologies (Drevo: ¶[0007]), which is the same automation Dirac’s service performs over its repository of stored models (Dirac: ¶[0030]).
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to have the layer modification of the combination of Dirac and Molchanov additionally update a per-layer hyperparameter as Drevo teaches, this being the application of a known technique to a known device ready for improvement to yield predictable results (MPEP 2143(I)(D)).
The suggestion/motivation for doing so is provided by Drevo itself, which teaches that the per-layer choices govern how well the resulting model performs and that leaving them unautomated is the deficiency its platform exists to cure, "A data scientist does not know apriori which methodology will result in the best performing model. To make the challenge more difficult, tuning a methodology can have a large impact on performance because a given methodology may have numerous parameters and design choices" (Drevo: ¶[0002]).
Allowable Subject Matter
Claims 9 and 17 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is the statement of reasons for the indication of allowable subject
matter: The prior art disclosed by the applicant and cited by the Examiner fail to teach
or suggest, alone or in combination, all the limitations of the independent and
intervening claims (claims 6 and 14), further including the particular notable limitations
of: identify the first trained machine learning model and a second trained model
based at least in part on the request; generate a first accuracy metric for the first
trained machine learning model based at least partly on processing of a portion
of the training data with the first trained machine learning model; generate a
second accuracy metric for the second trained model based at least partly on
processing of the portion of the training data with the second trained model; and
determine that the first accuracy metric indicates a higher level of accuracy than
the second accuracy metric.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALAN CHEN whose telephone number is (571)272-4143. The examiner can normally be reached M-F 10-7.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALAN CHEN/Primary Examiner, Art Unit 2125