Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner’s Note
Examiner notes that the claims were examined under the guidance of MPEP 2106 as possible judicial exceptions to eligibility under 35 U.S.C. 101. Examiner finds that the independent claims recite a mental process at Step 2A prong 1, namely, “selecting a path through the graph, wherein at least one input node is drawn from the plurality of input nodes depending on the probabilities assigned to the input nodes, wherein the path from the drawn input node along the edges to the output node is selected depending on the probabilities assigned to the edges,” which could be performed by a person using observation and evaluation, and with the aid of pencil and paper.
However, Examiner finds that the claims are not directed to the mental process, as the further elements of “creating a machine learning system depending on the selected path and training the created machine learning system” and “the probabilities of the edges and of the drawn input node of the path are adjusted” integrate the mental process into a practical application at step 2A prong 2.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 13–17 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 13:
Claim 13 recites in relevant part “(i) a first path is drawn through the graph from the input node along the edges to the additional node and a second path is drawn through the graph from the input node along the edges to the output node or (ii) the path is drawn through the graph from the input node along the edges via the additional node to the output node.” Examiner finds the phrase “the output node” indefinite in this limitation. Claim 13 previously recites “which additional node can serve as a further output node of the machine learning system,” and claim 11, upon which claim 13 depends, recites “providing a directed graph, wherein the directed graph includes a plurality of input nodes and at least one output node and a plurality of further nodes.” Examiners finds it unclear whether “the output node” of the claim is intended to be identified with the “additional node,” with one of the “at least one output node,” or with neither. The claim is therefore indefinite as a person having ordinary skill in the art would be unable to determine the metes and bound of the claim.
Regarding claim 14:
Claim 14 is rejected by dependence on claim 13. Further, claim 14 recites in relevant part “wherein the subset of the nodes is divided into sets of additional nodes, wherein each additional node is assigned a probability, which characterizes a probability with which the node from the set into which it is divided is drawn.” Examiner finds the phrase “which characterizes a probability with which the node from the set into which it is divided is drawn” indefinite, in particular, “the node from the set into which it is divided.” Examiner cannot identify a particular antecedent for the pronoun “it.” The claim is therefore indefinite as a person having ordinary skill in the art would be unable to determine the metes and bound of the claim.
Regarding claim 15:
Claim 15 is rejected by dependence on claims 13 and 14.
Regarding claim 16:
Claim 16 recites “the method according to claim 11, wherein the directed graph includes a first search space, wherein a resolution of data assigned to the nodes is continuously reduced, wherein the graph includes a second search space, which includes the additional nodes, wherein sets of additional nodes are respectively attached to a node of the first search space.”
In this limitation, the phrase “the additional nodes” occurs without an antecedent for “additional nodes.” Examiner finds this phrase indefinite as a person having ordinary skill in the art would be uncertain whether these additional nodes refers to some set of previously indicated nodes, to one or more of the “sets of additional nodes,” or neither.
Further, Examiner finds the limitation indefinite as to whether each “wherein” after the first is intended to indicate a narrowing of the previous phrase or is intended to be more broadly applied to “the method of claim 11” as a whole.
Further, Examiner finds the phrase “a resolution of data assigned to the nodes is continuously reduced” indefinite. The meaning of the “resolution of data assigned to the nodes” is unclear, for example, what data is being referred to, and what “resolution” is measuring. Further, the meaning of “continuously reduced” is unclear in the content of the claim, for example, whether the reduction occurs over time, over depths in the graph, or over some other variable.
A person having ordinary skill in the art would be unable to determine the metes and bounds of the claim.
Regarding claim 17:
Claim 17 recites in relevant part “the input nodes provide the following data: camera images and/or lidar data and/or radar data and/or ultrasonic data and/or thermal image data and/or microscopy data, including data from different perspectives.” Examiner finds the wording ambiguous as to whether “data from different perspectives” is included among the “microscopy data” or among the more general “following data” provided by the input nodes. Further, Examiner finds the wording ambiguous as to whether “data from different perspectives” is a required or optional component of the more general “data” provided by the input nodes. A person having ordinary skill in the art would be unable to determine the metes and bounds of claim. In further examination below, Examiner interprets “including data from different perspectives” as meaning that any of the following data provided by the inputs nodes may include data from different perspectives.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 11–13 and 17–19 rejected under 35 U.S.C. 103 over Houlsby et al., US Pre-Grant Publication No. 2022/0092416 (hereafter Houlsby) in view of Wang et al., US Pre-Grant Publication No. 2022/0300821 (hereafter Wang).
Regarding claim 11 and analogous claims 18–19:
Houlsby teaches:
“A computer-implemented method for creating a machine learning system for sensor data fusion, comprising the following steps”: Houlsby, paragraph 0005, “This specification describes a system implemented as computer programs [A computer-implemented method] on one or more computers in one or more locations that determines an architecture for a task neural network that is configured to perform a machine learning task”; Houlsby, paragraph 0089, “Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus.”
(bold only) “providing a directed graph, wherein the directed graph includes a plurality of input nodes and at least one output node and a plurality of further nodes, wherein the input and output nodes are connected via the further nodes using directed edges, wherein each respective edge of the edges is respectively assigned a probability, which characterizes a probability with which the respective edge is drawn, wherein each respective input node of the input nodes is also respectively assigned a probability”: Houlsby, paragraph 0034, “The search space of possible architectures for the task neural network is represented as a graph of nodes connected by edges [providing a directed graph]. Each node in the graph represents a decision point in selecting the architecture and each edge in the graph represents an action [wherein the directed graph includes … at least one output node and a plurality of further nodes]. In particular, the actions represented by outgoing edges from a given node are the possible decisions that can be made at the decision point represented by the given node [hence, a probability of an edge infers a probability onto the referenced node]. Each path through the graph that starts at an initial node in the graph and ends at a terminal node in the graph determines or represents an architecture for the task neural network”; Houlsby, paragraph 0085, “The system samples an action from the score distribution (step 506). In other words, the system selects the action, where the likelihood that each action is selected is defined by the score assigned to that action in the score distribution. When the scores are probabilities, the system samples the action by sampling from the probability distribution [wherein each respective edge of the edges is respectively assigned a probability, which characterizes a probability with which the respective edge is drawn, wherein each respective … node … is also respectively assigned a probability].”
(bold only) “selecting a path through the graph, wherein at least one input node is drawn from the plurality of input nodes depending on the probabilities assigned to the input nodes, wherein the path from the drawn input node along the edges to the output node is selected depending on the probabilities assigned to the edges”: Houlsby, paragraph 0034, “The search space of possible architectures for the task neural network is represented as a graph of nodes connected by edges. Each node in the graph represents a decision point in selecting the architecture and each edge in the graph represents an action. In particular, the actions represented by outgoing edges from a given node are the possible decisions that can be made at the decision point represented by the given node [hence, a probability of an edge infers a probability onto the referenced node]. Each path through the graph that starts at an initial node in the graph and ends at a terminal node in the graph determines or represents an architecture for the task neural network”; Houlsby, paragraph 0037, “As a general example, one node in the graph can represent a decision point that determines whether to add another layer to the architecture and the outgoing edges from that node can include a first edge that corresponds to adding a layer and a second edge that corresponds not adding a layer. A node connected by the first outgoing edge from the first node can represent a decision that determines what type of layer the new layer is, and another node connected by the second edge to the first node can be a terminal node that indicates that the architecture is finalized because no more layers are to be added”; Houlsby, paragraph 0038, “Thus, when generating a path through the graph, each edge in the path is a value for a different hyperparameter and the path defines a candidate architecture for the task neural network by specifying the values (edges) for the hyperparameters at each decision point”; Houlsby, paragraph 0085, “The system samples an action from the score distribution (step 506). In other words, the system selects the action, where the likelihood that each action is selected is defined by the score assigned to that action in the score distribution. When the scores are probabilities, the system samples the action by sampling from the probability distribution [selecting a path through the graph, wherein at least one … node is drawn from the plurality of input nodes depending on the probabilities assigned to the input nodes, wherein the path from the drawn … node along the edges to the output node is selected depending on the probabilities assigned to the edges].”
“creating a machine learning system depending on the selected path and training the created machine learning system”: Houlsby, paragraph 0074, “The system then trains the instance to perform the particular machine learning task by training the instance on some or all of the received training data using a machine learning training technique that is appropriate for the task [creating a machine learning system depending on the selected path and training the created machine learning system], e.g., stochastic gradient descent with backpropagation or backpropagation-through-time. When the path includes nodes that represent hyperparameters that impact training, the system performs the training in accordance with values for those hyperparameters in the path.”
(bold only) “wherein adjusted parameters of the trained machine learning system are stored in corresponding edges of the directed graph and the probabilities of the edges and of the drawn input node of the path are adjusted”: Houlsby, paragraph 0035–0036, “In particular, each decision point determines (at least in part) the value of some hyperparameter of the candidate architecture of the task neural network; edges from the decision point can represent available hyperparameter values for the decision point [wherein adjusted parameters of the trained machine learning system are stored in corresponding edges of the directed graph]. Generally, a hyperparameter is a value that is set prior to the commencement of the training of the task neural network and that impacts the operations performed by the task neural network or in training of the task neural network. The hyperparameters can include any values that impact any of: the number of layers in the task neural network, the operations performed by a given layer in the neural network (the type of layer (e.g., convolutional or fully-connected or max pooling or average pooling), the number of filters for a convolutional layer, dimension of each filter, type of convolution, number of hidden units for a fully connected layer, dimensionality of hidden state for a recurrent neural network layer), the connectivity between any two layers in the neural network ( e.g., which layer or layers receive the output generated by a given layer, whether a skip or residual connection is included between two layers, and so on), and hyperparameters of the training process (e.g., the optimizer used in the training, the update rule parameters used by the selected optimizer, the weight between different terms in the objective function being trained, and so on)”; Houlsby, paragraph 0042, “Generally, the system 100 determines the final architecture for the task neural network by training the controller neural network 110 to iteratively adjust the values of the controller parameters. At each iteration, the controller neural network can generate one or more paths, each representing an architecture to be trained and evaluated, and the controller parameters can be updated or adjusted based on the results of each evaluation [the probabilities of the edges and of the drawn … node of the path are adjusted]. In this way, the controller neural network can be used to design an improved task neural network for a specific task, such as image classification or other forms of image processing.”
“repeating the selecting, creating, and training steps several times”: Houlsby, paragraph 0046–0047, “By repeatedly updating the values of the controller parameters in this manner, the system 100 can train the controller neural network 110 to generate new paths that result in task neural networks that have increased performance on the particular task, i.e., to maximize the expected accuracy on the validation set of the architectures proposed by the controller neural network 110. Once the controller neural network 110 has been trained, the system 100 can select the architecture that had the best (for example, highest) performance measure as the final architecture of the task neural network or can generate a new path through the search space in accordance with the trained values of the controller parameters (i.e. using the trained controller neural network) and use the architecture defined by the new path as the final architecture of the task neural network [repeating the selecting, creating, and training steps several times].”
“creating the machine learning system depending on the directed graph”: Houlsby, paragraph 0047, “Once the controller neural network 110 has been trained, the system 100 can select the architecture that had the best (for example, highest) performance measure as the final architecture of the task neural network [creating the machine learning system depending on the directed graph] or can generate a new path through the search space in accordance with the trained values of the controller parameters (i.e. using the trained controller neural network) and use the architecture defined by the new path as the final architecture of the task neural network.”
Houlsby does not explicitly teach:
(bold only) “providing a directed graph, wherein the directed graph includes a plurality of input nodes and at least one output node and a plurality of further nodes, wherein the input and output nodes are connected via the further nodes using directed edges, wherein each respective edge of the edges is respectively assigned a probability, which characterizes a probability with which the respective edge is drawn, wherein each respective input node of the input nodes is also respectively assigned a probability”
(bold only) “selecting a path through the graph, wherein at least one input node is drawn from the plurality of input nodes depending on the probabilities assigned to the input nodes, wherein the path from the drawn input node along the edges to the output node is selected depending on the probabilities assigned to the edges”
(bold only) “wherein adjusted parameters of the trained machine learning system are stored in corresponding edges of the directed graph and the probabilities of the edges and of the drawn input node of the path are adjusted”
Wang teaches (bold only) “providing a directed graph, wherein the directed graph includes a plurality of input nodes and at least one output node and a plurality of further nodes, wherein the input and output nodes are connected via the further nodes using directed edges, wherein each respective edge of the edges is respectively assigned a probability, which characterizes a probability with which the respective edge is drawn, wherein each respective input node of the input nodes is also respectively assigned a probability,” (bold only) “selecting a path through the graph, wherein at least one input node is drawn from the plurality of input nodes depending on the probabilities assigned to the input nodes, wherein the path from the drawn input node along the edges to the output node is selected depending on the probabilities assigned to the edges,” and (bold only) “wherein adjusted parameters of the trained machine learning system are stored in corresponding edges of the directed graph and the probabilities of the edges and of the drawn input node of the path are adjusted”: Wang, paragraph 0040, “FIG. 1B shows a hybrid optimization pipeline architecture 100B, consistent with an illustrative embodiment. The model search result can be accelerated and improved in accuracy through searching both pipelines using multiple optimizing strategies in parallel. Pipelines 150, 175 are configured to operate according to different strategies. The pipeline 150 is configured for a Bayesian Optimization (BO) model training and selection. The pipeline 175 is configured to operate as an evolutionary Neural Architecture Search (NAS) pipeline. Pipeline 150 receives model candidates 155, performs hyperparameter optimization 157 followed by feature engineering and feature transformation selection 159 [plurality of input nodes ]. Pipeline 175 is configured for receiving candidates 177 from a neural network and utilizing the neural network model candidates 177 for the evolution of the neural network architecture and hyperparameter optimization 179. It is determined which pipeline output to select, for example, by pipeline ranking. An ensemble of pipeline operations may be provided.”
Wang and Houlsby are analogous arts as they are both related to algorithmic model search. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the multiple choices of inputs in the context of end-to-end model selection of Wang with the teachings of Houlsby to arrive at the present invention, in order to include input or feature selection to improve model search, as stated in Wang, paragraph 0003, “According to one embodiment, a computer-implemented method of automatically generating a machine learning model includes identifying one or more visualization features of a dataset associated with a machine learning model selection process. A plurality of candidate machine learning pipelines are configured for performing respective optimizing strategies in parallel based on the identified visualization features. A machine learning model is automatically generated based on at least one of the generated candidate machine learning pipelines.”
Regarding claim 12:
Houlsby as modified by Wang teaches “[t]he method according to claim 11.”
Houlsby further teaches:
“wherein the directed graph includes a plurality of output nodes”: Houlsby, paragraph 0040, “In particular, for each node in the path, the controller neural network 110 processes a controller input for the node to generate a score distribution over the actions represented by outgoing edges from the node. The system 100 then samples an action from the score distribution. If the node is not a terminal node, the system 100 adds, to the path, the node that is connected to by the outgoing edge represented by the sampled action. The outgoing edges for terminal nodes [wherein the directed graph includes a plurality of output nodes] do not connect to any other nodes in the graph and the system therefore terminates the path after sampling the action for the terminal node.”
“wherein, when selecting the path, at least one output node is selected from the plurality of output nodes depending on the assigned probabilities of the output nodes”: Houlsby, paragraph 0034, “The search space of possible architectures for the task neural network is represented as a graph of nodes connected by edges. Each node in the graph represents a decision point in selecting the architecture and each edge in the graph represents an action. In particular, the actions represented by outgoing edges from a given node are the possible decisions that can be made at the decision point represented by the given node. Each path through the graph that starts at an initial node in the graph and ends at a terminal node in the graph determines or represents an architecture for the task neural network [wherein, when selecting the path, at least one output node is selected from the plurality of output nodes depending on the assigned probabilities of the output nodes].”
(bold only) “wherein the probabilities assigned to the input nodes depend on the drawn output nodes”: Houlsby, paragraph 0045, “The controller parameter updating engine 130 then uses the results of the evaluations for the paths in the batch 112 to update the current values of the controller parameters to improve the expected performance of the architectures defined by the paths generated by the controller neural network 110 on the task [wherein the probabilities assigned to the input nodes depend on the drawn output nodes]. Evaluating the performance of trained instances and updating the current values of the controller parameters is described in more detail below with reference to FIG. 4.”
Wang teaches (bold only) “wherein the probabilities assigned to the input nodes depend on the drawn output nodes”: Wang, paragraph 0040, “FIG. 1B shows a hybrid optimization pipeline architecture 100B, consistent with an illustrative embodiment. The model search result can be accelerated and improved in accuracy through searching both pipelines using multiple optimizing strategies in parallel. Pipelines 150, 175 are configured to operate according to different strategies. The pipeline 150 is configured for a Bayesian Optimization (BO) model training and selection. The pipeline 175 is configured to operate as an evolutionary Neural Architecture Search (NAS) pipeline. Pipeline 150 receives model candidates 155, performs hyperparameter optimization 157 followed by feature engineering and feature transformation selection 159 [input nodes]. Pipeline 175 is configured for receiving candidates 177 from a neural network and utilizing the neural network model candidates 177 for the evolution of the neural network architecture and hyperparameter optimization 179. It is determined which pipeline output to select, for example, by pipeline ranking. An ensemble of pipeline operations may be provided.”
Wang and Houlsby are combinable for the rationale given under claim 11.
Regarding claim 13:
Houlsby as modified by Wang teaches “[t]he method according to claim 11.”
Houlsby further teaches:
“wherein a subset is determined from the plurality of further nodes, all of which satisfy a specified property with regard to a data resolution, wherein at least one additional node is selected from the subset”: Houlsby, paragraph 0054, “The search space 220 represents the decision points in generating the architecture as a graph. As can be seen from the search space 220, the search space 220 is represented as a graph with two terminal decision points: selecting the learning rate if the optimizer selected is not Adam and selecting the value for B1 if the optimizer selected is Adam [wherein a subset is determined from the plurality of further nodes, all of which satisfy a specified property with regard to a data resolution, wherein at least one additional node is selected from the subset, optimizer selection interpreted as a specified property].”
“which additional node can serve as a further output node of the machine learning system”: Houlsby, paragraph 0037, “As a general example, one node in the graph can represent a decision point that determines whether to add another layer to the architecture and the outgoing edges from that node can include a first edge that corresponds to adding a layer and a second edge that corresponds not adding a layer. A node connected by the first outgoing edge from the first node can represent a decision that determines what type of layer the new layer is, and another node connected by the second edge to the first node can be a terminal node that indicates that the architecture is finalized because no more layers are to be added [which additional node can serve as a further output node of the machine learning system].”
“wherein, during the selection: (i) a first path is drawn through the graph from the input node along the edges to the additional node and a second path is drawn through the graph from the input node along the edges to the output node or (ii) the path is drawn through the graph from the input node along the edges via the additional node to the output node”: Houlsby, paragraph 0034–0035, “The search space of possible architectures for the task neural network is represented as a graph of nodes connected by edges. Each node in the graph represents a decision point in selecting the architecture and each edge in the graph represents an action. In particular, the actions represented by outgoing edges from a given node are the possible decisions that can be made at the decision point represented by the given node. Each path through the graph that starts at an initial node in the graph and ends at a terminal node in the graph determines or represents an architecture for the task neural network [the path is drawn through the graph from the input node along the edges via the additional node to the output node]. In particular, each decision point determines (at least in part) the value of some hyperparameter of the candidate architecture of the task neural network; edges from the decision point can represent available hyperparameter values for the decision point. Generally, a hyperparameter is a value that is set prior to the commencement of the training of the task neural network and that impacts the operations performed by the task neural network or in training of the task neural network.”
Regarding claim 17:
Houlsby as modified by Wang teaches “[t]he method according to claim 11.”
Houlsby further teaches “wherein outputs of the machine learning system output a segmentation and/or object detection and/or depth estimation and/or gesture/behavior recognition, and the input nodes provide the following data: camera images and/or lidar data and/or radar data and/or ultrasonic data and/or thermal image data and/or microscopy data, including data from different perspectives”: Houlsby, paragraph 0018, “In some cases, the task neural network is a convolutional neural network that is configured to receive an input image [camera images] and to process the input image to generate a network output for the input image, i.e., to perform some kind of machine learning image processing task. For example, the particular machine learning task may be image classification and the output generated by the neural network for a given image may be scores for each of a set of object categories, with each score representing an estimated likelihood that the image contains an image of an object belonging to the category [wherein outputs of the machine learning system output … object detection].”
Claim 14 rejected under 35 U.S.C. 103 over Houlsby as modified by Wang in view of Benyahia et al., US Pre-Grant Publication No. 2020/0104688 (hereafter Benyahia).
Houlsby as modified by Wang teaches “[t]he method according to claim 13.”
Houlsby further teaches:
“wherein each additional node is assigned a probability, which characterizes a probability with which the node from the set into which it is divided is drawn”: Houlsby, paragraph 0085, “The system samples an action from the score distribution (step 506). In other words, the system selects the action, where the likelihood that each action is selected is defined by the score assigned to that action in the score distribution. When the scores are probabilities, the system samples the action by sampling from the probability distribution [wherein each additional node is assigned a probability, which characterizes a probability with which the node from the set into which it is divided is drawn].”
“wherein, when selecting the path, an additional node is respectively drawn randomly from each of the sets”: Houlsby, paragraph 0085, “The system samples an action from the score distribution (step 506). In other words, the system selects the action [an additional node is respectively drawn randomly from each of the sets], where the likelihood that each action is selected is defined by the score assigned to that action in the score distribution. When the scores are probabilities, the system samples the action by sampling from the probability distribution.”
“and wherein the probability assigned to the additional nodes is also adjusted during the training”: Houlsby, paragraph 0042, “Generally, the system 100 determines the final architecture for the task neural network by training the controller neural network 110 to iteratively adjust the values of the controller parameters. At each iteration, the controller neural network can generate one or more paths, each representing an architecture to be trained and evaluated, and the controller parameters can be updated or adjusted based on the results of each evaluation [the probability assigned to the additional nodes is also adjusted during the training]. In this way, the controller neural network can be used to design an improved task neural network for a specific task, such as image classification or other forms of image processing.”
Houlsby as modified by Wang does not explicitly teach “wherein the subset of the nodes is divided into sets of additional nodes.”
Benyahia teaches “wherein the subset of the nodes is divided into sets of additional nodes”: Benyahia, paragraph 0039, “The controller may be trained to suggest candidate models, e.g., from a subset of options of nodes for the model and possible subgraphs [wherein the subset of the nodes is divided into sets of additional nodes]. The controller may be updated (or trained) based on presenting the controller with an indication of a preferred model out of a subset of candidate models. Performance of the preferred model may be tested using a validation set. The controller may then be trained to suggest better candidate models by learning which models worked well previously. To train the controller, its policy may be updated by updating values for the controller parameters. The controller may provide a set of suggested candidate models, which may then be sequentially trained before being analyzed to help identify the preferred model. This process may be iterative.”
Benyahia and Houlsby are analogous arts as they are both related to algorithmic model search. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have combined the node subsets of Benyahia with the teachings of Houlsby to arrive at the present invention, in order to improve model search, as stated in Benyahia, paragraph 0021, “As the training of a network may be very time-consuming, initially selecting suitable neurons may provide a significant reduction in the time, and resources, taken to provide a neural network that performs satisfactorily for a selected task. In addition to increasing the efficiency of training a neural network, the ability to select a suitable configuration of neurons may make the difference between the provision of a neural network that may solve a technical problem, and one that cannot.”
Claim 15 rejected under 35 U.S.C. 103 over Houlsby as modified by Wang and Benyahia in view of Ning Wang et al., “NAS-FCOS: Fast Neural Architecture Search for Object Detection,” 2020, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (hereafter Ning).
Houlsby as modified by Wang and Benyahia teaches “[t]he method according to claim 13.”
Houlsby further teaches:
(bold only) “wherein each task-specific head is assigned a probability, which characterizes a probability with which the task-specific head is drawn”: Houlsby, paragraph 0085, “The system samples an action from the score distribution (step 506). In other words, the system selects the action, where the likelihood that each action is selected is defined by the score assigned to that action in the score distribution. When the scores are probabilities, the system samples the action by sampling from the probability distribution [wherein each … is assigned a probability, which characterizes a probability with which … is drawn].”
(bold only) “wherein, when selecting the path, one of the task-specific heads is drawn from a plurality of task-specific heads depending on the probabilities assigned to the task-specific heads”: Houlsby, paragraph 0038, “Thus, when generating a path through the graph, each edge in the path is a value for a different hyperparameter and the path defines a candidate architecture for the task neural network by specifying the values (edges) for the hyperparameters at each decision point”; Houlsby, paragraph 0085, “The system samples an action from the score distribution (step 506). In other words, the system selects the action, where the likelihood that each action is selected is defined by the score assigned to that action in the score distribution. When the scores are probabilities, the system samples the action by sampling from the probability distribution [wherein, when selecting the path, one … is drawn from a plurality of task-specific heads depending on the probabilities assigned]
(bold only) “and wherein the probability assigned to the task-specific heads is also adjusted during the training”: Houlsby, paragraph 0042, “Generally, the system 100 determines the final architecture for the task neural network by training the controller neural network 110 to iteratively adjust the values of the controller parameters. At each iteration, the controller neural network can generate one or more paths, each representing an architecture to be trained and evaluated, and the controller parameters can be updated or adjusted based on the results of each evaluation [the probability assigned … is also adjusted during the training]. In this way, the controller neural network can be used to design an improved task neural network for a specific task, such as image classification or other forms of image processing.”
Houlsby as modified by Wang and Benyahia does not explicitly teach:
(bold only) “wherein each task-specific head is assigned a probability, which characterizes a probability with which the task-specific head is drawn”
(bold only) “wherein, when selecting the path, one of the task-specific heads is drawn from a plurality of task-specific heads depending on the probabilities assigned to the task-specific heads”
(bold only) “and wherein the probability assigned to the task-specific heads is also adjusted during the training”
Ning teaches (bold only) “wherein each task-specific head is assigned a probability, which characterizes a probability with which the task-specific head is drawn,” (bold only) “wherein, when selecting the path, one of the task-specific heads is drawn from a plurality of task-specific heads depending on the probabilities assigned to the task-specific heads,” and (bold only) “and wherein the probability assigned to the task-specific heads is also adjusted during the training”: Ning, section 3.2.2, “Prediction head h maps each feature in the pyramid P to the output of corresponding y, which in FCOS and RetinaNet, consists of four 3×3 convolutions [task-specific head ]. To explore the potential of the head, we therefore extend a sequential search space for its generation. Specifically, our head is defined as a sequence of six basic operations. Compared with candidate operations in the FPN structures, the head search space has two slight differences. First, we add standard convolution modules (including conv1x1 and conv3x3) to the head sampling pool for better comparison. Second, we follow the design of FCOS by replacing all the Batch Normalization (BN) layers to Group Normalization (GN) [25] in the operations sampling pool of head, considering that head needs to share weights between different levels, which causes BN invalid. The final output of head is the output of the last (sixth) layer.”
Ning and Houlsby are analogous arts as they are both related to algorithmic model search. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have applied the extension of the model search space to model heads from Ning with the teachings of Houlsby to arrive at the present invention, in order to improve model selection results, as stated in Ning, section 1, paragraph 5, “In this work, we propose a fast and memory-efficient NAS method for searching both FPN and head architectures, with carefully designed proxy tasks, search space and evaluation strategies, which is able to find top-performing architectures over 3,000 architectures using 28 GPU-days only.”
Claim 16 rejected under 35 U.S.C. 103 over Houlsby as modified by Wang in view of Ning.
Houlsby as modified by Wang teaches “[t]he method according to claim 11.”
Houlsby further teaches:
“wherein the directed graph includes a first search space”: Houlsby, paragraph 0009, “The described techniques on the other hand, make use of a graph search space [wherein the directed graph includes a first search space]. Thus the sequence of decisions defining an architecture is not fixed, but is determined dynamically by the actions selected at each decision.”
“wherein a resolution of data assigned to the nodes is continuously reduced”: Houlsby, paragraph 0075, “In some implementations, the system parallelizes the training of the instances to decrease the overall training time for the controller neural network. The system can train each instance for a specified amount of time or a specified number of training iterations or until convergence [wherein a resolution of data assigned to the nodes is continuously reduced, interpreted as including reducing the error of the generated model].”
Houlsby as modified by Wang does not explicitly teach “wherein the graph includes a second search space, which includes the additional nodes, wherein sets of additional nodes are respectively attached to a node of the first search space.”
Ning teaches “wherein the graph includes a second search space, which includes the additional nodes, wherein sets of additional nodes are respectively attached to a node of the first search space”: Ning, section 3.2.2, “Prediction head h maps each feature in the pyramid P to the output of corresponding y [wherein sets of additional nodes are respectively attached to a node of the first search space], which in FCOS and RetinaNet, consists of four 3×3 convolutions. To explore the potential of the head, we therefore extend a sequential search space for its generation [a second search space]. Specifically, our head is defined as a sequence of six basic operations. Compared with candidate operations in the FPN structures, the head search space has two slight differences. First, we add standard convolution modules (including conv1x1 and conv3x3) to the head sampling pool for better comparison. Second, we follow the design of FCOS by replacing all the Batch Normalization (BN) layers to Group Normalization (GN) [25] in the operations sampling pool of head, considering that head needs to share weights between different levels, which causes BN invalid. The final output of head is the output of the last (sixth) layer.”
Ning and Houlsby are analogous arts as they are both related to algorithmic model search. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to have applied the extension of the model search space to model heads from Ning with the teachings of Houlsby to arrive at the present invention, in order to improve model selection results, as stated in Ning, section 1, paragraph 5, “In this work, we propose a fast and memory-efficient NAS method for searching both FPN and head architectures, with carefully designed proxy tasks, search space and evaluation strategies, which is able to find top-performing architectures over 3,000 architectures using 28 GPU-days only.”
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Appel et al., US Pre-Grant Publication No. 2022/0188663, discloses a method of neural architecture search using a body-and-head model paradigm.
Vasudevan et al., US Pre-Grant Publication No. 2019/0026639, discloses a method of neural architecture search implemented as a graph search, in which parameters of the graph search are updated during the search.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to VINCENT SPRAUL whose telephone number is (703) 756-1511. The examiner can normally be reached M-F 9:00 am - 5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MICHAEL HUNTLEY can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/VAS/Examiner, Art Unit 2129
/MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129