Prosecution Insights
Last updated: October 01, 2026
Application No. 16/998,545

SYSTEMS AND METHODS FOR CLASSIFYING DATA USING HIERARCHICAL CLASSIFICATION MODEL

Non-Final OA §103
Filed
Aug 20, 2020
Examiner
FEITL, LEAH M
Art Unit
2147
Tech Center
2100 — Computer Architecture & Software
Assignee
Capital One Services LLC
OA Round
7 (Non-Final)
23%
Grant Probability
At Risk
7-8
OA Rounds
0m
Est. Remaining
28%
With Interview

Examiner Intelligence

Grants only 23% of cases
23%
Career Allowance Rate
21 granted / 93 resolved
-32.4% vs TC avg
Moderate +6% lift
Without
With
+5.8%
Interview Lift
resolved cases with interview
Typical timeline
4y 3m
Avg Prosecution
24 currently pending
Career history
127
Total Applications
across all art units

Statute-Specific Performance

§101
29.9%
-10.1% vs TC avg
§103
47.6%
+7.6% vs TC avg
§102
7.0%
-33.0% vs TC avg
§112
13.9%
-26.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 93 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 03/04/2026 has been entered. Status of Claims This action is in response to the amendments filed 03/04/2026. Claims 1-2, 4-5, 13, 17, and 20 have been amended, claims 1-6 and 8-21 are currently pending. Response to Arguments Applicant’s arguments regarding the prior art rejection have been fully considered but are moot because of the new ground(s) of rejection. Applicant argues that the Wang reference does not teach a hierarchical child model wherein a child model is “configured to classify received data directly. . .without first classifying the received data by the top-level machine learning model” and a top-level model is located at least one level above the one or more child machine learning models. Examiner notes that given further consideration, the Wang reference is no longer relied upon and the Cheng reference has been brought in to more clearly teach a hierarchical model that can skip levels in the hierarchy. Examiner notes that at least section 3.2 of Cheng teaches an example where the “recreation” level, which is the parent model to the “sports” and “dancing” models, can be skipped but would still be considered at “top-level” model over the “sports” and “dancing” child models. The prior art rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-6, and 8-13, 15-18, and 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over Yan et al (US 20160117587 A1, herein Yan) in view of Ding et al (US 20210073686 A1, herein Ding), in further view of Cheng et al, “Hierarchical Classification of Documents with Error Control”, herein Cheng), in further view of Gonnet (US 20170220665 A1, herein Gonnet). Regarding claim 1, Yan teaches a system for classifying data (para. [0016] recites “Example methods and systems are directed to hierarchical deep CNNs for image classification”), the system comprising: a memory unit storing instructions; and one or more processors configured to execute the instructions to perform operations comprising: receiving a training dataset (para. [0021] recites “FIG. 1 is a network diagram illustrating a network environment 100 suitable for creating and using hierarchical deep CNNs for image classification, according to some example embodiments. The network environment 100 includes e-commerce servers 120 and 140, an HD-CNN server 130, and devices 150A, 150B, and 150C, all communicatively coupled to each other via a network 170. The devices 150A, 150B, and 150C may be collectively referred to as "devices 150," or generically referred to as a "device 150." Para. [0025] recites “the HD-CNN server 130 receives data regarding an item of interest to a user. For example, a camera attached to the device 150A can take an image of an item the user 160 wishes to sell and transmit the image over the network 170 to the HD-CNN server 130” (i.e., receiving a dataset for training a classification model)); training a top-level machine learning model based on the training dataset and a set of possible categories (para. [0023] recites “The HD-CNN server 130 creates an HD-CNN for classifying images, uses an HD-CNN to classify images, or both. For example, the HD-CNN server 130 can create an HD-CNN for classifying images based on a training set or a preexisting HD-CNN can be loaded onto the HD-CNN server 130”. Para. [0017] recites “A hierarchical deep CNN (HD-CNN) follows the coarse-to-fine classification strategy and modular design principle. For any given class label, it is possible to define a set of easy classes and a set of confusing classes. Accordingly, an initial coarse classifier CNN can separate the easily separable classes from one another. Subsequently, the challenging classes are routed to downstream fine CNNs that focus solely on confusing classes. In some example embodiments, HDCNN improves classification performance over standard deep CNN models. As with a CNN, the structure of an HD-CNN (e.g., the structure of each component CNN, the number of fine classes, and so on) may be determined by a designer, while the parameters of each layer of each CNN may be determined through training.” Para. [0039] recites “FIG. 5 is a block diagram illustrating relationships between components of the classification module 250, according to some example embodiments. A single standard deep CNN can be used as the building block of the fine prediction components of an HD-CNN. As shown in FIG. 5, a coarse category CNN 520 predicts the probabilities over coarse categories. Multiple branching CNNs 540-550 are independently added”. Para. [0049] recites “In operation 610, the coarse category identification module 220 divides a set of training samples into a training set and an evaluation set. For example, a dataset consisting of N, training samples {xi, y’i}, for i in the range of! to N, is divided into two parts, train_train and train_ val” (i.e., training a top-level image classification model in a hierarchical model)). However, while Yan teaches pretraining and training steps of a machine learning model (see at least paragraph [0031]), Yan does not explicitly teach retraining a machine learning model. Ding teaches retraining a machine learning model (Ding para. [0022] recites “FIG. 1B depicts a self-structured hierarchical machine learning classifier”. Ding para. [0025] recites “because the Main Classifiers 160A and 160B are trained separately on reduced data sets, they can each be refined, retrained, and/or replaced more rapidly, as compared to training a single classifier to evaluate all of the classes (i.e., models in a hierarchy can be re-trained)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by applying the performance evaluation and retraining methods from Ding to models in the hierarchical classification system from Yan. Yan and Ding are both directed to using neural networks for hierarchical classification of image data, but Yan does not explicitly teach retraining models. One of ordinary skill would understand how to apply the performance evaluation method from Ding to determine when to retrain the models for top-level and conjoined categories from Yan. Yan in view of Ding teaches retraining the top-level machine learning model based on an updated set of the possible categories to obtain an updated top-level machine learning model, the updated possible set of categories comprising: one or more conjoined categories that combine possible categories of the set of possible categories, that are being misclassified by the top-level machine learning model (Ding para. [0025] recites “because the Main Classifiers 160A and 160B are trained separately on reduced data sets, they can each be refined, retrained, and/or replaced more rapidly, as compared to training a single classifier to evaluate all of the classes” (i.e., retraining at least a top-level model in a hierarchy on data sets that can be). Ding para. [0018] recites “with respect to the classes "A" and "B," the Credibility Assessments 110A and 110B have determined that the Main Classifier 105 does not meet the predefined performance criteria. Thus, the system has generated and trained a Secondary Classifier 115 for these classes”. Ding para. [0019] recites “the system identifies a subset of the original training data to be used to train the Secondary Classifier 115. In one such embodiment, the system identifies training data that corresponds to the poor-performing classes, and uses this subset of exemplars to train the Secondary Classifier 115. In the illustrated embodiment, this would involve identifying the training exemplars labeled as class "A" and class "B," and training the Secondary Classifier 115 using only these exemplars (e.g., excluding exemplars labeled "N" or other classes)” (i.e., determining that a performance criterion is not met and adjusting, or transforming, the classes, such as the coarse, or conjoined categories from paragraph [0030] of Yan and the example from paragraph [0038] of Yan where items like “bus” belong to a different coarse, or top-level model from a coarse, or top-level model including categories that may be conjoined from items like “apple” or “orange” corresponding to fine, or child models)). However, the combination of Yan and Ding does not teach classifying received data directly by one or more child machine learning models without first classifying the received data by a top-level machine learning model. Cheng teaches that classifying received data directly by one or more child machine learning models without first classifying the received data by a top-level machine learning model (section 3.2 para. 2 recites “We run three classification methods in parallel. The first and third classifications are hierarchical classifications of traditional sense. The second classification is performed by dynamically skipping some levels in the class taxonomy. For example, to classify a document based on the taxonomy in Fig. 1, we can perform an additional classification by first skipping level 1 (i.e., {Recreation, Business}). Say the three classifications end up with class `Dancing'. We then classify it at node `Recreation'. But this time we skip the left part of level 2, i.e., {Sports, Dancing}” (i.e., a child model can be used directly without first classifying with a parent node)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by applying the method of skipping levels in a hierarchy from Cheng to the hierarchical classification model from Yan (as modified by Ding). Cheng and Yan are both directed to methods for building hierarchical classification models. One of ordinary skill in the art would benefit from using the method from Cheng to skip levels in the hierarchical convolutional neural network model, as Cheng teaches that “Skipping some levels has the effect of (partially) globalizing the information for the classification, and therefore can possibly reducing the misclassification rate. The more levels skipped, the more likely it is to reduce misclassification rate.” in section 3.2. Yan in view of Ding and Cheng teaches constructing, based on generating the one or more child machine learning models for the one or more conjoined categories, a hierarchical classification machine learning model that includes the updated top-level machine learning model and the one or more child machine learning models and that is configured to classify received data directly by the one or more child machine learning models without first classifying the received data by the top-level machine learning model, wherein the updated top-level machine learning model is located at least one level above the one or more child machine learning models in the hierarchical classification machine learning model (Yan para. [0017] recites “A hierarchical deep CNN (HD-CNN) follows the coarse-to-fine classification strategy and modular design principle. For any given class label, it is possible to define a set of easy classes and a set of confusing classes. Accordingly, an initial coarse classifier CNN can separate the easily separable classes from one another. Subsequently, the challenging classes are routed to downstream fine CNNs that focus solely on confusing classes” (i.e., constructing a hierarchical classification model that includes a top-level model, such as the retrained top model from at least paragraph [0025] of Ding, and at least one generated child model for at least one conjoined category below the top-level model). Cheng section 3.2 para. 2 recites “The second classification is performed by dynamically skipping some levels in the class taxonomy. For example, to classify a document based on the taxonomy in Fig. 1, we can perform an additional classification by first skipping level 1” (i.e., a child model, such as the fine model from Yan, can receive data directly without first classifying data with its parent model, such as the coarse model from Yan)). However, while the Yan reference teaches using a distance matrix to determine the set of possible categories (see at least paragraph [0051]), the combination of Yan, Ding, and Cheng does not explicitly teach using a subset of one or more data samples of the possible categories having a distance less than a threshold distance to a clustering center of one or more clusters, to determine a category of the one or more categories. Gonnet teaches using a subset of one or more data samples of the possible categories having a distance less than a threshold distance to a clustering center of one or more clusters, to determine a category of the one or more categories (para. [0048] recites “a distance may be computed, using chi-square or another statistical measure. This distance may be utilized in performing classifications to form groupings 104 (e.g., clusters) based on identified similarities between aspects of information. The distances based on statistical measures can be input or provided in the form of a distance matrix. For a group of documents, the pairs of distances is arranged in a matrix of distances (e.g., an all-against-all matrix of distances)”. Para. [0074] recites “The distance matrix may be processed in various ways to arrange, in the form of an informational topology/ontology 106 (e.g., a tree, a linked list, a directed graph) having various weights and/or indicia that may indicate, among other aspects, similar elements of information 102, the strength of determined similarity between elements of information 102, etc. The informational topology/ontology 106 may be hierarchical, and the relationships between connected nodes can be associated with linkage values to indicate the significance of relationships between the nodes” (i.e., determining a category in a potential hierarchy of categories using data having a distance less than a threshold distance to a given cluster center)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by utilizing the distance matrix calculation methods from Gonnet to modify the distance matrix used in Yan (as modified by Ding and Cheng). Yan and Gonnet are both directed to methods which can be used to cluster data in a hierarchical format. One of ordinary skill in the art would understand how to modify Yan’s distance matrix to make use of the methods from Gonnet of calculating a distance matrix to cluster training data into categories that could be utilized by a hierarchical model. Yan in view of Ding, Cheng, and Gonnet teaches generating, using a subset of one or more data samples of the possible categories having a distance less than a threshold distance to a clustering center of one or more clusters, one or more child machine learning models for the one or more conjoined categories and trained to consider additional feature parameters of the data sample used to generate the one or more child machine learning models (Gonnet para. [0074] recites “The distance matrix may be processed in various ways to arrange, in the form of an informational topology/ontology 106 (e.g., a tree, a linked list, a directed graph) having various weights and/or indicia that may indicate, among other aspects, similar elements of information 102, the strength of determined similarity between elements of information 102, etc. The informational topology/ontology 106 may be hierarchical, and the relationships between connected nodes can be associated with linkage values to indicate the significance of relationships between the nodes” (i.e., constructing a hierarchical model from categories determined by computed distances between clusters). Yan para. [0060] – [0061] recite “In operation 730, a loop is begun to process each of the C' fine category components. Accordingly, the operations 740 and 750 are performed for each fine category component. For example, when four coarse categories are identified, the loop will iterate over each of four fine category components (i.e., a fine, or child model, is iterated over fine categories). The pretrain module 230 makes a copy of the prototypical fine category component for the fine category component, in operation 740. Thus, all fine category components are initialized into the same state. The fine category component is further trained on the portion of the dataset corresponding to the coarse category of the fine category component. For example, the subset of the dataset {xi, yi} where P(yi) is the coarse category may be used. Once all fine category components and the coarse category component have been trained, the HD-CNN is constructed” (i.e., generating at least a first child model for at least one output conjoined category in the hierarchical model, wherein the child model further trained on, or considers, additional features)); Regarding claim 2, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the top-level machine learning model is a classification machine learning model comprising a neural network (Yan para. [0016] recites “Example methods and systems are directed to hierarchical deep CNNs for image classification” (i.e., the hierarchical classification machine learning model includes a neural network as the top-level model)). Regarding claim 3, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the operations further comprise: storing the hierarchical classification machine learning model in the storage device (Yan para. [0032]-[0033] recite “The classification module 250 is configured to receive and process image data. The image data may be a two-dimensional image, a frame from a continuous video stream, a three-dimensional image, a depth image, an infrared image, a binocular image, or any suitable combination thereof. For example, an image may be received from a camera. To illustrate, a camera may take a picture and send it to the classification module 250. The classification module 250 determines a fine category for the image by using an HDCNN (e.g., by determining a coarse category or coarse category weights using a coarse category CNN and determining the fine category using one or more fine category CNNs). The HD-CNN may have been generated using the pretrain module 230, the fine-tune module 240, or both. Alternatively, the HD-CNN may have been provided from an external source. The storage module 260 is configured to store and retrieve data generated and used by the coarse category identification module 220, the pretrain module 230, the fine-tune module 240, and the classification module 250. For example, the HD-CNN generated by the pretrain module 230 can be stored by the storage module 260 for retrieval by the fine-tune module 240. Information regarding categorization of an image, generated by the classification module 250, can also be stored by the storage module 260” (i.e., storing a hierarchical machine learning model)). Regarding claim 4, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the operations further comprise: classifying, by the retrained top-level machine learning model, the training dataset (Yan para. [0039] recites “FIG. 5 is a block diagram illustrating relationships between components of the classification module 250, according to some example embodiments. A single standard deep CNN can be used as the building block of the fine prediction components of an HD-CNN. As shown in FIG. 5, a coarse category CNN 520 predicts the probabilities over coarse categories. Multiple branching CNNs 540-550 are independently added. In some example embodiments, branching CNNs 540-550 share the branching shallow layers 530. The coarse category CNN 520 and the multiple branching CNNs 540-550 each receive the input image and operate on it in parallel. Although each branching CNN 540-550 receives the input image and gives a probability distribution over the full set of fine categories, the result of each branching CNN 540-550 is only valid for a subset of categories. The multiple full predictions from branching CNNs 540-550 are linearly combined by the probabilistic averaging layer 560 to form the final fine category prediction, weighted by the corresponding coarse category probabilities”. Ding para. [0025] recites “because the Main Classifiers 160A and 160B are trained separately on reduced data sets, they can each be refined, retrained, and/or replaced more rapidly, as compared to training a single classifier to evaluate all of the classes” (i.e., classifying the dataset by at least the re-trained top-level machine learning model)). Regarding claim 5, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the operations further comprise: generating a confusion matrix based on a result of classifying the training dataset by the top-level machine learning model; and generating the one or more conjoined categories based on the confusion matrix (Yan para. [0050] recites “In operation 630, the coarse category identification module 220 plots a confusion matrix based on train_ val. The confusion matrix is of size CxC. The columns of the matrix correspond to the predicted fine categories and the rows of the matrix correspond to the actual fine categories in train_ val”. Yan para. [0051] recites “The coarse category identification module 220 generates a distance matrix, D, by subtracting each element of the confusion matrix from 1 and zeroing the diagonal elements of D. The distance matrix is made symmetric by taking the average of D with DT, the transposition of D. After these operations are performed, each element Dij measures the ease with which category i is distinguished from category j”. Yan para. [0052] recites “In operation 640, low-dimensional feature representations {fi} for i in the range of 1 to C, are obtained for the fine categories” (i.e., using a confusion matrix based on the classification results and using the matrix to determine conjoined categories)). Regarding claim 6, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the operations further comprise: generating a plurality of confidence levels for respective data elements in the training dataset; and generating the one or more conjoined categories based on the confidence levels (Ding para. [0023] recites “two or more classes (e.g., classes "A and "B") are identified as requiring 95% or better accuracy, while two or more other classes (e.g., classes "C" and "D") require 75% accuracy. In an embodiment, the required accuracy for a given class is defined by one or more users. As illustrated, the system has trained a Group Classifier 155 to classify the Input 150 into the appropriate group, such that the Input 150 can be routed to the appropriate classifier. To do so, in one embodiment, the system identifies training exemplars corresponding to each group and labels them appropriately (e.g., labeling exemplars corresponding to classes "A" and "B" with a first group label, and labeling exemplars corresponding to classes "C" and "D" with a second group label). The Group Classifier 155 can then be trained on this group-labelled data. Similarly, the Main Classifiers 160A-B can be trained using the exemplars belonging to the corresponding classes, as discussed above. That is, continuing the above example, the Main Classifier 160A is trained based on data corresponding to classes "A" and "B," while the Main Classifier 160B is trained based on data corresponding to classes "C" and "D" (i.e., generating a conjoined category based on calculated confidence levels)). Regarding claim 8, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the operations further comprise: excluding from the updated set of possible categories, the possible categories joined to generate the one or more conjoined categories (Yan para. [0056] recites “In operation 710, the pretrain module 230 trains a coarse category CNN on a set of coarse categories. For example, the set of coarse categories may have been identified using the process 600. The fine categories of the training dataset are replaced with the coarse categories using the mapping P(y)=y'. In an example embodiment, the dataset {xi , y’i} , for i in the range of 1 to N, is used to train a standard deep CNN model. The trained model becomes the coarse category component of the HD-CNN (e.g., the coarse category CNN 520)” (i.e., replacing a single category with a conjoined category)). Regarding claim 9, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the operations further comprise: generating a first child machine learning model of the one or more child machine learning models for the one or more conjoined categories (Yan para. [0060] – [0061] recite “In operation 730, a loop is begun to process each of the C' fine category components. Accordingly, the operations 740 and 750 are performed for each fine category component. For example, when four coarse categories are identified, the loop will iterate over each of four fine category components (i.e., a fine, or child machine learning model, is iterated over fine categories). The pretrain module 230 makes a copy of the prototypical fine category component for the fine category component, in operation 740. Thus, all fine category components are initialized into the same state. The fine category component is further trained on the portion of the dataset corresponding to the coarse category of the fine category component. For example, the subset of the dataset {xi, yi} where P(yi) is the coarse category may be used. Once all fine category components and the coarse category component have been trained, the HD-CNN is constructed” (i.e., generating at least a first child machine learning model for at least one output conjoined category in the hierarchical machine learning model)). Regarding claim 10, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein generating the one or more child machine learning models comprises: generating a first child machine learning model of the one or more child machine learning models for a first conjoined category of the one or more conjoined categories; classifying, by the first child machine learning model, data in the first conjoined category based on the updated set of possible categories (Yan para. [0060] – [0061] recite “In operation 730, a loop is begun to process each of the C' fine category components. Accordingly, the operations 740 and 750 are performed for each fine category component. For example, when four coarse categories are identified, the loop will iterate over each of four fine category components. The pretrain module 230 makes a copy of the prototypical fine category component for the fine category component, in operation 740. Thus, all fine category components are initialized into the same state. The fine category component is further trained on the portion of the dataset corresponding to the coarse category of the fine category component. For example, the subset of the dataset {xi, yi} where P(yi) is the coarse category may be used. Once all fine category components and the coarse category component have been trained, the HD-CNN is constructed”. Ding para. [0025] recites “because the Main Classifiers 160A and 160B are trained separately on reduced data sets, they can each be refined, retrained, and/or replaced more rapidly, as compared to training a single classifier to evaluate all of the classes” (i.e., generating at least a first child machine learning model to classify at least a first conjoined category)). determining that a performance criterion is not satisfied based on the result of the classifying by the first child machine learning model; in response to determining that the performance criterion is not satisfied based on the result of the classifying by the first child machine learning model, determining that a stopping condition is not met based on a result of the classifying by the first child machine learning model (Ding para. [0046] recites “the method 400 proceeds to block 450, where the ML Application 235 similarly evaluates the secondary classifier. In an embodiment, the ML Application 235 does so using the second testing set. This evaluation can include, for example, generating a confusion matrix, evaluating class-specific precision, and the like. At block 455, the ML Application 235 similarly determines whether the secondary classifier has unsatisfactory performance with respect to any of the classes. If so, in one embodiment, the method 400 returns to block 430. This process can be iterated until all classes are adequately predicted. In another embodiment, the ML Application 235 requests additional data and retrains the secondary classifier until performance is adequate” (i.e., determining whether a performance criterion is met and determining whether a stopping condition has been met)); in response to determining that the stopping condition is not met, generating a second conjoined category based on the result of the classifying by the first child machine learning model (Ding para. [0045] recites “At block 430, the ML Application 235 selects one of the identified class( es) that has poor results. At block 435, the ML Application 235 identifies the data from the training set that corresponds to the selected class. For example, if the class "flower" performed poorly, the ML Application 235 identifies each exemplar in the training set that is labeled "flower." This subset of the training set is to be used to train supplementary classifier(s)”. Para. [0050] recites “If a secondary classifier is available, the method 500 continues to block 525 where the ML Application 235 processes the received input using this secondary classifier (e.g., Secondary Classifier 115). The method 500 then returns to block 515 to perform a credibility assessment, as discussed above. In some embodiments, this process repeats until a credible classification is returned, or no additional classifiers remain” (i.e., creating another conjoined category for a child model when a stopping condition has not been met)); updating a set of child categories to be output by the first child machine learning model based on the second conjoined category; and retraining the first child machine learning model based on the updated set of child categories to be output by the first child machine learning model (Ding para. [0045] recites “At block 440, the ML Application 235 determines whether there is at least one additional unsatisfactory class for which training data has not been identified and retrieved. If so, the method 400 returns to block 430. Otherwise, the method 400 continues to block 445. At block 445, the ML Application 235 uses the identified subset of the training data set (e.g., the exemplars identified at block 435) to train a secondary classifier. In some embodiments, the method 400 then proceeds to block 460 to finalize the model architecture”. Ding para. [0046] recites “At block 455, the ML Application 235 similarly determines whether the secondary classifier has unsatisfactory performance with respect to any of the classes. If so, in one embodiment, the method 400 returns to block 430. This process can be iterated until all classes are adequately predicted. In another embodiment, the ML Application 235 requests additional data and retrains the secondary classifier until performance is adequate” (i.e., updating or retraining at least a first child machine learning model with updated potential categories to be output by the updated child machine learning model)). Regarding claim 11, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 10, wherein determining that the performance criterion is not satisfied comprises: generating a confusion matrix based on the result of the classifying by the first child machine learning model; and determining that the performance criterion is not satisfied based on the confusion matrix (Ding para. [0035] recites “the Credibility Component 250 evaluates the quality or performance of the ML Models 255 in order to instruct the ML Component 245 how to proceed. For example, in one embodiment, the Credibility Component 250 generates a confusion matrix by evaluating one or more test sets of data using the initial ML Model 255, in order to identify class(es) that demonstrate sufficient performance and/or inadequate performance”. Ding para. [0036] recites “the Credibility Component 250 compares each determined class-specific quality to predefined threshold(s) or criteria (e.g., provided by a user) in order to identify underperforming classes (e.g., classes which the ML Model 255 has not adequately learned). The Credibility Component 250 can then return an indication of these classes, such that the ML Component 245 can train one or more additional ML Models 255 to better-evaluate input data. In an embodiment, once these classifiers are trained, the Credibility Component 250 similarly evaluates each in order to determine whether additional training and/or models are required”. Ding para. [0046] recites “At block 455, the ML Application 235 similarly determines whether the secondary classifier has unsatisfactory performance with respect to any of the classes. If so, in one embodiment, the method 400 returns to block 430. This process can be iterated until all classes are adequately predicted. In another embodiment, the ML Application 235 requests additional data and retrains the secondary classifier until performance is adequate” (i.e., using a confusion matrix to determine whether a performance criterion has been met by a retrained child machine learning model)). Regarding claim 12, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 10, wherein determining that the performance criterion is not satisfied comprises: generating a plurality of confidence levels for respective data elements in the one or more conjoined categories; and determining that the performance criterion is not satisfied based on the confidence levels (Ding para. [0048] – [0050] recite “the ML Application 235 evaluates the credibility of the output based on the confidence score generated by the main classifier. If, at block 515, the ML Application 235 determines that the output is credible, the method 500 proceeds to block 530, where the ML Application 235 returns the generated classification as the final output of the architecture (e.g., as Output Classification 120B). In contrast, if the ML Application 235 determines that the output is not credible (e.g., because it belongs to a suspect class and/or because the confidence was below a defined threshold), the method 500 continues to block 520. At block 520, the ML Application 235 determines whether the hierarchical architecture includes a secondary classifier that has been trained to recognize the class predicted by the main classifier. If not, the method 500 proceeds to block 530 where the generated output is returned. If a secondary classifier is available, the method 500 continues to block 525 where the ML Application 235 processes the received input using this secondary classifier (e.g., Secondary Classifier 115). The method 500 then returns to block 515 to perform a credibility assessment, as discussed above. In some embodiments, this process repeats until a credible classification is returned, or no additional classifiers remain” (i.e., determining whether a performance criterion has been met based on the confidence levels for different models in the tree structure)). Regarding claim 13, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 10, wherein determining that the stopping condition is not met comprises: determining that the result of the classifying by the first child machine learning model is an improvement over the result of the classifying by the top-level machine learning model; and in response to determining that the result of the classifying by the first child machine learning model is an improvement, determining that the stopping condition is not met (Ding para. [0021] recites “in addition to an output classification, each model further provides a confidence in this prediction. In some embodiments, the Credibility Assessments 110A-N evaluate this confidence for a given Input 102 in order to determine whether to return the classification, or to provide the Input 102 to a Secondary Classifier 115”. Ding para. [0046] recites “At block 455, the ML Application 235 similarly determines whether the secondary classifier has unsatisfactory performance with respect to any of the classes. If so, in one embodiment, the method 400 returns to block 430. This process can be iterated until all classes are adequately predicted” (i.e., determining whether the classification results of the child machine learning model are an improvement over the results from the top-level machine learning model, and continuing the process of creating and utilizing child models until no further improvements over unsatisfactory performance are determined)). Regarding claim 15, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the one or more conjoined categories includes a first conjoined category, and wherein the operations further comprise: adding a second conjoined category into the updated set of possible categories (Ding para. [0023] recites “the system has trained a Group Classifier 155 to classify the Input 150 into the appropriate group, such that the Input 150 can be routed to the appropriate classifier. To do so, in one embodiment, the system identifies training exemplars corresponding to each group and labels them appropriately (e.g., labeling exemplars corresponding to classes "A" and "B" with a first group label, and labeling exemplars corresponding to classes "C" and "D" with a second group label). The Group Classifier 155 can then be trained on this group-labelled data. Similarly, the Main Classifiers 160A-B can be trained using the exemplars belonging to the corresponding classes, as discussed above. That is, continuing the above example, the Main Classifier 160A is trained based on data corresponding to classes "A" and "B," while the Main Classifier 160B is trained based on data corresponding to classes "C" and "D"” (i.e., adding a second conjoined category to the list of possible categories)). Regarding claim 16, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 15, wherein the operations further comprise: deleting other possible categories, joined by the second conjoined category, from the updated set of possible categories (Yan para. [0056] recites “In operation 710, the pretrain module 230 trains a coarse category CNN on a set of coarse categories. For example, the set of coarse categories may have been identified using the process 600. The fine categories of the training dataset are replaced with the coarse categories using the mapping P(y)=y'. In an example embodiment, the dataset {xi , y’i} , for i in the range of 1 to N, is used to train a standard deep CNN model. The trained model becomes the coarse category component of the HD-CNN (e.g., the coarse category CNN 520)” (i.e., replacing a single category with a conjoined category). Ding para. [0023] recites “the system has trained a Group Classifier 155 to classify the Input 150 into the appropriate group, such that the Input 150 can be routed to the appropriate classifier. To do so, in one embodiment, the system identifies training exemplars corresponding to each group and labels them appropriately (e.g., labeling exemplars corresponding to classes "A" and "B" with a first group label, and labeling exemplars corresponding to classes "C" and "D" with a second group label). The Group Classifier 155 can then be trained on this group-labelled data. Similarly, the Main Classifiers 160A-B can be trained using the exemplars belonging to the corresponding classes, as discussed above. That is, continuing the above example, the Main Classifier 160A is trained based on data corresponding to classes "A" and "B," while the Main Classifier 160B is trained based on data corresponding to classes "C" and "D"” (i.e., replacing a single category with a conjoined category, wherein a second conjoined category can be different from a first conjoined category)). Claim 17 is a system claim and its limitation is included in claim 1. The only difference is that the preamble of claim requires “a system for generating a hierarchical classification model” rather than “a system for classifying data” as in the preamble of claim 1. Therefore, claim 17 is rejected for the same reasons as claim 1. Regarding claim 18, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 17, wherein generating the child machine learning model comprises: generating a preliminary child machine learning model; classifying, by the preliminary child machine learning model, data in the first conjoined category (Yan para. [0038] recites “FIG. 4 is a group diagram illustrating categorized images, according to some example embodiments. In FIG. 4, twenty-seven images have been correctly classified as depicting an apple (group 410), an orange (group 420), or a bus (group 430). The groups 410-430 are referred to herein as Apple, Orange, and Bus. By inspection, it is relatively easy to tell a member of Apple from a member of Bus, while distinguishing a member of Apple from a member of Orange is more difficult. Images from Apple and Orange can have a similar shape, texture, and color, so correctly telling one from the other is harder. In contrast, images from Bus often have a visual appearance distinct from those in Apple, and classification can be expected to be easier. In fact, both categories, Apple and Orange, can be defined as belonging to the same coarse category (i.e., a conjoined category) while Bus belongs to a different coarse category”. Yan para. [0039] recites “The coarse category CNN 520 and the multiple branching CNNs 540-550 each receive the input image and operate on it in parallel. Although each branching CNN 540-550 receives the input image and gives a probability distribution over the full set of fine categories, the result of each branching CNN 540-550 is only valid for a subset of categories. The multiple full predictions from branching CNNs 540-550 are linearly combined by the probabilistic averaging layer 560 to form the final fine category prediction, weighted by the corresponding coarse category probabilities” (i.e., generating a child machine learning model and classifying data in the conjoined category using the child machine learning model)); generating a second conjoined category based on a result of the classifying by the preliminary child machine learning model, the second conjoined category being generated by combining two possible categories in the set of possible categories (Ding para. [0045] recites “At block 430, the ML Application 235 selects one of the identified class(es) that has poor results. At block 435, the ML Application 235 identifies the data from the training set that corresponds to the selected class. For example, if the class "flower" performed poorly, the ML Application 235 identifies each exemplar in the training set that is labeled "flower." This subset of the training set is to be used to train supplementary classifier(s)”. Ding para. [0050] recites “If a secondary classifier is available, the method 500 continues to block 525 where the ML Application 235 processes the received input using this secondary classifier (e.g., Secondary Classifier 115). The method 500 then returns to block 515 to perform a credibility assessment, as discussed above. In some embodiments, this process repeats until a credible classification is returned, or no additional classifiers remain” (i.e., creating another conjoined category for a child machine learning model when a stopping condition has not been met)) updating the set of child categories to be output by the child machine learning model based on the second conjoined category; and retraining the preliminary child machine learning model based on the updated set of possible categories to be output by the child machine learning model (Ding para. [0045] recites “At block 440, the ML Application 235 determines whether there is at least one additional unsatisfactory class for which training data has not been identified and retrieved. If so, the method 400 returns to block 430. Otherwise, the method 400 continues to block 445. At block 445, the ML Application 235 uses the identified subset of the training data set (e.g., the exemplars identified at block 435) to train a secondary classifier. In some embodiments, the method 400 then proceeds to block 460 to finalize the model architecture”. Ding para. [0045] recites “At block 455, the ML Application 235 similarly determines whether the secondary classifier has unsatisfactory performance with respect to any of the classes. If so, in one embodiment, the method 400 returns to block 430. This process can be iterated until all classes are adequately predicted. In another embodiment, the ML Application 235 requests additional data and retrains the secondary classifier until performance is adequate” (i.e., retraining a child machine learning model with an updated set of potential categories)). Claim 20 is a method claim and its limitation is included in claim 1. The only difference is that claim 20 requires a method (para. [0006] recites “Example methods and systems are directed to hierarchical deep CNNs for image classification”). Therefore, claim 20 is rejected for the same reasons as claim 1. Claim 21 is a method claim and its limitation is included in claim 10. Claim 21 is rejected for the same reasons as claim 10. Claims 14 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Yan et al (US 20160117587 A1, herein Yan) in view of Ding et al (US 20210073686 A1, herein Ding), in further view of Cheng et al, “Hierarchical Classification of Documents with Error Control”, herein Cheng), in further view of Gonnet (US 20170220665 A1, herein Gonnet), in further view of Shotton et al (US 20150134576 A1, herein Shotton). Regarding claim 14, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 1, wherein the operations further comprise: generating a first child machine learning model, of the one or more child machine learning models, for the one or more conjoined categories (Ding para. [0045] recites “At block 440, the ML Application 235 determines whether there is at least one additional unsatisfactory class for which training data has not been identified and retrieved. If so, the method 400 returns to block 430. Otherwise, the method 400 continues to block 445. At block 445, the ML Application 235 uses the identified subset of the training data set (e.g., the exemplars identified at block 435) to train a secondary classifier. In some embodiments, the method 400 then proceeds to block 460 to finalize the model architecture”. Yan para. [0053] recites “The coarse category identification module 220 clusters (in operation 650) the C fine categories into C' coarse categories. The clustering may be performed using affinity propagation, k-means clustering, or other clustering algorithms” (i.e., generating a child machine learning model, wherein a child machine learning model can classify data in a coarser, or conjoined category)). However, the combination of Yan, Ding, Cheng, and Gonnet does not explicitly teach in response to determining that a stopping condition is met, deleting the first child machine learning model (para. [0065] teaches “when the stopping criteria at met at box 812 of FIG. 8, the temporary child nodes are deleted and replaced by direct branches from the parent nodes to the second layer of child nodes, according to the branching patterns of the temporary child nodes” (i.e., deleting a child machine learning model when a stopping criterion is met)). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by utilizing the method of deleting child models from Shotton to modify the child models from the hierarchical model in Yan (as modified by Ding, Cheng, and Gonnet). Yan and Shotton are both directed to methods of building out hierarchical classification models (see the hierarchical model structure in at least figure 3 of Shotton). One of ordinary skill in the art would recognize that utilizing the ability to delete child models from Shotton would provide greater flexibility to find a best path from root to leaf in the model and modify the hierarchical structure from Yan when necessary. Regarding claim 19, the combination of Yan, Ding, Cheng, and Gonnet teaches the system of claim 18. However, the combination of Yan, Ding, Cheng, and Gonnet does not explicitly teach wherein the operations further comprise deleting the preliminary child machine learning model. Shotton teaches wherein the operations further comprise deleting the preliminary child machine learning model (para. [0065] teaches “when the stopping criteria at met at box 812 of FIG. 8, the temporary child nodes are deleted and replaced by direct branches from the parent nodes to the second layer of child nodes, according to the branching patterns of the temporary child nodes (i.e., deleting a temporary, or preliminary, child machine learning model)). See claim 14 for motivation to combine. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20200210721 A1 (Goel et al) teaches a method for refining a hierarchical classification model such that sub-class models may be skipped. US 8296330 B2 (Bennett et al) teaches a hierarchical model that uses knowledge from children and cousin models from the bottom of the hierarchy to classify items at the parent and uses a top-down approach is used to further refine the classification of items once the top of the hierarchy is reached. US 20050188108 A1 (Carter et al) teaches a hierarchical tree model with augmentation to enable alternative paths and includes a skip-parent module. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEAH M FEITL whose telephone number is (571) 272-8350. The examiner can normally be reached on M-F 0900-1700 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /L.M.F./ Examiner, Art Unit 2147 /MARC S SOMERS/ Primary Examiner, Art Unit 2159
Read full office action

Prosecution Timeline

Show 22 earlier events
Sep 23, 2025
Response Filed
Jan 12, 2026
Final Rejection mailed — §103
Feb 10, 2026
Interview Requested
Feb 19, 2026
Examiner Interview Summary
Feb 19, 2026
Applicant Interview (Telephonic)
Mar 04, 2026
Request for Continued Examination
Mar 13, 2026
Response after Non-Final Action
Sep 04, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705248
PREDICTION OF AN EVENT AFFECTING A PHYSICAL SYSTEM
7y 7m to grant Granted Aug 11, 2026
Patent 12682009
CONTENT TARGETING USING CONTENT CONTEXT AND USER PROPENSITY
5y 6m to grant Granted Jul 14, 2026
Patent 12670304
METHODS AND APPARATUSES FOR RESOURCE-OPTIMIZED FERMIONIC LOCAL SIMULATION ON QUANTUM COMPUTER FOR QUANTUM CHEMISTRY
3y 6m to grant Granted Jun 30, 2026
Patent 12619874
Stochastic Gradient Boosting For Deep Neural Networks
2y 2m to grant Granted May 05, 2026
Patent 12572720
METHODS AND APPARATUSES FOR RESOURCE-OPTIMIZED FERMIONIC LOCAL SIMULATION ON QUANTUM COMPUTER FOR QUANTUM CHEMISTRY
5y 0m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

7-8
Expected OA Rounds
23%
Grant Probability
28%
With Interview (+5.8%)
4y 3m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 93 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month