Prosecution Insights
Last updated: October 02, 2026
Application No. 18/385,472

TESSELLATED SIMPLEXED PATTERN IDENTIFICATION PROCESSOR

Non-Final OA §101§103
Filed
Oct 31, 2023
Examiner
KASSIM, IMAD MUTEE
Art Unit
Tech Center
Assignee
Bank of America Corporation
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
130 granted / 175 resolved
+14.3% vs TC avg
Strong +31% interview lift
Without
With
+31.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
22 currently pending
Career history
194
Total Applications
across all art units

Statute-Specific Performance

§101
23.2%
-16.8% vs TC avg
§103
47.7%
+7.7% vs TC avg
§102
12.1%
-27.9% vs TC avg
§112
12.0%
-28.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 175 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 9-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non- statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because the claims recite “neural network representing a data universe”, but recite no hardware to perform the claimed steps. The claims lack the necessary physical articles or objects to constitute a machine or manufacture within the meaning of 35 USC 101. One of ordinary skill in the art may conclude that the steps associated with digital content generation may reasonably be implemented as mere software routines since no requisite computer hardware, such as a processor and memory, is recited as elements of the claimed invention The claims lack the necessary physical articles or objects to constitute a machine or manufacture within the meaning of 35 USC 101. They are clearly not a series of steps or acts to be a process nor are they a combination of chemical compounds to be a composition of matter. As such, they fail to fall within a statutory category. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The analysis of the claims’ subject matter eligibility will follow the 2019 Revised Patent Subject Matter Eligibility Guidance, 84 Fed. Reg. 50-57 (January 7, 2019) (“2019 PEG”). With respect to claim 1. Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Is the claim to a process, machine, manufacture, or composition of matter? Yes—claim 1 recites a method, which is a process. Step 2A, prong one: Does the claim recite an abstract idea, law of nature or natural phenomenon? Yes—the limitations identified below each, under its broadest reasonable interpretation, covers mental processes abstract idea grouping (concepts performed in the human mind (including an observation, evaluation, judgment, opinion)), see MPEP 2106.04(a)(2), subsection III and the 2019 PEG, but for the recitation of generic computer components: “creating a neural network that represents a data universe, wherein: the neural network comprises a plurality of neurons; each neuron within the neural network represents a data point; each of the neurons in the neural network being sorted in a hierarchical tree, the sorting within the neural network based on attributes of the data points; and the hierarchical tree includes a plurality of decision forks, each decision fork included in the hierarchical tree representing a differentiator between a data point type that categorizes the data points.”: (Mental processes- building or creating a data structure that organizes and sorts data). This falls within the mental process grouping of abstract ideas that can be performed in the human mind, or by a human with pencil and paper. Thus, Claim 1 recites an abstract idea. Step 2A, prong two: Does the claim recite additional elements that integrate the judicial exception into a practical application? No—the judicial exception is not integrated into a practical application because: The limitations, “operating a neural network on one or more processors”, “neurons”, “hierarchical tree”: Merely reciting the words “apply it” (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No—there are no additional limitations beyond the mental processes identified above. The limitation treated above, are directed to the well-understood, routine, and conventional activity of storing and retrieving information in memory. See MPEP § 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). It also includes limitations that Merely reciting the words “apply it” (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). The additional element is insignificant application, which is similar to examples of activities that the courts have found to be insignificant extra-solution activity, in accordance with MPEP 2106.05(g), Insignificant Extra-Solution Activity. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claim 2. Step 1: A method, as above. Step 2A Prong 1: The claim recites that “converting the additional data point to a neuron; and adding the neuron to the hierarchical tree at a bottom edge of the hierarchical tree.”: This limitation merely specifies mental processes- concept of observation and evaluation. Step 2A Prong 2, Step 2B: The limitations “receiving an additional data point to append to the hierarchical tree; receiving metadata relating to a categorization of the additional data point;”, involves the mere gathering of data, which is insignificant extra-solution activity. See MPEP § 2106.05(g). This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. Claim 3. Step 1: A method, as above. Step 2A Prong 1: The claim recites that “flattening out the hierarchical tree into a flattened neural network, where each decision fork in the hierarchical tree is parallel to each other decision fork in the hierarchical tree.”: This limitation merely specifies mental processes- concept of observation and evaluation of reorganizing data structure or information. Step 2A Prong 2, Step 2B: This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. Claim 4. Step 1: A method, as above. Step 2A Prong 1: The claim recites that “wherein the flattened neural network is not greater than four neuron layers”: This limitation merely specifies mental processes- concept of observation and evaluation of reorganizing data structure or information. Step 2A Prong 2, Step 2B: This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. Claim 5. Step 1: A method, as above. Step 2A Prong 1: The claim recites that “wherein the four layers comprise: a first layer corresponding to an input layer; a second layer corresponding to a decision query included in the decision fork; a third layer corresponding to a decision response included in the decision fork; and a fourth layer corresponding to an output layer that combines the decision responses from the third layer.”: This limitation merely specifies mental processes- concept of observation and evaluation of reorganizing data structure or information. Step 2A Prong 2, Step 2B: This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. Claim 6. Step 1: A method, as above. Step 2A Prong 1: The claim recites that “wherein the second layer and the third layer are hidden layers.”: This limitation merely specifies mental processes- concept of observation and evaluation of reorganizing data structure or information. Step 2A Prong 2, Step 2B: This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. Claim 7. Step 1: A method, as above. Step 2A Prong 1: mental process of claim 5. Step 2A Prong 2, Step 2B: The limitations “searching the flattened neural network using parallel processing;”, involves Merely reciting the words “apply it” (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. Claim 8. Step 1: A method, as above. Step 2A Prong 1: The claim recites that “changing an output of the flattened neural network based on a less than a predetermined amount of added data points.”: This limitation merely specifies mental processes- concept of observation and evaluation of changing output based on criteria. Step 2A Prong 2, Step 2B: This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. With respect to claim 9. Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Is the claim to a process, machine, manufacture, or composition of matter? No—claim 9 recites a neural network representing a data universe. See rejection above. Step 2A, prong one: Does the claim recite an abstract idea, law of nature or natural phenomenon? Yes—the limitations identified below each, under its broadest reasonable interpretation, covers mental processes abstract idea grouping (concepts performed in the human mind (including an observation, evaluation, judgment, opinion)), see MPEP 2106.04(a)(2), subsection III and the 2019 PEG, but for the recitation of generic computer components: “a plurality of neurons; wherein: each neuron included in the plurality of neurons represents a data point; each neuron included in the neural network is sorted in a hierarchical tree based on attributes of the data points; the hierarchical tree includes the plurality of neurons and a plurality of decision forks; and each decision fork included in the hierarchical tree represents a differentiator between a data point type that categorizes the data points.”: (Mental processes- building or creating a data structure that organizes and sorts data). This falls within the mental process grouping of abstract ideas that can be performed in the human mind, or by a human with pencil and paper. Thus, Claim 9 recites an abstract idea. Step 2A, prong two: Does the claim recite additional elements that integrate the judicial exception into a practical application? No—the judicial exception is not integrated into a practical application. The limitations, “a neural network”, “neurons”, “hierarchical tree”: Merely reciting the words “apply it” (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No—there are no additional limitations beyond the mental processes identified above. The limitation treated above, are directed to the well-understood, routine, and conventional activity of storing and retrieving information in memory. See MPEP § 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). It also includes limitations that Merely reciting the words “apply it” (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). The additional element is insignificant application, which is similar to examples of activities that the courts have found to be insignificant extra-solution activity, in accordance with MPEP 2106.05(g), Insignificant Extra-Solution Activity. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Thus, considering the additional elements individually and in combination and the claims as a whole, the additional elements do not provide significantly more than the abstract idea. This claim is not patent eligible. Claims 10-11 Step 1: The claims recite neural network representing a data universe; therefore, they do not fall into any of the four statutory categories. Step 2A Prong 1: The claims recite the same mental processes as claims 2-3, respectively. Step 2A Prong 2: This judicial exception is not integrated into a practical application. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Claim 12. Step 1: The claims recite neural network representing a data universe; therefore, they do not fall into any of the four statutory categories. Step 2A Prong 1: The claim recites that “wherein the flattened neural network replaces the neural network.”: This limitation merely specifies mental processes- concept of observation and evaluation of reorganizing data structure or information. Step 2A Prong 2, Step 2B: This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. Claims 13-17 Step 1: The claims recite neural network representing a data universe; therefore, they do not fall into any of the four statutory categories. Step 2A Prong 1: The claims recite the same mental processes as claims 4-8, respectively. Step 2A Prong 2: This judicial exception is not integrated into a practical application. Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Claim 18. Step 1: The claims recite neural network representing a data universe; therefore, they do not fall into any of the four statutory categories. Step 2A Prong 1: The claim recites that “wherein the predetermined amount of added data points is ten.”: This limitation merely specifies mental processes- concept of observation and evaluation. Step 2A Prong 2, Step 2B: This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. With respect to claim 19. Claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Is the claim to a process, machine, manufacture, or composition of matter? Yes—claim 1 recites a method, which is a process. Step 2A, prong one: Does the claim recite an abstract idea, law of nature or natural phenomenon? Yes—the limitations identified below each, under its broadest reasonable interpretation, covers mental processes abstract idea grouping (concepts performed in the human mind (including an observation, evaluation, judgment, opinion)), see MPEP 2106.04(a)(2), subsection III and the 2019 PEG, but for the recitation of generic computer components: “the neural network comprises a plurality of neurons; each neuron within the neural network represents a data point; each of the neurons in the neural network being sorted in a hierarchical tree, the sorting within the neural network based on attributes of the data points; and the hierarchical tree includes a plurality of decision forks, each decision fork included in the hierarchical tree representing a differentiator between a data point type that categorizes the data points;… converting the additional data point to a neuron; and adding the neuron to the hierarchical tree at a bottom edge of the hierarchical tree; flattening out the hierarchical tree into a flattened neural network, where each decision fork in the hierarchical tree is parallel to each other decision fork in the hierarchical tree; and replacing the neural network with the flattened neural network.”: (Mental processes- building or creating a data structure that organizes and sorts data). This falls within the mental process grouping of abstract ideas that can be performed in the human mind, or by a human with pencil and paper. Thus, Claim 19 recites an abstract idea. Step 2A, prong two: Does the claim recite additional elements that integrate the judicial exception into a practical application? No—the judicial exception is not integrated into a practical application. “receiving an additional data point to append to the hierarchical tree; receiving metadata relating to a categorization of the additional data point;” involves the mere gathering of data, which is insignificant extra-solution activity. See MPEP § 2106.05(g). The limitations, “operating a neural network on one or more processors”, “neurons”, “hierarchical tree”: Merely reciting the words “apply it” (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). The generic computer components in these steps are recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception? No—there are no additional limitations beyond the mental processes identified above. The limitation treated above, are directed to the well-understood, routine, and conventional activity of storing and retrieving information in memory. See MPEP § 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). It also includes limitations that Merely reciting the words “apply it” (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). The additional element is insignificant application, which is similar to examples of activities that the courts have found to be insignificant extra-solution activity, in accordance with MPEP 2106.05(g), Insignificant Extra-Solution Activity. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. Claim 20. Step 1: A method, as above. Step 2A Prong 1: The claim recites that “wherein the flattened neural network is not greater than four neuron layers”: This limitation merely specifies mental processes- concept of observation and evaluation of reorganizing data structure or information. Step 2A Prong 2, Step 2B: This judicial exception is not integrated into a practical application. Mere recitation of generic computer components neither integrates the judicial exception into a practical application nor provides an inventive concept. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. (“HINET: HIERARCHICAL CLASSIFICATION WITH NEURAL NETWORK”, Workshop track - ICLR 2017) in view of Mugali et al. (US 20180329935 A1). Regarding claim 1. Wu teaches a method for operating a neural network on one or more processors (see page 2, section 2.1, “whereby each node is an object with pointers to its child and parent, and takes up large memory, neural network models the connections as compact matrix which takes up much less memory.”), the method comprising: creating a neural network that represents a data universe (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”), wherein: the neural network comprises a plurality of neurons (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”, also see page 1, section 2.1, “Using layers of neurons to model hierarchy is very efficient and flexible. It can be easily used to model a Directed Acyclic Graph (DAG) (Figure 1a) or a Tree (Figure 1b) by masking out the unnecessary connections.”); each neuron within the neural network represents a data point (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”); each of the neurons in the neural network being sorted in a hierarchical tree, see page 1, section 2.1, “Using layers of neurons to model hierarchy is very efficient and flexible. It can be easily used to model a Directed Acyclic Graph (DAG) (Figure 1a) or a Tree (Figure 1b) by masking out the unnecessary connections.”); and the hierarchical tree includes a plurality of decision forks (see page 1, section 1, “hierarchical classification can be done by training a classifier on the flattened labels (Babbar et al. (2013)) or by training a classifier at each hierarchical node (Silla Jr & Freitas (2011)), whereby each hierarchical node is a decision maker of which subsequent node to route to.”, i.e. each node in the tree layer acts as a fork determine the route to next layer), each decision fork included in the hierarchical tree representing a differentiator between a data point type that categorizes the data points (see page 1, section 2, “HiNet has different procedures for training and inference. During training, as illustrated in Figure 2, the model is forced to learn MAP (Maximum a Posteriori) hypothesis over predictions at different hierarchical levels independently. Since the hierarchical layers contain shared information as child node is conditioned on the parent node, we employ a combined cost function over errors across different levels. A combined cost allows travelling of information across levels which is equivalent to transfer learning between levels.”, i.e. mathematical routing via posterior probability acts as the differentiator to categorize data points across levels). Wu do not specifically teach each of the neurons in the neural network being sorted in a hierarchical tree, the sorting within the neural network based on attributes of the data points. Mugali teaches each of the neurons in the neural network being sorted in a hierarchical tree, the sorting within the neural network based on attributes of the data points (see ¶ 36, “the various components and subsystems of the resource management and classification system 130 may be configured to receive and classify content resources into various classification hierarchies… the resource classification system 130 may receive large quantities of content resources such as documents or web pages, may use one or more classification algorithms (e.g., different combinations of machine-learning algorithms) to analyze and classify their respective resources into a taxonomies (stored as tree or hierarchy data structures) based on content classification.”, also ¶ 42, “multiple different machine-learning algorithms may be assigned to different data partitions, thus likely resulting in classification hierarchies that have different tree-structure arrangements of levels and nodes.”. also see ¶ 50, “various technical taxonomies, audience taxonomies, intent taxonomies, and the like may be generated based on various different analyses of the content resources (e.g., web pages), such as keyword based algorithms configured to analyze content and/or metadata, search engine referral classification algorithms, etc.”). Both Wu and Mugali pertain to the problem of neural network hierarchical classification, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Wu and Mugali to teach the above limitations. The motivation for doing so would be “the present disclosure describes techniques for distributed storage of network session data in hierarchical data structures stored on multiple servers and/or physical storage devices, and techniques for analyzing and classifying the distributed hierarchical structures. Such techniques may include executing different machine-learning algorithms on different servers and/or different storage devices, and generating node mapping data between a plurality of different hierarchical structures and a top-level derivative hierarchy that references the underlying hierarchical structures in order to access and manage the different distributed taxonomies within the underlying hierarchical structures.” (see Mugali Abstract). Regarding claim 2. Wu and Mugali teaches the method of claim 1, Wu further teaches see page 2, section 2.1, “whereby each node is an object with pointers to its child and parent, and takes up large memory, neural network models the connections as compact matrix which takes up much less memory. In order to model hierarchies of different length, we append a stop neuron (red neuron in Figure 1) at each layer. So a top-down path will end when it reaches the stop neuron.”. Mugali teaches further comprising: receiving an additional data point to append to the hierarchical tree; receiving metadata relating to a categorization of the additional data point (see ¶ 50, “various technical taxonomies, audience taxonomies, intent taxonomies, and the like may be generated based on various different analyses of the content resources (e.g., web pages), such as keyword based algorithms configured to analyze content and/or metadata, search engine referral classification algorithms, etc.”, also see ¶ 73, “In step 801, the resource management and classification system 130 may receive resource contents from one or more data sources. For example, content resources such as web pages or other documents may be received application servers 120, back-end storage systems 125, and/or any other data sources. In step 802, the system 130 may execute one or more machine-learning classification algorithms (and/or non-machine learning algorithms) to analyze and classify the content resources received in step 801.”); converting the additional data point to a neuron; and adding the neuron to the hierarchical tree at a bottom edge of the hierarchical tree (see ¶ 73, “In step 803, before the new content resources (e.g., web pages) may be added into the partitioned classification hierarchy 710, the resource classification system 130 first may determine which partition (or tree) 720 within the distributed tree structure will store the data corresponding to the content resource. The determination of a partition/tree in step 803 may be based on the node/branch of the content resource, determined in step 802. In step 804, the appropriate partition may be updated to reflect the new content resource. In this way, the distributed tree structure may be maintained without needing to access any of the other partitions of the distributed tree structure during the update.”, also see ¶ 60, “it could create a new logical node for the topic “basketball” at level one in the derivative taxonomy (e.g., the top-level logical hierarchy), without affecting the structures in the underlying data hierarchies.”). The motivation utilized in the combination of claim 1, super, applies equally as well to claim 2. Regarding claim 3. Wu and Mugali teaches the method of claim 1, Wu further teaches further comprising flattening out the hierarchical tree into a flattened neural network, where each decision fork in the hierarchical tree is parallel to each other decision fork in the hierarchical tree (see page 1, section 1, “Therefore for large number of categories, the hierarchy tree is flattened to produce single labels.”, also see page 4, table 1 and , section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset… We see that the number of parameters in the classification layer for Flatten Network is exponential to the maximum length of the trace. Thus for very deep hierarchies, the number of parameters in Flatten Network will be exponentially large while HiNet is always polynomial”). Regarding claim 4. Wu and Mugali teaches the method of claim 3, Wu further teaches wherein the flattened neural network is not greater than four neuron layers (see page 4, table 1 and , section 3, k: dimension of the first feature layer connected to the first hierarchical layer. n: dimension of each hierarchical layer. h: height of the hierarchy. For a fully dense hierarchy, the total number of classes is nh). Regarding claim 5. Wu and Mugali teaches the method of claim 4, Wu further teaches wherein the four layers comprise: a first layer corresponding to an input layer; a second layer corresponding to a decision query included in the decision fork; a third layer corresponding to a decision response included in the decision fork; and a fourth layer corresponding to an output layer that combines the decision responses from the third layer (see page 2, section 2, “Figure 2 shows the model for transfer learning with combined cost function. Given an input feature X with multiple levels of outputs fy(1); y(2); : : : ; y(n)g, where the outputs may have inter-level dependencies”, also section 2.3, “During inference, as illustrated in Figure 4, the model will output a normalized probability distribution at each level”, also see page 4, table 1 and , section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset). Regarding claim 6. Wu and Mugali teaches the method of claim 5, Wu further teaches wherein the second layer and the third layer are hidden layers (see page 2, section 2, “Figure 2 shows the model for transfer learning with combined cost function. Given an input feature X with multiple levels of outputs fy(1); y(2); : : : ; y(n)g, where the outputs may have inter-level dependencies”, also section 2.3, “During inference, as illustrated in Figure 4, the model will output a normalized probability distribution at each level”, also see page 4, table 1 and , section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset). Regarding claim 7. Wu and Mugali teaches the method of claim 5, Wu in page 2, section 2.1, “network models the connections as compact matrix which takes up much less memory.” Mugali further teaches further comprising searching the flattened neural network using parallel processing (see ¶ 41, “After determining one or more data partitions in step 201, steps 202-205 may be performed separately (e.g., sequentially or in parallel) for each of the data partitions.”, also ¶ 49, “For deep neural networks, training large-scale and/or complex neural networks (e.g., classifying very large sets of large and complex web pages and/or other content resources) a deep learning cluster may use multiple processors and/or servers so that the network may be trained within a reasonable time period. In some cases, such training processes may be distributed over multiple GPUs and/or CPUs, and/or may use cloud memory and/or proprietary racks.”). The motivation utilized in the combination of claim 1, super, applies equally as well to claim 7. Regarding claim 8. Wu and Mugali teaches the method of claim 5, Wu in page 2, section 2.1, “For each output level from network f_k (X) = y(k) and its corresponding label ~y(k). The combined cost is defined …allows the parameters _k from different levels to exchange knowledge.” Mugali further teaches further comprising changing an output of the flattened neural network based on a less than a predetermined amount of added data points (see ¶ 75, “in a tree-based data structure, the system 130 may determine whether a URL node already exists on a particular day. If so, the system 130 may simply increment a counter, and if not, the system 130 may ignore it. Thus, a technical advantage in such tree-based systems is that if a URL is seen multiple different times (e.g., 10 million times), it doesn't take up any more space in the data structure. Additionally, the existence of the tree path can be used to efficiently perform additional tasks, such as tracking something that happens at least N times in the tree. Additionally, tree-based joins can be used to create and update the taxonomies/partitions of the classification hierarchy.”). The motivation utilized in the combination of claim 1, super, applies equally as well to claim 8. Regarding claim 9. Wu teaches a neural network representing a data universe (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”), the neural network comprising: a plurality of neurons (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”, also see page 1, section 2.1, “Using layers of neurons to model hierarchy is very efficient and flexible. It can be easily used to model a Directed Acyclic Graph (DAG) (Figure 1a) or a Tree (Figure 1b) by masking out the unnecessary connections.”); wherein: each neuron included in the plurality of neurons represents a data point (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”); each neuron included in the neural network is see page 1, section 2.1, “Using layers of neurons to model hierarchy is very efficient and flexible. It can be easily used to model a Directed Acyclic Graph (DAG) (Figure 1a) or a Tree (Figure 1b) by masking out the unnecessary connections.”); the hierarchical tree includes the plurality of neurons and a plurality of decision forks (see page 1, section 1, “hierarchical classification can be done by training a classifier on the flattened labels (Babbar et al. (2013)) or by training a classifier at each hierarchical node (Silla Jr & Freitas (2011)), whereby each hierarchical node is a decision maker of which subsequent node to route to.”, i.e. each node in the tree layer acts as a fork determine the route to next layer); and each decision fork included in the hierarchical tree represents a differentiator between a data point type that categorizes the data points (see page 1, section 2, “HiNet has different procedures for training and inference. During training, as illustrated in Figure 2, the model is forced to learn MAP (Maximum a Posteriori) hypothesis over predictions at different hierarchical levels independently. Since the hierarchical layers contain shared information as child node is conditioned on the parent node, we employ a combined cost function over errors across different levels. A combined cost allows travelling of information across levels which is equivalent to transfer learning between levels.”, i.e. mathematical routing via posterior probability acts as the differentiator to categorize data points across levels). Wu do not specifically teach each neuron included in the neural network is sorted in a hierarchical tree based on attributes of the data points. Mugali teaches each neuron included in the neural network is sorted in a hierarchical tree based on attributes of the data points (see ¶ 36, “the various components and subsystems of the resource management and classification system 130 may be configured to receive and classify content resources into various classification hierarchies… the resource classification system 130 may receive large quantities of content resources such as documents or web pages, may use one or more classification algorithms (e.g., different combinations of machine-learning algorithms) to analyze and classify their respective resources into a taxonomies (stored as tree or hierarchy data structures) based on content classification.”, also ¶ 42, “multiple different machine-learning algorithms may be assigned to different data partitions, thus likely resulting in classification hierarchies that have different tree-structure arrangements of levels and nodes.”. also see ¶ 50, “various technical taxonomies, audience taxonomies, intent taxonomies, and the like may be generated based on various different analyses of the content resources (e.g., web pages), such as keyword based algorithms configured to analyze content and/or metadata, search engine referral classification algorithms, etc.”). Both Wu and Mugali pertain to the problem of neural network hierarchical classification, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Wu and Mugali to teach the above limitations. The motivation for doing so would be “the present disclosure describes techniques for distributed storage of network session data in hierarchical data structures stored on multiple servers and/or physical storage devices, and techniques for analyzing and classifying the distributed hierarchical structures. Such techniques may include executing different machine-learning algorithms on different servers and/or different storage devices, and generating node mapping data between a plurality of different hierarchical structures and a top-level derivative hierarchy that references the underlying hierarchical structures in order to access and manage the different distributed taxonomies within the underlying hierarchical structures.” (see Mugali Abstract). Regarding claim 10. Wu and Mugali teaches the neural network of claim 9, Wu further teaches see page 2, section 2.1, “whereby each node is an object with pointers to its child and parent, and takes up large memory, neural network models the connections as compact matrix which takes up much less memory. In order to model hierarchies of different length, we append a stop neuron (red neuron in Figure 1) at each layer. So a top-down path will end when it reaches the stop neuron.”. Mugali teaches wherein the neural network is operable to: receive an additional data point to append to the hierarchical tree; receive metadata relating to a categorization of the additional data point (see ¶ 50, “various technical taxonomies, audience taxonomies, intent taxonomies, and the like may be generated based on various different analyses of the content resources (e.g., web pages), such as keyword based algorithms configured to analyze content and/or metadata, search engine referral classification algorithms, etc.”, also see ¶ 73, “In step 801, the resource management and classification system 130 may receive resource contents from one or more data sources. For example, content resources such as web pages or other documents may be received application servers 120, back-end storage systems 125, and/or any other data sources. In step 802, the system 130 may execute one or more machine-learning classification algorithms (and/or non-machine learning algorithms) to analyze and classify the content resources received in step 801.”); convert the additional data point to a neuron; and add the neuron to the hierarchical tree at a bottom edge of the hierarchical tree based on the categorization of the additional data point (see ¶ 73, “In step 803, before the new content resources (e.g., web pages) may be added into the partitioned classification hierarchy 710, the resource classification system 130 first may determine which partition (or tree) 720 within the distributed tree structure will store the data corresponding to the content resource. The determination of a partition/tree in step 803 may be based on the node/branch of the content resource, determined in step 802. In step 804, the appropriate partition may be updated to reflect the new content resource. In this way, the distributed tree structure may be maintained without needing to access any of the other partitions of the distributed tree structure during the update.”, also see ¶ 60, “it could create a new logical node for the topic “basketball” at level one in the derivative taxonomy (e.g., the top-level logical hierarchy), without affecting the structures in the underlying data hierarchies.”). The motivation utilized in the combination of claim 9, super, applies equally as well to claim 10. Regarding claim 11. Wu and Mugali teaches the neural network of claim 9, Wu further teaches wherein the neural network is operable to convert the hierarchical tree into a flattened neural network, where each decision fork in the hierarchical tree is parallel to each other decision fork in the hierarchical tree (see page 1, section 1, “Therefore for large number of categories, the hierarchy tree is flattened to produce single labels.”, also see page 4, table 1 and , section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset… We see that the number of parameters in the classification layer for Flatten Network is exponential to the maximum length of the trace. Thus for very deep hierarchies, the number of parameters in Flatten Network will be exponentially large while HiNet is always polynomial”). Regarding claim 12. Wu and Mugali teaches the neural network of claim 11, Wu further teaches wherein the flattened neural network replaces the neural network (see page 4, section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset. From the results, HiNet out-performs Flatten Network for both a Tree hierarchical dataset with much lesser parameters. We see that the number of parameters in the classification layer for Flatten Network is exponential to the maximum length of the trace. Thus for very deep hierarchies, the number of parameters in Flatten Network will be exponentially large while HiNet is always polynomial. This makes HiNet not only better architecture in terms of accuracy but also way more efficient in parameters space.”). Regarding claim 13. Wu and Mugali teaches the neural network of claim 12, Wu further teaches wherein the flattened neural network is not greater than four neuron layers (see page 4, table 1 and , section 3, k: dimension of the first feature layer connected to the first hierarchical layer. n: dimension of each hierarchical layer. h: height of the hierarchy. For a fully dense hierarchy, the total number of classes is nh). Regarding claim 14. Wu and Mugali teaches the neural network of claim 13, Wu further teaches wherein the four neuron layers comprise: a first layer corresponding to an input layer; a second layer corresponding to a decision query included in the decision forks; a third layer corresponding to a decision response included in the decision forks, said decision response being selected from the plurality of neurons; and a fourth layer corresponding to an output layer that combines the decision responses from the third layer (see page 2, section 2, “Figure 2 shows the model for transfer learning with combined cost function. Given an input feature X with multiple levels of outputs fy(1); y(2); : : : ; y(n)g, where the outputs may have inter-level dependencies”, also section 2.3, “During inference, as illustrated in Figure 4, the model will output a normalized probability distribution at each level”, also see page 4, table 1 and , section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset). Regarding claim 15. Wu and Mugali teaches the neural network of claim 14, Wu further teaches wherein the second layer and the third layer are hidden layers (see page 2, section 2, “Figure 2 shows the model for transfer learning with combined cost function. Given an input feature X with multiple levels of outputs fy(1); y(2); : : : ; y(n)g, where the outputs may have inter-level dependencies”, also section 2.3, “During inference, as illustrated in Figure 4, the model will output a normalized probability distribution at each level”, also see page 4, table 1 and , section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset). Regarding claim 16. Wu and Mugali teaches the neural network of claim 13, Wu in page 2, section 2.1, “network models the connections as compact matrix which takes up much less memory.” Mugali further teaches wherein the flattened neural network is searched using parallel processing (see ¶ 41, “After determining one or more data partitions in step 201, steps 202-205 may be performed separately (e.g., sequentially or in parallel) for each of the data partitions.”, also ¶ 49, “For deep neural networks, training large-scale and/or complex neural networks (e.g., classifying very large sets of large and complex web pages and/or other content resources) a deep learning cluster may use multiple processors and/or servers so that the network may be trained within a reasonable time period. In some cases, such training processes may be distributed over multiple GPUs and/or CPUs, and/or may use cloud memory and/or proprietary racks.”). The motivation utilized in the combination of claim 9, super, applies equally as well to claim 16. Regarding claim 17. Wu and Mugali teaches the neural network of claim 13, Wu in page 2, section 2.1, “For each output level from network f_k (X) = y(k) and its corresponding label ~y(k). The combined cost is defined …allows the parameters _k from different levels to exchange knowledge.” Mugali further teaches wherein an output of the flattened neural network is changed based on less than a predetermined amount of added data points (see ¶ 75, “in a tree-based data structure, the system 130 may determine whether a URL node already exists on a particular day. If so, the system 130 may simply increment a counter, and if not, the system 130 may ignore it. Thus, a technical advantage in such tree-based systems is that if a URL is seen multiple different times (e.g., 10 million times), it doesn't take up any more space in the data structure. Additionally, the existence of the tree path can be used to efficiently perform additional tasks, such as tracking something that happens at least N times in the tree. Additionally, tree-based joins can be used to create and update the taxonomies/partitions of the classification hierarchy.”). The motivation utilized in the combination of claim 9, super, applies equally as well to claim 17. Regarding claim 18. Wu and Mugali teaches the neural network of claim 17, Wu in page 2, section 2.1, “For each output level from network f_k (X) = y(k) and its corresponding label ~y(k). The combined cost is defined …allows the parameters _k from different levels to exchange knowledge.” Mugali further teaches wherein the predetermined amount of added data points is ten (see ¶ 75, “in a tree-based data structure, the system 130 may determine whether a URL node already exists on a particular day. If so, the system 130 may simply increment a counter, and if not, the system 130 may ignore it. Thus, a technical advantage in such tree-based systems is that if a URL is seen multiple different times (e.g., 10 million times), it doesn't take up any more space in the data structure. Additionally, the existence of the tree path can be used to efficiently perform additional tasks, such as tracking something that happens at least N times in the tree. Additionally, tree-based joins can be used to create and update the taxonomies/partitions of the classification hierarchy.”, i.e. data point being 10 is a design choice and obvious based on teaching of Mugali of “the system 130 may determine whether a URL node already exists on a particular day. If so, the system 130 may simply increment a counter, and if not, the system 130 may ignore it. “). The motivation utilized in the combination of claim 9, super, applies equally as well to claim 18. Regarding claim 19. Wu teaches a method for operating a neural network on one or more processors (see page 2, section 2.1, “whereby each node is an object with pointers to its child and parent, and takes up large memory, neural network models the connections as compact matrix which takes up much less memory.”), the method comprising: creating a neural network (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”), wherein: the neural network comprises a plurality of neurons (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”, also see page 1, section 2.1, “Using layers of neurons to model hierarchy is very efficient and flexible. It can be easily used to model a Directed Acyclic Graph (DAG) (Figure 1a) or a Tree (Figure 1b) by masking out the unnecessary connections.”); each neuron within the neural network represents a data point (see page 1, section 1, “we model the large hierarchical labels with layers of neurons directly. Unlike the traditional structural modeling with classifier at each node, here we represent each label in the hierarchy simply as a neuron.”); each of the neurons in the neural network being sorted in a hierarchical tree, see page 1, section 2.1, “Using layers of neurons to model hierarchy is very efficient and flexible. It can be easily used to model a Directed Acyclic Graph (DAG) (Figure 1a) or a Tree (Figure 1b) by masking out the unnecessary connections.”); and the hierarchical tree includes a plurality of decision forks (see page 1, section 1, “hierarchical classification can be done by training a classifier on the flattened labels (Babbar et al. (2013)) or by training a classifier at each hierarchical node (Silla Jr & Freitas (2011)), whereby each hierarchical node is a decision maker of which subsequent node to route to.”, i.e. each node in the tree layer acts as a fork determine the route to next layer), each decision fork included in the hierarchical tree representing a differentiator between a data point type that categorizes the data points (see page 1, section 2, “HiNet has different procedures for training and inference. During training, as illustrated in Figure 2, the model is forced to learn MAP (Maximum a Posteriori) hypothesis over predictions at different hierarchical levels independently. Since the hierarchical layers contain shared information as child node is conditioned on the parent node, we employ a combined cost function over errors across different levels. A combined cost allows travelling of information across levels which is equivalent to transfer learning between levels.”, i.e. mathematical routing via posterior probability acts as the differentiator to categorize data points across levels); see page 2, section 2.1, “whereby each node is an object with pointers to its child and parent, and takes up large memory, neural network models the connections as compact matrix which takes up much less memory. In order to model hierarchies of different length, we append a stop neuron (red neuron in Figure 1) at each layer. So a top-down path will end when it reaches the stop neuron.”); flattening out the hierarchical tree into a flattened neural network, where each decision fork in the hierarchical tree is parallel to each other decision fork in the hierarchical tree (see page 1, section 1, “Therefore for large number of categories, the hierarchy tree is flattened to produce single labels.”, also see page 4, table 1 and , section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset… We see that the number of parameters in the classification layer for Flatten Network is exponential to the maximum length of the trace. Thus for very deep hierarchies, the number of parameters in Flatten Network will be exponentially large while HiNet is always polynomial”); and replacing the neural network with the flattened neural network (see page 4, section 3, “We compared HiNet with a Flatten Network which have the same architecture except the output layer for HiNet is hierarchical as illustrated in Figure 2 and flatten for Flatten Network. The number of outputs for Flatten Network corresponds to the number of classes in the dataset. From the results, HiNet out-performs Flatten Network for both a Tree hierarchical dataset with much lesser parameters. We see that the number of parameters in the classification layer for Flatten Network is exponential to the maximum length of the trace. Thus for very deep hierarchies, the number of parameters in Flatten Network will be exponentially large while HiNet is always polynomial. This makes HiNet not only better architecture in terms of accuracy but also way more efficient in parameters space.”). Wu do not specifically teach each of the neurons in the neural network being sorted in a hierarchical tree, the sorting within the neural network based on attributes of the data points; receiving an additional data point to append to the hierarchical tree; receiving metadata relating to a categorization of the additional data point; converting the additional data point to a neuron; and adding the neuron to the hierarchical tree at a bottom edge of the hierarchical tree. Mugali teaches each of the neurons in the neural network being sorted in a hierarchical tree, the sorting within the neural network based on attributes of the data points (see ¶ 36, “the various components and subsystems of the resource management and classification system 130 may be configured to receive and classify content resources into various classification hierarchies… the resource classification system 130 may receive large quantities of content resources such as documents or web pages, may use one or more classification algorithms (e.g., different combinations of machine-learning algorithms) to analyze and classify their respective resources into a taxonomies (stored as tree or hierarchy data structures) based on content classification.”, also ¶ 42, “multiple different machine-learning algorithms may be assigned to different data partitions, thus likely resulting in classification hierarchies that have different tree-structure arrangements of levels and nodes.”. also see ¶ 50, “various technical taxonomies, audience taxonomies, intent taxonomies, and the like may be generated based on various different analyses of the content resources (e.g., web pages), such as keyword based algorithms configured to analyze content and/or metadata, search engine referral classification algorithms, etc.”); receiving an additional data point to append to the hierarchical tree; receiving metadata relating to a categorization of the additional data point (see ¶ 50, “various technical taxonomies, audience taxonomies, intent taxonomies, and the like may be generated based on various different analyses of the content resources (e.g., web pages), such as keyword based algorithms configured to analyze content and/or metadata, search engine referral classification algorithms, etc.”, also see ¶ 73, “In step 801, the resource management and classification system 130 may receive resource contents from one or more data sources. For example, content resources such as web pages or other documents may be received application servers 120, back-end storage systems 125, and/or any other data sources. In step 802, the system 130 may execute one or more machine-learning classification algorithms (and/or non-machine learning algorithms) to analyze and classify the content resources received in step 801.”); converting the additional data point to a neuron; and adding the neuron to the hierarchical tree at a bottom edge of the hierarchical tree (see ¶ 73, “In step 803, before the new content resources (e.g., web pages) may be added into the partitioned classification hierarchy 710, the resource classification system 130 first may determine which partition (or tree) 720 within the distributed tree structure will store the data corresponding to the content resource. The determination of a partition/tree in step 803 may be based on the node/branch of the content resource, determined in step 802. In step 804, the appropriate partition may be updated to reflect the new content resource. In this way, the distributed tree structure may be maintained without needing to access any of the other partitions of the distributed tree structure during the update.”, also see ¶ 60, “it could create a new logical node for the topic “basketball” at level one in the derivative taxonomy (e.g., the top-level logical hierarchy), without affecting the structures in the underlying data hierarchies.”). Both Wu and Mugali pertain to the problem of neural network hierarchical classification, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Wu and Mugali to teach the above limitations. The motivation for doing so would be “the present disclosure describes techniques for distributed storage of network session data in hierarchical data structures stored on multiple servers and/or physical storage devices, and techniques for analyzing and classifying the distributed hierarchical structures. Such techniques may include executing different machine-learning algorithms on different servers and/or different storage devices, and generating node mapping data between a plurality of different hierarchical structures and a top-level derivative hierarchy that references the underlying hierarchical structures in order to access and manage the different distributed taxonomies within the underlying hierarchical structures.” (see Mugali Abstract). Regarding claim 20. Wu and Mugali teaches method of claim 19, Wu further teaches wherein the flattened neural network is not greater than four neuron layers (see page 4, table 1 and , section 3, k: dimension of the first feature layer connected to the first hierarchical layer. n: dimension of each hierarchical layer. h: height of the hierarchy. For a fully dense hierarchy, the total number of classes is nh). Related prior arts: PERERA et al. (US 20200285943 A1) teaches hierarchical classification optimizing method involves receiving training data set (302) and a hierarchical classification ontology data structure (301), generating neural network architecture based on training data set and the hierarchical classification ontology data structure, training neural network architecture to classify the training data set to leaf nodes at the lower leaf tier (LLT) output and parent nodes at the parent tier (PT) output, determining surface that passes through each path from a root to a leaf node in the hierarchical ontology data structure, and training a classifier model for a cognitive system (320) using the surface and the training data set. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to IMAD M KASSIM whose telephone number is (571)272-2958. The examiner can normally be reached 10:30AM-5:30PM, M-F (E.S.T.). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297 - 4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /IMAD KASSIM/Primary Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Oct 31, 2023
Application Filed
Jul 10, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743620
PREDICTING SOUND PLEASANTNESS USING REGRESSION PREDICTION MACHINE LEARNING MODEL
3y 10m to grant Granted Sep 22, 2026
Patent 12737674
COMBINING MODEL OUTPUTS INTO A COMBINED MODEL OUTPUT
4y 4m to grant Granted Sep 15, 2026
Patent 12711394
FEDERATED LEARNING
3y 5m to grant Granted Aug 18, 2026
Patent 12699901
FAST QUANTISED TRAINING OF TRAINABLE MODULES
4y 6m to grant Granted Aug 04, 2026
Patent 12694966
MACHINE LEARNING MODELING TO PREDICT HEURISTIC PARAMETERS FOR RADIATION THERAPY TREATMENT PLANNING
4y 10m to grant Granted Jul 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
99%
With Interview (+31.3%)
3y 8m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 175 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month