Prosecution Insights
Last updated: October 02, 2026
Application No. 17/932,212

CLASSIFICATION USING A MACHINE LEARNING MODEL TRAINED WITH TRIPLET LOSS

Non-Final OA §101§103
Filed
Sep 14, 2022
Examiner
BEVERIDGE, CONNOR HAMMOND
Art Unit
1687
Tech Center
1600 — Biotechnology & Organic Chemistry
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Non-Final)
0%
Grant Probability
At Risk
2-3
OA Rounds
1m
Est. Remaining
0%
With Interview

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 1 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
4y 2m
Avg Prosecution
32 currently pending
Career history
21
Total Applications
across all art units

Statute-Specific Performance

§101
30.1%
-9.9% vs TC avg
§103
59.5%
+19.5% vs TC avg
§102
3.3%
-36.7% vs TC avg
§112
6.5%
-33.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of the Claims Claims 1-7, 9-12, 14-18 20-23 are currently pending and under exam herein. Claims 1-7, 9-12, 14-18 20-23 are rejected. Priority The instant application does not claim priority from another application. Therefore, the effective filing date of the instant application is 9/14/2022. Drawings The Drawings filed on 9/14/2022 were considered. Information Disclosure Statement The information disclosure statement (IDS) submitted on 07/02/2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement has been considered by the examiner. Claim Objections Claims 1 and 9 are objected to as they do not contain proper claim annotations of what has been added and removed. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-7, 9-12, 14-18 20-23 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite: (a) mathematical concepts, (e.g., mathematical relationships, formulas or equations, mathematical calculations); and (b) mental processes, i.e., concepts performed in the human mind, (e.g., observation, evaluation, judgement, opinion). Subject matter eligibility evaluation in accordance with MPEP 2106: Eligibility Step 1: Claims 1-7, 9-12, 14-18 20-23 are directed to a classification using a machine learning model trained with triplet loss [Step 1: YES] Eligibility Step 2A: First it is determined in Prong One whether a claim recites a judicial exception, and if so, then it is determined in Prong Two whether the recited judicial exception is integrated into a practical application of that exception. Eligibility Step 2A Prong One: In determining whether a claim is directed to a judicial exception, examination is performed that analyzes whether the claim recites a judicial exception, i.e., whether a law of nature, natural phenomenon, or abstract idea is set forth or described in the claim. Independent claim 1 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: generating a masked sequence by masking one or more letters of an unannotated sequence string from a corpus containing unannotated sequence strings (mental process) selecting an anchor sequence from a plurality of taxonomy-annotated sequences, wherein each of the plurality of taxonomy-annotated sequences is annotated with an expected taxonomy location within a taxonomy; (mental process) selecting a taxonomically similar sequence that shares a defined number of levels of the taxonomy with the anchor sequence; (mental process) selecting a taxonomically dissimilar sequence that does not share the defined number of levels of the taxonomy with the anchor sequence; (mental process) inferring an embedding vector from the pre-trained language model for each of the anchor sequence, the taxonomically similar sequence, and the taxonomically dissimilar sequence, wherein the pre-trained language model is a component of a classifier model; (mental process) computing a triplet loss by evaluating a triplet loss function with the embedding vectors; (mental process and/or mathematical concept) inferring a taxonomy location prediction from additional layers of the classifier model for each of the embedding vectors; (mental process and/or mathematical concept) converting each expected taxonomy location of the anchor sequence, the taxonomically similar sequence, and the taxonomically dissimilar sequence to taxonomy location hash value; (mental process and/or mathematical concept) generating a total cross-entropy loss by adding the results of invoking a cross-entropy loss function on each of the taxonomy location predictions and corresponding taxonomy location hash value; (mental process and/or mathematical concept) and training the classifier model by computing a composite loss value by adding the total cross-entropy loss and the triplet loss (mental process and/or mathematical concept) and applying backpropagation to the classifier model with the composite loss value, wherein the backpropagation is applied through the additional layers and continues through the pre-trained language model (mental process and/or mathematical concept) Dependent claim 2 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein the sequence string comprises a first sequence string, further comprising: inferring a location within the taxonomy of a second sequence string by providing the second sequence string as input to the classifier model (mental process and/or mathematical concept) Dependent claim 3 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein the sequence string is encoded as a string of letters, further including: converting the sequence string to one or more numeric values using a byte paired encoding operation (mental process and/or mathematical concept) Dependent claim 4 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein the taxonomy comprises a hierarchy of classification levels (mental process and/or mathematical concept) Dependent claim 5 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: herein each expected taxonomy location is comprised of a plurality of identifiers of classification levels within the hierarchy of classification levels. (mental process and/or mathematical concept) Dependent claim 6 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein the additional layers of the classifier model include a SoftMax layer that outputs probabilities associated with locations within the taxonomy, and wherein a taxonomy location prediction is inferred for an embedding vector by selecting a highest probability location within the taxonomy from the SoftMax layer (mental process and/or mathematical concept) Dependent claim 7 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein each expected taxonomy location is converted to a single number (mental process and/or mathematical concept) Independent claim 9 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: select an anchor sequence from a plurality of taxonomy-annotated sequences, wherein each of the plurality of taxonomy-annotated sequences is annotated with an expected taxonomy location within a taxonomy, wherein the taxonomy comprises a hierarchy of classification levels and each expected taxonomy location is converted to a single number, and wherein each of the taxonomy location predictions is a single number; (mental process and/or mathematical concept) select taxonomically similar sequence that shares a defined number of levels of the taxonomy with the anchor sequence; (mental process and/or mathematical concept) select a taxonomically dissimilar sequence that does not share the defined number of levels of the taxonomy with the anchor sequence; (mental process and/or mathematical concept) infer an embedding vector from the language model for each of the anchor sequence, the taxonomically similar sequence, and the taxonomically dissimilar sequence; (mental process and/or mathematical concept) compute a triplet loss by evaluating a triplet loss function with the embedding vectors; (mental process and/or mathematical concept) infer a taxonomy location prediction from additional layers of the classifier model for each of the embedding vectors; (mental process and/or mathematical concept) converting each expected taxonomy location of the anchor sequence, the taxonomically similar sequence, and the taxonomically dissimilar sequence to taxonomy location hash value; (mental process and/or mathematical concept) generating a total cross-entropy loss by adding the results of invoking a cross- entropy loss function on each of the taxonomy location predictions and corresponding taxonomy location hash value (mental process and/or mathematical concept) and training the classifier model based on the total cross-entropy loss and the triplet loss (mental process and/or mathematical concept) Dependent claim 10 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein the sequence string comprises a first sequence string, wherein the instructions further cause the one or more processors to: infer a location within the taxonomy of a second sequence string by providing the second sequence string as input to the classifier model (mental process and/or mathematical concept) Dependent claim 11 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein the sequence string is encoded as a string of letters, wherein the instructions further cause the one or more processors to: convert the sequence string to one or more numeric values using a byte paired encoding operation (mental process and/or mathematical concept) Independent claim 13 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: providing the sequence data to a machine learning model, the machine learning model trained with masking language modeling on a corpus of polynucleotide sequences and trained with triplet loss on a corpus of labeled polynucleotide sequences, wherein the machine learning model is trained using a total loss function that sums cross-entropy loss and triplet loss; and (mental process and/or mathematical concept) receiving from the machine learning model a plurality of taxonomic classifications for the plurality of organisms. (mental process and/or mathematical concept) Dependent claim 17 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein the plurality of taxonomic classifications comprises at least one of, domain, kingdom, phylum, class, order, family, genus, or species. (this just limits the abstract idea) Dependent claim 18 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein the machine learning model comprises an attention-based bi-directional transformer layer. (mathematical concept, describes a specific math concept) Dependent claim 20 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: further comprising: generating two-dimensional representations of higher-dimensionality vectors representing the taxonomic classifications for the plurality of organisms; and displaying the two-dimensional representations on a display device. (mathematical concept) Dependent claim 21 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein each of the taxonomy location predictions is a single number (mathematical concept) Dependent claim 22 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein training the classifier model based on the total cross-entropy loss and the triplet loss comprises: computing a composite loss value by adding the total cross-entropy loss and the triplet loss; (mathematical concept) and applying backpropagation to the classifier model with the composite loss value, wherein the backpropagation is applied through the additional layers and continues through the pre-trained language model. (mathematical concept) Dependent claim 23 recites the following steps which fall within the mental processes and/or mathematical concepts groupings of abstract ideas: wherein converting each expected taxonomy location to a single number comprises applying a hash function that maps each distinct taxonomy location within the hierarchy of classification levels to a unique numeric value (mathematical concept) The abstract ideas recited in the claims are evaluated under the broadest reasonable interpretation (BRI) of the claim limitations when read in light of and consistent with the specification. As noted in the foregoing section, the claims are determined to contain limitations that can practically be performed in the human mind with the aid of a pencil and paper, and therefore recite judicial exceptions from the mental process grouping of abstract ideas. Additionally, the recited limitations that are identified as judicial exceptions from the mathematical concepts grouping of abstract ideas are abstract ideas irrespective of whether or not the limitations are practical to perform in the human mind. Therefore, claims 1-7, 9-12, 14-18 20-23 recite an abstract idea as the dependent claims will inherit the abstract ideas from the independent claims. [Step 2A Prong One: YES] Eligibility Step 2A Prong Two: In determining whether a claim is directed to a judicial exception, further examination is performed that analyzes if the claim recites additional elements that when examined as a whole integrates the judicial exception(s) into a practical application (MPEP 2106.04(d)). A claim that integrates a judicial exception into a practical application will apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception. The claimed additional elements are analyzed to determine if the abstract idea is integrated into a practical application (MPEP 2106.04(d)(I); MPEP 2106.05(a-h)). If the claim contains no additional elements beyond the abstract idea, the claim fails to integrate the abstract idea into a practical application (MPEP 2106.04(d)(III)). The judicial exceptions identified in Eligibility Step 2A Prong One are not integrated into a practical application because of the reasons noted below. The additional element in independent claim 1 includes: A method comprising pre-training a language model with the masked sequence thereby creating a pre-trained language model The additional element in Independent claim 9 includes: A computing device comprising: one or more processors; a memory in communication with the one or more processors, the memory having computer-readable instructions stored thereupon which, when executed by the one or more processors, cause the computing device to: train a language model of a classifier model with a sequence string; The additional element in Independent claim 13 includes: A method of performing metagenomic analysis on an environmental sample, the method comprising: collecting the environmental sample, the environmental sample comprising genomic material from a plurality of organisms generating sequence data from the genomic material, wherein the sequence data is not correlated with taxonomic categories of the plurality of organisms The additional element in dependent claim 14 includes: wherein the sequence data comprises at least 1,000,000 unique reads (this limits the data gathered) The additional element in dependent claim 15 includes: wherein the environmental sample is collected from a body of water, soil, or the digestive tract of a vertebrate. The additional element in dependent claim 16 includes: wherein the sequence data comprises sequences of 16S ribosomal RNA genes. The additional elements of collecting the environmental sample, the environmental sample comprising genomic material from a plurality of organisms (Claim 13), generating sequence data from the genomic material, wherein the sequence data is not correlated with taxonomic categories of the plurality of organisms (claim 13), wherein the sequence data comprises at least 1,000,000 unique reads (Claim 14), wherein the environmental sample is collected from a body of water, soil, or the digestive tract of a vertebrate (claim 15), wherein the sequence data comprises sequences of 16S ribosomal RNA genes (Claim 16) are insignificant extra-solution activity that are part of the data gathering process used in the recited judicial exceptions (see MPEP 2106.05(g)). The additional elements of a method comprising (Claim 1), pre-training a language model with the masked sequence thereby creating a pre-trained language model (Claim 1), a computing device comprising: one or more processors; a memory in communication with the one or more processors, the memory having computer-readable instructions stored thereupon which, when executed by the one or more processors, cause the computing device to (Claim 9), train a language model of a classifier model with a sequence string; (Claim 9) A method of performing metagenomic analysis on an environmental sample, the method comprising: (Claim 13) fail to integrate a judicial exception into a practical application merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). Claims 1-7, 9-12, 14-18 20-23 do not recite any elements in addition to the judicial exception, and thus are part of the judicial exception. Thus, the additionally recited elements merely invoke a computer as a tool, and/or amount to insignificant extra-solution data gathering activity, and as such, when all limitations in claims 1-7, 9-12, 14-18 20-23 have been considered as a whole, the claims are deemed to not recite any additional elements that would integrate a judicial exception into a practical application, and therefore claims 1-7, 9-12, 14-18 20-23 are directed to an abstract idea (MPEP 2106.04(d)). [Step 2A Prong Two: NO] Eligibility Step 2B: Because the claims recite an abstract idea, and do not integrate that abstract idea into a practical application, the claims are probed for a specific inventive concept. The judicial exception alone cannot provide that inventive concept or practical application (MPEP 2106.05). Identifying whether the additional elements beyond the abstract idea amount to such an inventive concept requires considering the additional elements individually and in combination to determine if they amount to significantly more than the judicial exception (MPEP 2106.05A i-vi). The claims do not include any additional elements that are sufficient to amount to significantly more than the judicial exception(s) because of the reasons noted below. The additional elements recited in claims 1-7, 9-12, 14-18 20-23 are identified above, and carried over from Step 2A: Prong Two along with their conclusions for analysis at Step 2B. Any additional element or combination of elements that was considered to be insignificant extra-solution activity at Step 2A: Prong Two was re-evaluated at Step 2B, because if such re-evaluation finds that the element is unconventional or otherwise more than what is well-understood, routine, conventional activity in the field, this finding may indicate that the additional element is no longer considered to be insignificant; and all additional elements and combination of elements were evaluated to determine whether any additional elements or combination of elements are other than what is well-understood, routine, conventional activity in the field, or simply append well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception, per MPEP 2106.05(d). The additional elements of collecting the environmental sample, the environmental sample comprising genomic material from a plurality of organisms (Claim 13), generating sequence data from the genomic material, wherein the sequence data is not correlated with taxonomic categories of the plurality of organisms (claim 13), wherein the sequence data comprises at least 1,000,000 unique reads (Claim 14), wherein the environmental sample is collected from a body of water, soil, or the digestive tract of a vertebrate (claim 15), wherein the sequence data comprises sequences of 16S ribosomal RNA genes (Claim 16) are conventional and part of the data gathering process used in the recited judicial exceptions (see MPEP 2106.05(g)). Evidence for conventionality is shown by Liang et al. (citation below) which acquires sequencing data containing millions of reads. The additional elements of a method comprising (Claim 1), pre-training a language model with the masked sequence thereby creating a pre-trained language model (Claim 1), a computing device comprising: one or more processors; a memory in communication with the one or more processors, the memory having computer-readable instructions stored thereupon which, when executed by the one or more processors, cause the computing device to (Claim 9), train a language model of a classifier model with a sequence string; (Claim 9) A method of performing metagenomic analysis on an environmental sample, the method comprising (Claim 13) are conventional fail to integrate a judicial exception into a practical application merely reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including instructions to implement an abstract idea on a computer, or merely using a computer as a tool to perform an abstract idea, as discussed in MPEP § 2106.05(f). When taken alone, all additional elements in claims 1-7, 9-12, 14-18 20-23 do not amount to significantly more than the above-identified judicial exception(s). Even when evaluated as a combination, the additional elements fail to transform the exception(s) into a patent-eligible application of that exception. Thus, claims 1-7, 9-12, 14-18 20-23 are deemed to not contribute an inventive concept, i.e., amount to significantly more than the judicial exception(s) (MPEP 2106.05(II)). [Step 2B: NO] Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claims 1-2, 4-7, 9-10, 13-15 and 17-18, 22 are rejected under 35 U.S.C. 103 as being unpatentable over Ji et al. (Ji et a. DNABERT: Pre-Trained Bidirectional Encoder Representations from Transformers Model for DNA-Language in Genome. Bioinformatics 2021, 37 (15)) in further view of Memon et al. (Memon et al. HECNet: A Hierarchical Approach to Enzyme Function Classification Using a Siamese Triplet Network. Bioinformatics 2020, 36 (17), 4583–4589.) in further view of Liang et al. (Liang et al. DeepMicrobes: Taxonomic Classification for Metagenomics with Deep Learning. NAR Genomics and Bioinformatics 2020, 2 (1)) in view of Ruan van der Merwe (TRIPLET ENTROPY LOSS: IMPROVING THE GENERALISATION OF SHORT SPEECH LANGUAGE IDENTIFICATION SYSTEMS, Ruan van der Merwe, arxiv, 3 Dec 2020) in view of Wood et al. (Wood, D. E.; Lu, J.; Langmead, B. Improved Metagenomic Analysis with Kraken 2. Genome Biology 2019, 20 (1)). The italicized text corresponds to the instant claim limitations. With respect to the limitations of Claim 1, 9, 22, Ji et al. teaches DNABERT which uses tokenized k-mer sequences as input, which also contains a CLS token (a tag representing meaning of entire sentence), a SEP token (sentence separator) and MASK tokens (to represent masked k-mers in pre-training). The input passes an embedding layer and is fed to 12 Transformer blocks. The first output among last hidden states will be used for sentence-level classification while outputs for individual masked token used for token-level classification. Et, It and Ot denote the positional, input embedding and last hidden state at token t, respectively. DNABert generates a masked version (Masked token) of the DNA sequence (pg. 2113, Figure 1 (b), A method comprising: generating a masked sequence by masking one or more letters of an unannotated sequence string from a corpus containing unannotated sequence strings; (Claim 1)) Ji et al. also teaches DNABERT-TF achieved near perfect performances (∼0.99) on binary classification of individual TFs. Using input sequences with a much wider context (500 bp), DNABERT-TF effectively distinguished the two TAp73 isoforms with an accuracy of 0.828. In summary, DNABERT-TF can accurately identify even very similar TFBSs based on the distinct context windows. DNABERT was trained to accurately predict transcription factor binding sites as a classification problem. (pg. 2116, col. 1, paragraph 2, and pg. 2113 Figure 1, pre-training a language model with the masked sequence thereby creating a pre-trained language model (Claim 1) training a language model of a classifier model with the masked sequence ; (Claim 1) wherein the pre-trained language model is a component of a classifier model (Claim 1) train a language model of a classifier model with a sequence string (110) (Claim 9). Ji et al. also teaches using a computer to train DNABert (pg. 2113, col. 2, paragraph 1, computing device comprising: one or more processors; a memory in communication with the one or more processors, the memory having computer-readable instructions stored thereupon which, when executed by the one or more processors, cause the computing device to (Claim 9) Ji et al. also teaches details of architecture and characteristics of DNABERT model.(a) Differences between RNN, CNN and Transformer in understanding contexts. T1 to 5 denotes embedded tokens which were input into models to develop hidden states (white boxes, orange box is the current token of interest). RNN propagates information through all hidden states, and CNN takes local information in developing each representation. In contrast, Transformers develop global contextual embedding via self-attention. (b) DNABERT uses tokenized k-mer sequences as input, which also contains a CLS token (a tag representing meaning of entire sentence), a SEP token (sentence separator) and MASK tokens (to represent masked k-mers in pre-training). The input passes an embedding layer and is fed to 12 Transformer blocks. The first output among last hidden states will be used for sentence-level classification while outputs for individual masked token used for token-level classification. Et, It and Ot denote the positional, input embedding and last hidden state at token t, respectively. (c) DNABERT adopts general-purpose pre-training which can then be fine-tuned for multiple purposes using various task-specific data. (d) Example overview of global attention patterns across 12 attention heads showing DNABERT correctly focusing on two important regions corresponding to known binding sites within sequence (boxed regions, where self-attention converged) (Figure 1, Caption) DNA Bert performs backpropagation and they pre-trained DNABERT with a composite loss that includes cross-entropy loss. computing a composite loss value by adding the total cross-entropy loss and the triplet loss; and applying backpropagation to the classifier model with the composite loss value, wherein the backpropagation is applied through the additional layers and continues through the pre-trained language model (Claim 1) and applying backpropagation to the classifier model with the composite loss value, wherein the backpropagation is applied through the additional layers and continues through the pre-trained language model. (Claim 22) With respect to the limitations of Claim 13, Ji et al. teaches that for each sequence, we randomly mask regions of k contiguous tokens that constitute 15% of the sequence and let DNABERT to predict the masked sequences based on the remainder (pg. 2114, col. 1, paragraph 3, providing the sequence data to a machine learning model, the machine learning model trained with masking language modeling on a corpus of polynucleotide sequences (Claim 13) With respect to the limitations of Claim 18, Ji et al. teaches DNABert and that transformers develop global contextual embedding via self-attention. BERT stands for Bidirectional encoder representations from transformers, wherein the machine learning model comprises an attention-based bi-directional transformer layer (Claim 18) Ji et al. does not explicitly teach from a plurality of taxonomy-annotated sequences, wherein each of the plurality of taxonomy-annotated sequences is annotated with an expected taxonomy location within a taxonomy (Claim 1). selecting an anchor sequence (Claim 1). selecting a taxonomically similar sequence that shares a defined number of levels of the taxonomy with the anchor sequence ; (Claim 1) selecting a taxonomically dissimilar sequence that does not share the defined number of levels of the taxonomy with the anchor sequence; (Claim 1) inferring an embedding vector from pre-trained the language model for each of the anchor sequence, the taxonomically similar sequence, and the taxonomically dissimilar sequence; (Claim 1) computing a triplet loss by evaluating a triplet loss function with the embedding vectors; (Claim 1)) inferring a taxonomy location prediction from additional layers of the classifier model for each of the embedding vectors converting each expected taxonomy location of the anchor sequence the taxonomically similar sequence , and the taxonomically dissimilar sequence to taxonomy location hash value; generating a total cross-entropy loss by adding the results of invoking a cross-entropy loss function on each of the taxonomy location predictions and corresponding taxonomy location hash value ; and training the classifier model based on the total cross-entropy loss and the triplet loss . (Claim 1) wherein the sequence string comprises a first sequence string, further comprising: inferring a location within the taxonomy of a second sequence string by providing the second sequence string as input to the classifier model (Claim 2) wherein the taxonomy comprises a hierarchy of classification levels (Claim 4). wherein each expected taxonomy location is comprised of a plurality of identifiers of classification levels within the hierarchy of classification levels (Claim 5). wherein the additional layers of the classifier model include a SoftMax layer that outputs probabilities associated with locations within the taxonomy, and wherein a taxonomy location prediction is inferred for an embedding vector by selecting a highest probability location within the taxonomy from the SoftMax layer (Claim 6) select an anchor sequence from a plurality of taxonomy-annotated sequences (Claim 9) select taxonomically similar sequence that shares a defined number of levels of the taxonomy with the anchor sequence; select a taxonomically dissimilar sequence that does not share the defined number of levels of the taxonomy with the anchor sequence (Claim 9) infer an embedding vector from pre-trained the language model for each of the anchor sequence ), the taxonomically similar sequence , and the taxonomically dissimilar sequence (Claim 9) compute a triplet loss by evaluating a triplet loss function (802) with the embedding vectors (Claim 9) infer a taxonomy location prediction from additional layers of the classifier model for each of the embedding vectors; converting each expected taxonomy location of the anchor sequence, the taxonomically similar sequence, and the taxonomically dissimilar sequence to taxonomy location hash value; generating a total cross-entropy loss by adding the results of invoking a cross- entropy loss function on each of the taxonomy location predictions and corresponding taxonomy location hash value; and training the classifier model based on the total cross-entropy loss and the triplet loss. (Claim 9) wherein the sequence string comprises a first sequence string, wherein the instructions further cause the one or more processors to: infer a location within the taxonomy of a second sequence string by providing the second sequence string as input to the classifier model (Claim 10) a method of performing metagenomic analysis on an environmental sample, the method comprising: collecting the environmental sample, the environmental sample comprising genomic material from a plurality of organism; generating sequence data from the genomic material, wherein the sequence data is not correlated with taxonomic categories of the plurality of organisms (Claim 13). trained with triplet loss (Claim 13) receiving from the machine learning model a plurality of taxonomic classifications for the plurality of organisms. (Claim 13) wherein the sequence data comprises at least 1,000,000 unique reads (Claim 14) wherein the environmental sample is collected from a body of water, soil, or the digestive tract of a vertebrate (Claim 15) wherein the plurality of taxonomic classifications comprises at least one of, domain, kingdom, phylum, class, order, family, genus, or species. (Claim 17) wherein training the classifier model based on the total cross-entropy loss and the triplet loss comprises: computing a composite loss value by adding the total cross-entropy loss and the triplet loss (Claim 22) With respect to the limitations of Claims 1 and 9, Memon et al. teaches ‘Hierarchical Enzyme Classification Network (HECNet) and that each branch of their machine learning model takes as input the positive, anchor and negative enzyme, respectively. Each enzyme and its subsequent features pass through the CNN, LSTM and the fully connected components and ultimately, generate the embeddings. The three embeddings generated by the three branches of STNet are used to compute the triplet. (pg. 4586 Figure3, selecting an anchor sequence (Claim 1) select an anchor sequence (Claim 9).Memon et al. also teaches that each branch of their machine learning model takes as input the positive, anchor and negative enzyme, respectively. Each enzyme and its subsequent features pass through the CNN, LSTM and the fully connected components and ultimately, generate the embeddings. The three embeddings generated by the three branches of STNet are used to compute the triplet. (pg. 4586 Figure3) and each consisted of three enzymes, two of which belonged to the same class while the third belonged to a different class. Each example is called a triplet. The enzymes were split into the training and test sets. After this, triplets of both sets were made separately thus ensuring that an enzyme in a triplet belonging to the test set would never be found in a triplet of the training set. Each triplet contains an anchor enzyme, a positive enzyme and a negative enzyme. The anchor and positive enzymes belonged to the same class while the negative enzyme belonged to a different class. Most of the classes at the fourth level had a small number of enzymes. Therefore, we used all the possible triplets (25 477 869) from the training set with the restriction that the parent class (1, 4, 5, 6, 7, sub-classes of 2 and 3) of the positive and negative enzyme is the same (pg. 4585, col. 2 paragraph 3, selecting a taxonomically similar sequence that shares a defined number of levels of the taxonomy (250) with the anchor sequence ; (Claim 1) selecting a taxonomically dissimilar sequence that does not share the defined number of levels of the taxonomy (250) with the anchor sequence (Claim 1) select taxonomically similar sequence that shares a defined number of levels of the taxonomy (250) with the anchor sequence ; select a taxonomically dissimilar sequence that does not share the defined number of levels of the taxonomy (250) with the anchor sequence (Claim 9) Memon et al. also teaches the goal of STNet is to find feature embeddings in such a way that the embeddings of the anchor and positive enzyme are close together, while the embeddings of the anchor and negative enzyme are far apart, in terms of euclidean distance (pg. 4585 col 2. Paragraph 6-4586, col 1, paragraph 1, inferring an embedding vector from pre-trained the language model for each of the anchor sequence , the taxonomically similar sequence , and the taxonomically dissimilar sequence ); (Claim 1) infer an embedding vector from the language model for each of the anchor sequence , the taxonomically similar sequence, and the taxonomically dissimilar sequence ; (Claim 9) Memon et al. also teaches the three embeddings generated by the three branches of STNet are used to compute the triplet loss (pg. 4586 Figure 3, computing a triplet loss by evaluating a triplet loss function with the embedding vectors ; (Claim 1) compute a triplet loss by evaluating a triplet loss function with the embedding vectors (Claim 9) Memon et al. also teaches that after training STNet for a class, the anchor embeddings of that class are fed into the feedforward neural network whose last layer corresponds to the number of neurons which are equal to the number of classes at level 4 of that particular class. Finally, a softmax loss function is used to train the network (pg. 4586, Figure 4) and Memon et al. also teaches weighted categorical cross entropy loss function to train the CNN and RNN modules. The penalties of the loss function were weighted according to the class distribution which meant that smaller classes were given more importance than the larger classes. During training, the training error generated was back propagated to each module. This error would weigh more on the features that improve the overall performance and less on the less significant features thus providing an end-to-end feature selection. As a result, the weights of both modules would be adjusted to adopt the change.(pg. 4585, col. 2, paragraph 3, inferring a taxonomy location prediction from additional layers (340) of the classifier model for each of the embedding vectors converting each expected taxonomy location of the anchor sequence , the taxonomically similar sequence , and the taxonomically dissimilar sequence to taxonomy location hash value ; generating a total cross-entropy loss by adding the results of invoking a cross-entropy loss function on each of the taxonomy location predictions and corresponding taxonomy location hash value ; and training the classifier model based on the total cross-entropy loss and the triplet loss (Claim 1) infer a taxonomy location prediction from additional layers (340) of the classifier model for each of the embedding vectors ; converting each expected taxonomy location of the anchor sequence , the taxonomically similar sequence , and the taxonomically dissimilar sequence to taxonomy location hash value ; generating a total cross-entropy loss by adding the results of invoking a cross- entropy loss function on each of the taxonomy location predictions and corresponding taxonomy location hash value ; and training the classifier model based on the total cross-entropy loss and the triplet loss . (Claim 9) With respect to the limitations of Claim 4, Memon et al. teaches that the loss function depends on how deep we go into the EC number’s hierarchy. The values of the hierarchical weights and margins are dependent on the class of the negative example. ‘mi’ and ‘hwi’ are dependent on this depth. Suppose we are going to level 4 and we encounter two different classes. They have the same digits up to level 3. So, the negative class, in this case, has a difference of only one digit. Hence, the penalty will be smaller and ‘hwi‘ will be 0.5 for level 4. For level 3, the penalty will be greater at 0.7. And finally, at level 2, the penalty will be at a maximum of 1. The value of margin for levels 4, 3 and 2 was assigned as 0.1, 0.15 and 0.2, respectively following the above mentioned strategy. The enzymes are classified based on a functional hierarchy. (pg. 4586, paragraph 3, col. 1, wherein the taxonomy comprises a hierarchy of classification levels (Claim 4). With respect to the limitations of Claim 5, Memon et al. teaches a novel computational approach to predict an enzyme’s function up to the fourth level of the Enzyme Commission (EC) Number. Many studies have attempted to predict an enzyme’s function. Yet, no approach has properly tackled the fourth and final level of the EC number. The fourth level holds great significance as it gives us the most specific information of how an enzyme performs its function. There are multiple identifiers for each enzyme classification level (pg. 4583, paragraph 1) and a softmax loss function is used to train the network. Each class is assigned a singular neuron and therefore has a singular number (pg. 4586, Figure 4, wherein each expected taxonomy location is comprised of a plurality of identifiers of classification levels within the hierarchy of classification levels (Claim 5)). With respect to the limitations of Claims 6 and 7, Memon et al. teaches that after training STNet for a class, the anchor embeddings of that class are fed into the feedforward neural network whose last layer corresponds to the number of neurons which are equal to the number of classes at level 4 of that particular class. Finally, a softmax loss function is used to train the network. Each class is assigned a singular neuron and therefore has a singular number. When viewed in light of DNAbert which does use a composite loss a person of ordinary skill in the art would (pg. 4586, Figure 4, wherein the additional layers of the classifier model include a SoftMax layer that outputs probabilities associated with locations within the taxonomy, and wherein a taxonomy location prediction is inferred for an embedding vector by selecting a highest probability location within the taxonomy from the SoftMax layer (Claim 6) wherein each expected taxonomy location is converted to a single number (Claim 7). wherein the classifier model is trained by applying backpropagation to the classifier model with a composite loss value, wherein the composite loss value is computed by adding the total cross-entropy loss and the triplet loss (Claim 8) With respect to the limitations of Claim 13, Memon et al. teaches the three embeddings generated by the three branches of STNet are used to compute the triplet loss (pg. 4586 Figure 3, with triplet loss (Claim 13) With respect to the limitations of Claims 1 and 9, Liang et al. teaches that machine learning techniques provide a possible solution to bypass the curation of a taxonomic tree. Previous machine learning algorithms for taxonomic classification mainly utilize handcrafted sequence composition features such as oligonucleotide frequency and another classification algorithm DeepMicrobes was trained on the previously defined complete bacterial repertoire of the human gut microbiota. The repertoire is composed of 2505 species, most of which are identified by meta genome assembly of human gut microbiomes. DeepMicrobes performs taxonomic classification on the species and genus level (pg. 1 col. 2, paragraph 2-3 and pg. 2 Figure 1, from a plurality of taxonomy-annotated sequences, wherein each of the plurality of taxonomy-annotated sequences is annotated with an expected taxonomy location within a taxonomy (250) (Claim 1)), from a plurality of taxonomy-annotated sequences (260), wherein each of the plurality of taxonomy-annotated sequences (260) is annotated with an expected taxonomy location within a taxonomy (250) (Claim 9)) With respect to the limitations of Claims 2 and 10, Liang et al. teaches the method of reanalyzing a gut microbiome dataset from the Integrative Human Microbiome Project (iHMP) using DeepMicrobes and discovered potential uncultured species signatures in inflammatory bowel diseases in order to classify taxonomy (pg. 2, col. 1, paragraph 1, wherein the sequence string comprises a first sequence string, further comprising: inferring a location within the taxonomy of a second sequence string by providing the second sequence string as input to the classifier model (Claim 2) wherein the sequence string comprises a first sequence string, wherein the instructions further cause the one or more processors to: infer a location within the taxonomy of a second sequence string by providing the second sequence string as input to the classifier model (Claim 10). With respect to the limitations of Claim 13, Liang et al. teaches that shotgun metagenomic sequencing provides unprecedented insight into the critical functional roles of microorganisms in human health and the environment. One of the fundamental analysis steps in metagenomic data interpretation is to assign individual reads to their taxon-of-origin, which is termed taxonomic classification (pg. 1, col. 1, paragraph 1, a method of performing metagenomic analysis on an environmental sample, the method comprising: collecting the environmental sample, the environmental sample comprising genomic material from a plurality of organisms; generating sequence data from the genomic material, wherein the sequence data is not correlated with taxonomic categories of the plurality of organisms (Claim 13) , Liang et al. also shows that DeepMicrobes surpasses state-of-the-art taxonomic classification tools in genus or species identification and performs at least comparably in abundance estimation on the gut-derived data. We reanalyzed a gut microbiome dataset from the Integrative Human Microbiome Project (iHMP) using DeepMicrobes and discovered potential uncultured species signatures in inflammatory bowel diseases (pg 1, col. 1, paragraph 3 - pg. 2, col. 1, paragraph 1,receiving from the machine learning model a plurality of taxonomic classifications for the plurality of organisms (Claim 13) With respect to the limitations of Claim 14, Liang et al. teaches that DeepMicrobes surpasses state-of-the-art taxonomic classification tools in genus or species identification and performs at least comparably in abundance estimation on the gut-derived data. We reanalyzed a gut microbiome dataset from the Integrative Human Microbiome Project (iHMP) using DeepMicrobes and discovered potential uncultured species signatures in inflammatory bowel diseases. The iHMP contains more than 1,000,000 reads. (pg 1, col. 1, paragraph 3 - pg. 2, col. 1, paragraph 1, wherein the sequence data comprises at least 1,000,000 unique reads (Claim 14)) With respect to the limitations of Claim 15, Liang et al. teaches using data from gut microbiome to train DeepMicrobes (pg. 6, col. 1, paragraph 3, wherein the environmental sample is collected from a body of water, soil, or the digestive tract of a vertebrate (Claim 15) With respect to the limitations of Claim 17, Liang et al. teaches that DeepMicrobes can make species and genus level predictions (pg. 3, col. 2, paragraph 4, wherein the plurality of taxonomic classifications comprises at least one of, domain, kingdom, phylum, class, order, family, genus, or species (Claim 17). With respect to the limitations of Claims 1, 13, 19, 22, Merwe et al. sums cross entropy loss and triplet loss in order to improve model performance and generalization (Equation 3, computing a composite loss value by adding the total cross-entropy loss and the triplet loss; and applying backpropagation to the classifier model with the composite loss value, wherein the backpropagation is applied through the additional layers and continues through the pre-trained language model (Claim 1) wherein the machine learning model is trained using a total loss function that sums cross-entropy loss and triplet loss (Claim 13) wherein the machine learning model is trained using a total loss function that sums cross-entropy loss and triplet loss (Claim 19). wherein training the classifier model based on the total cross-entropy loss and the triplet loss comprises: computing a composite loss value by adding the total cross-entropy loss and the triplet loss (claim 22) With respect to the limitations of Claims 9, 21, 23, Wood et al. teaches while Kraken 1 used the taxonomy provided by the user without modification, Kraken 2 makes some modifications to its internal representation of the taxonomy that causes that representation to differ from the user provided taxonomy. First, Kraken 2 finds a minimal set of nodes in the user-provided taxonomy. This minimal set consists of all nodes to which a reference sequence is assigned, as well as all of those nodes’ ancestors; vertices between nodes in this set remain as they were in the user-provided taxonomy, maintaining the tree structure in the internal representation. Kraken 2 then assigns nodes in the minimal set sequentially increasing internal taxonomy ID numbers using a breadth-first search (BFS) beginning at the root, with the root having an internal ID number of 1. This BFS provides a guarantee that ancestor nodes will have smaller internal ID numbers than their descendants; an example of this numbering is shown in Additional file 2: Figure S3. Kraken 2 stores a mapping of its internal taxonomy numbers to the exter nal taxonomy ID numbers to make its results more easily interpretable, and performs all output using the external taxonomy ID numbers. (Internal taxonomy of a Kraken 2 database) Kraken 2 begins building a CHT by first estimating the number of distinct minimizers present in the reference library for the selected values of k, ℓ,ands. This is done through a form of zeroth frequency moment estimation where Kraken 2 creates a small set structure implemented with a traditional hash table. In this set Q,we insert only the distinct minimizers that satisfy the criterion h(m) mod (Populating the Kraken 2 hash table. Kraken 2 classifies sequence fragments similarly to Kra ken 1, with modifications to facilitate minimizer- and hash-based subsampling. For each k-mer in an input se quence, Kraken 2 finds its minimizer and, if it is distinct from the previous k-mer’s minimizer, uses it as a key to probe the CHT. If the minimizer matches a key in the CHT, Kraken 2 considers the associated LCA value to be the k-mer’s LCA (Fig. 1b). Classification then proceeds in the same manner as Kraken 1, taking note of how many k-mer hits mapped to each taxon, construct ing a pruned classification tree, and using the leaf of the maximally scoring root-to-leaf path of that tree to classify the sequence. If hash-based subsampling was used to build the CHT, each minimizer has its hash code compared against the table’s maximum allowable hash code, and minimizers with higher-than-allowed hash codes are not searched against the CHT. Any k-mer containing an ambiguous nucleotide code is also not searched against the CHT. (Classification of a sequence fragment with Kraken 2, wherein the taxonomy comprises a hierarchy of classification levels and each expected taxonomy location is converted to a single number, and wherein each of the taxonomy location predictions is a single number (Claim 9, when viewed in light of Memon et al. predictions subclass via a single number this representation system can be applied to the machine learning method )wherein each of the taxonomy location predictions is a single number. (Claim 21), wherein converting each expected taxonomy location to a single number comprises applying a hash function that maps each distinct taxonomy location within the hierarchy of classification levels to a unique numeric value (Claim 23) A person of ordinary skill in the art would be motivated to modify the use the masking and embedding of sequences using DNABert taught by Ji et al. with the hierarchical classification model as well as using a single number to represent a prediction taught by Memon et al. and the DeepMicrobe model classification of hierarchical classification taught by Liang et al. in order to create a better machine learning model to classify taxonomy in a hierarchical manner. A person of ordinary skill in the art would also be motivated to combine triplet and cross entropy loss to improve generalization as taught by Merwe et al. The representation of a hierarchy as a single value was taught by Wood et al. A person having ordinary skill would understand how to apply this representation to a machine learning model. A person of ordinary skill in the art is using known methods in order to improve DNABert classification and apply it to a different classification problem. Additionally, there are a limited number of machine learning and transformer architectures to try and a person of ordinary skill in the art would be motivated to combine methods from multiple machine learning models in order to optimize their classification. There is a reasonable expectation of success because they are all machine learning models that work on biological data and a person of ordinary skill in the art is just taking different components without changing how they function therefore there is a reasonable expectation they will continue to work when combined as they work individually. Claims 16 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Ji et al. in view of Memon et al. in further view of Liang et al. in view of Ruan van der Merwe in view of Wood et al. as applied to claims 1-2, 4-7, 9-10, 13-15 and 17-18, 22 above in further view of Xu et al. (Xu et al. A T-SNE Based Classification Approach to Compositional Microbiome Data. Frontiers in Genetics 2020, 11.) The italicized text corresponds to the instant claim limitations. The limitations of claims 1-2, 4-7, 9-10, 13-15 and 17-18, 22 have been taught Ji et al. in view of Memon et al. in further view of Liang et al. in view of Ruan van der Merwe in view of Wood et al. above Ji et al. in view of Memon et al. in further view of Liang et al. in view of Ruan van der Merwe in view of Wood et al. does not explicitly teach wherein the sequence data comprises sequences of 16S ribosomal RNA genes. (Claim 16) further comprising: generating two-dimensional representations of higher-dimensionality vectors representing the taxonomic classifications for the plurality of organisms; and displaying the two-dimensional representations on a display device. (Claim 20) However, these limitations were known in the art at the time of the effective filing date of the invention, as taught by Xu et al. With respect to the limitations of Claim 16, Xu et al. teaches a method to train machine learning algorithms on microbiota data were generated from the Miseq platform by sequencing the V3-V4 hypervariable region of microbial 16S rDNA ( pg. 5, col. 1, paragraph 1, wherein the sequence data comprises sequences of 16S ribosomal RNA genes. (Claim 16) With respect to the limitations of Claim 20, Xu et al. teaches the use of t-SNE to reduce dimensions for classification and visualization (pg. 5, col. 1 paragraph 2 – col. 2 paragraph 4, further comprising: generating two-dimensional representations of higher-dimensionality vectors representing the taxonomic classifications for the plurality of organisms; and displaying the two-dimensional representations on a display device. (Claim 20) A person of ordinary skill in the art would be motivated to modify the hierarchical classification model taught by Ji et al. in view of Memon et al. in further view of Liang et al. in view of Ruan van der Merwe in view of Wood et al. as applied to claims . in view of Ruan van der Merwe in view of Wood et al. above with the knowledge of Byte Pair Encoding (BPE) algorithm taught by Asgari et al. above with the knowledge of t-SNE taught by Xu et al. in order to visualize or reduce dimensions for their classification model. T-sne is a common technique in the field of sequence analysis. There is a reasonable expectation of success because t-SNE is just added to the standard workflow and works in the same manner as it does individually. In addition, t-SNE is a mathematical process and should always work to reduce dimensionality. Claims 3 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Ji et al. in view of Memon et al. in further view of Liang et al. in view of Ruan van der Merwe in view of Wood et al. as applied to claims 1-2, 4-7, 9-10, 13-15 and 17-18, 22 above in further view of Asgari et al. (Asgari et al, Nucleotide-pair encoding of 16S rRNA sequences for host phenotype and biomarker detection, bioRxiv, July 19, 2018) The italicized text corresponds to the instant claim limitations. The limitations of claims 1-2, 4-7, 9-10, 13-15 and 17-18, 22 have been taught Ji et al. in view of Memon et al. in further view of Liang et al. in view of Ruan van der Merwe in view of Wood et al. above Ji et al. in view of Memon et al. in further view of Liang et al. in view of Ruan van der Merwe in view of Wood et al. does not explicitly teach wherein the sequence string is encoded as a string of letters, further including: converting the sequence string to one or more numeric values using a byte paired encoding operation (Claim 3) wherein the instructions further cause the one or more processors to: convert the sequence string to one or more numeric values using a byte paired encoding operation (Claim 11). However, these limitations were known in the art at the time of the effective filing date of the invention, as taught by Asgari et al. teaches With respect to the limitations of Claim 3 and 11, Asgari et al. teaches that idea of Nucleotide-pair Encoding (NPE) is inspired by the Byte Pair Encoding (BPE) algorithm, a simple universal text compression scheme, which has been also used for compressed pattern matching in genomics. Although BPE had lost its popularity for a long time in compression, only recently it again became popular, but for a different reason, i.e. word segmentation in machine translation in natural language processing (NLP). BPE became a common approach for a data-driven unsupervised segmentation of words into their frequent subwords, which facilitate open vocabulary neural network machine translation and improve the quality of translation by reducing the vocabulary size. In this work, we adapt the BPE algorithm for splitting biological sequences into frequent variable length subsequences called Nucleotide-pair Encoding (NPE). We propose NPE as general purpose segmentation for the biological sequence (DNA, RNA, and proteins). In contrast to the use of BPE in NLP for vocabulary size reduction, we use this method to increase the size of symbols from 4 nucleotides to a large set of variable length biomarkers (pg. 4, paragraph 4, wherein the sequence string is encoded as a string of letters, further including: converting the sequence string to one or more numeric values using a byte paired encoding operation (Claim 3) wherein the instructions further cause the one or more processors to: convert the sequence string to one or more numeric values using a byte paired encoding operation (Claim 11)). A person of ordinary skill in the art would be motivated to modify the hierarchical classification model taught by Ji et al. in view of Memon et al. in further view of Liang et al. in view of Ruan van der Merwe in view of Wood et al. as applied to claims . in view of Ruan van der Merwe in view of Wood et al. above with the knowledge of Byte Pair Encoding (BPE) algorithm taught by Asgari et al. to improve embedding and classification accuracy. BTE is a common technique in the field of machine learning. There is a reasonable expectation of success because BTE is just added to the standard workflow and works in the same manner as it does individually. A person of ordinary skill in the art would also be motivated to try the few options such as BTE in order to improve a machine learning model. Response to Arguments Examiner finds applicant’s arguments persuasive regarding Memon et al. Examiner agrees that a new ground of rejection is needed which necessitates the office action be non-final. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Connor Beveridge whose telephone number is 571-272-2099. The examiner can normally be reached Monday - Thursday 9 am - 5 pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Karlheinz Skowronek can be reached at 571-272-9047. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /C.H.B./Examiner, Art Unit 1687 /Karlheinz R. Skowronek/Supervisory Patent Examiner, Art Unit 1687
Read full office action

Prosecution Timeline

Sep 14, 2022
Application Filed
Apr 02, 2026
Non-Final Rejection mailed — §101, §103
May 10, 2026
Interview Requested
May 21, 2026
Examiner Interview Summary
Jun 25, 2026
Response Filed
Sep 15, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
0%
Grant Probability
0%
With Interview (+0.0%)
4y 2m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month