Prosecution Insights
Last updated: August 17, 2026
Application No. 17/953,293

COMBINED AND TRANSFER LEARNING OF A VARIANT PATHOGENICITY PREDICTOR USING GAPPED AND NON-GAPPED PROTEIN SAMPLES

Non-Final OA §101§103
Filed
Sep 26, 2022
Priority
Oct 06, 2021 — provisional 63/253,122 +2 more
Examiner
MOHANTA, PRAMOD KUMAR
Art Unit
1686
Tech Center
1600 — Biotechnology & Organic Chemistry
Assignee
Illumina Inc.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
3 currently pending
Career history
2
Total Applications
across all art units

Statute-Specific Performance

§101
20.0%
-20.0% vs TC avg
§103
80.0%
+40.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
lkDETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status of the Claims Claims 1-30 have been examined. Claims 1-30 are rejected. Priority Domestic priority data as claimed by applicant, this application claims benefit of 63/281,592 11/19/2021 and claims benefit of 63/281,579 11/19/2021 and claims benefit of 63/253,122 10/06/2021. The effective filing date considered is 10/06/2021. Information Disclosure Statement The information disclosure statement (IDS) submitted on 02/09/2023 and 08/03/2023 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Nucleotide and/or Amino Acid Sequence Disclosures REQUIREMENTS FOR PATENT APPLICATIONS CONTAINING NUCLEOTIDE AND/OR AMINO ACID SEQUENCE DISCLOSURES Items 1) and 2) provide general guidance related to requirements for sequence disclosures. 37 CFR 1.821(c) requires that patent applications which contain disclosures of nucleotide and/or amino acid sequences that fall within the definitions of 37 CFR 1.821(a) must contain a "Sequence Listing," as a separate part of the disclosure, which presents the nucleotide and/or amino acid sequences and associated information using the symbols and format in accordance with the requirements of 37 CFR 1.821 - 1.825. This "Sequence Listing" part of the disclosure may be submitted: In accordance with 37 CFR 1.821(c)(1) via the USPTO patent electronic filing system (see Section I.1 of the Legal Framework for Patent Electronic System (https://www.uspto.gov/PatentLegalFramework), hereinafter "Legal Framework") as an ASCII text file, together with an incorporation-by-reference of the material in the ASCII text file in a separate paragraph of the specification as required by 37 CFR 1.823(b)(1) identifying: the name of the ASCII text file; ii) the date of creation; and iii) the size of the ASCII text file in bytes; In accordance with 37 CFR 1.821(c)(1) on read-only optical disc(s) as permitted by 37 CFR 1.52(e)(1)(ii), labeled according to 37 CFR 1.52(e)(5), with an incorporation-by-reference of the material in the ASCII text file according to 37 CFR 1.52(e)(8) and 37 CFR 1.823(b)(1) in a separate paragraph of the specification identifying: the name of the ASCII text file; the date of creation; and the size of the ASCII text file in bytes; In accordance with 37 CFR 1.821(c)(2) via the USPTO patent electronic filing system as a PDF file (not recommended); or In accordance with 37 CFR 1.821(c)(3) on physical sheets of paper (not recommended). When a “Sequence Listing” has been submitted as a PDF file as in 1(c) above (37 CFR 1.821(c)(2)) or on physical sheets of paper as in 1(d) above (37 CFR 1.821(c)(3)), 37 CFR 1.821(e)(1) requires a computer readable form (CRF) of the “Sequence Listing” in accordance with the requirements of 37 CFR 1.824. If the "Sequence Listing" required by 37 CFR 1.821(c) is filed via the USPTO patent electronic filing system as a PDF, then 37 CFR 1.821(e)(1)(ii) or 1.821(e)(2)(ii) requires submission of a statement that the "Sequence Listing" content of the PDF copy and the CRF copy (the ASCII text file copy) are identical. If the "Sequence Listing" required by 37 CFR 1.821(c) is filed on paper or read-only optical disc, then 37 CFR 1.821(e)(1)(ii) or 1.821(e)(2)(ii) requires submission of a statement that the "Sequence Listing" content of the paper or read-only optical disc copy and the CRF are identical. Specific deficiencies and the required response to this Office Action are as follows: Specific deficiency - This application fails to comply with the requirements of 37 CFR 1.821 - 1.825 because it does not contain a "Sequence Listing" as a separate part of the disclosure or a CRF of the “Sequence Listing.”. Required response - Applicant must provide: A "Sequence Listing" part of the disclosure; together with An amendment specifically directing its entry into the application in accordance with 37 CFR 1.825(a)(2); A statement that the "Sequence Listing" includes no new matter as required by 37 CFR 1.821(a)(4); and A statement that indicates support for the amendment in the application, as filed, as required by 37 CFR 1.825(a)(3). If the "Sequence Listing" part of the disclosure is submitted according to item 1) a) or b) above, Applicant must also provide: A substitute specification in compliance with 37 CFR 1.52, 1.121(b)(3) and 1.125 inserting the required incorporation-by-reference paragraph, consisting of: A copy of the previously-submitted specification, with deletions shown with strikethrough or brackets and insertions shown with underlining (marked-up version); A copy of the amended specification without markings (clean version); and A statement that the substitute specification contains no new matter. If the "Sequence Listing" part of the disclosure is submitted according to item 1) c) or d) above, applicant must also provide: A CRF in accordance with 37 CFR 1.821(e)(1) or 1.821(e)(2) as required by 1.825(a)(5); and A statement according to item 2) a) or b) above. Specific deficiency – Nucleotide and/or amino acid sequences appearing in the drawings are not identified by sequence identifiers in accordance with 37 CFR 1.821(d). Sequence identifiers for nucleotide and/or amino acid sequences must appear either in the drawings or in the Brief Description of the Drawings. Required response – Applicant must provide: Replacement and annotated drawings in accordance with 37 CFR 1.121(d) inserting the required sequence identifiers; AND/OR A substitute specification in compliance with 37 CFR 1.52, 1.121(b)(3) and 1.125 inserting the required sequence identifiers into the Brief Description of the Drawings, consisting of: A copy of the previously-submitted specification, with deletions shown with strikethrough or brackets and insertions shown with underlining (marked-up version); A copy of the amended specification without markings (clean version); and A statement that the substitute specification contains no new matter. Drawings The drawings are objected because figures 38-41 and 62-63 include enumerated amino acid sequences without providing sequence identification number (see above). Corrected drawing sheets in compliance with 37 CFR l.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as 'amended." If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either ''Replacement Sheet'' or ''New Sheet'' pursuant to 37 CFR l.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Subject Matter Eligibility Analysis Step 1: Statutory Category? Claims 1, 26 and 30 recite a series of steps and therefore, is a process. (Step 1, Yes). Step 2A - Prong One: Judicial Exception Recited? Independent claim 1 is directed to a computer-implemented method of training a pathogenicity predictor, which recites the following active steps : accessing a gapped training set that includes respective gapped protein samples, which is an additional element pertaining to judicial exception. Accessing a non-gapped training set that includes non-gapped benign protein samples and non-gapped pathogenic protein samples, which is an additional element pertaining to judicial exception. Generating gapped spatial representations for the gapped protein samples, the step is performed using mathematical algorithm, therefore is a mathematical concept. Generating respective non- gapped spatial representations for the non-gapped benign protein samples, the step is performed using mathematical algorithm, therefore is a mathematical concept. Training a pathogenicity predictor over one or more training cycles, which is an additional information pertaining to abstract idea. Generating a trained pathogenicity predictor using gapped spatial representations, which is viewed as training a machine learning model involving mathematical modeling, therefore is a mathematical concept. Using the trained pathogenicity classifier to determine pathogenicity of variants, which is an additional information pertaining to judicial exception. Claim 2 recites gapped protein samples are labelled with respective gapped ground truth sequences, which is an additional information pertaining to judicial exception. Claim 3 recites a particular gapped ground truth sequence for a particular gapped protein sample has a benign label for a particular amino acid class, which is an additional information pertaining to judicial exception. Claim 4 recites particular gapped protein sample has respective pathogenic labels for respective remaining amino acid classes, which is an additional information pertaining to judicial exception. Claim 5 recites a particular non-gapped benign protein sample includes a benign alternate amino acid at a particular position substituted by a benign nucleotide variant, which is an additional information pertaining to judicial exception. Claim 6 recites a particular non-gapped pathogenic protein sample includes a pathogenic alternate amino acid at a particular position substituted by a pathogenic nucleotide variant, which is an additional information pertaining to judicial exception. Claim 7 recites particular non-gapped benign protein sample is labelled with a benign ground truth sequence that has a benign label for a particular amino acid class that corresponds to the benign alternate amino acid, which is an additional information pertaining to judicial exception. Claim 8 recites the benign ground truth sequence respective masked labels for respective remaining amino acid classes that correspond to amino acids that are different from the benign alternate amino acid, which is an additional information pertaining to judicial exception. Claim 9 recites the particular non-gapped pathogenic protein sample is labelled with a pathogenic ground truth sequence that has a pathogenic label for a particular amino acid class that corresponds to the pathogenic alternate amino acid. Claim 10 recites the pathogenic ground truth sequence has respective masked labels for respective remaining amino acid classes that correspond to amino acids that are different from the pathogenic alternate amino acid, which is an additional information pertaining to judicial exception. Claim 11 recites using a sample indicator to indicate to the pathogenicity predictor whether a current training example is a gapped spatial representation for a gapped protein sample, or a non-gapped spatial representation for a non-gapped protein sample. The step involves indicating information to a model, can theoretically be performed in the human mind or with a pen and paper. Therefore, is a mental step. Claim 12 recites masking the benign label for the particular amino acid class that corresponds to the reference amino acid at the particular position in the particular gapped protein, the concept involves mathematical algorithm, therefore is a mathematical step. Claim 13 recites non-gapped benign protein samples are derived from common human and non-human primate nucleotide variants, which is an additional information pertaining to abstract ideas. Claim 14 recites non-gapped pathogenic protein samples are derived from combinatorically simulated nucleotide variants, which is an additional information pertaining to abstract ideas. Claim 15 recites pathogenicity predictor generates an amino acid class-wise output sequence in response to processing a training example, the concept involves mathematical calculation, therefore is a mathematical concept. Claim 16 recites measuring performance of the trained pathogenicity predictor (model) between training cycles over a validation set, which is a mathematical concept of calculation. Claim 17 recites validation set includes a pair of gapped and non-gapped spatial representations for each held-out protein sample, which is an additional information pertaining to judicial exception. Claim 18 recites trained pathogenicity predictor generates a first amino acid class-wise output sequence for the gapped spatial representation in the pair, and a second amino acid class-wise output sequence for the non-gapped spatial representation, Which are viewed as a mathematical concepts because trained pathogenicity predictor (trained model) generates output using spatial data. Claim 19 recites final pathogenicity score is based on an average of the first and second pathogenicity scores, which is viewed as averaging numbers , therefore is a mathematical concept of calculation. Claim 20 recites some of the training cycles use a same of number of gapped spatial representations and non-gapped spatial representations, which is an additional information pertinent to judicial exception. Claim 21 recites some of the training cycles use batches of training examples that have a same of number of gapped spatial representations and non-gapped spatial representations, which is an additional information pertinent to judicial exception. Claim 22 recites a masked label does not contribute to error determination, and therefore does not contribute to training of the pathogenicity predictor. which is an additional information pertinent to judicial exception. Claim 23 recites the masked label is zeroed-out, which is an additional information pertinent to judicial exception. Claim 24 recites gapped spatial representations are weighted differently from the non-gapped spatial representations, which is an additional element pertinent to judicial exception. Processing the non-gapped spatial representations varies from a contribution of the non-gapped spatial representations to gradient updates applied to the parameters of the pathogenicity predictor in response to the pathogenicity predictor processing the non-gapped spatial representations, which is an additional element pertinent to judicial exception. Claim 25 recites variation is determined by pre-defined weights, which is an additional element pertinent to judicial exception. Independent claim 26 is directed to a method, which includes following active steps: Training a pathogenicity classifier (model) on a gapped training set, which is providing instruction to pathogenicity classifier (model), therefore is viewed as an additional element pertinent to judicial exception. Generating a trained pathogenicity classifier, the step involves mathematical algorithm, therefore is a mathematical concept. Training the trained pathogenicity classifier on a non-gapped training set, which is providing instruction to trained pathogenicity classifier (model), therefore is viewed as an additional element pertinent to judicial exception. Generating a retrained pathogenicity classifier (model), which is viewed as mathematical algorithm, therefore is a mathematical concept. Determine pathogenicity of variants using the retrained pathogenicity classifier to, which is an additional element pertaining to judicial exception. Claim 27 recites measuring performance of the trained pathogenicity predictor between training cycles over a first validation set , which is an additional element pertaining to judicial exception. Claim 28 recites including measuring performance of the retrained pathogenicity predictor between training cycles over a second validation set, which is an additional element pertaining to judicial exception. Claim 29 recites the retrained pathogenicity predictor generates a first amino acid class-wise output sequence for the pair in response to processing the pair, which is viewed as generating an output using a mathematical model/algorithm, therefore is a mathematical concept. Independent claim 30 is directed to a computer-implemented method, which includes following active steps: Training a pathogenicity predictor (model), which is an additional element pertaining to the judicial exception. Accessing a gapped training set that includes respective gapped protein samples, which is an additional element pertaining to the judicial exception. Accessing a non-gapped training set that includes non-gapped benign protein samples and non-gapped pathogenic protein samples, which is an additional element pertaining to the judicial exception. Generating respective gapped spatial representations for the gapped protein samples, and generating respective non- gapped spatial representations for the non-gapped benign protein samples and the non-gapped pathogenic protein samples; which are additional elements pertaining to the judicial exception. Training a pathogenicity predictor over one or more training cycles, which is an additional element pertaining to judicial exception. Generating a trained pathogenicity predictor, wherein each of the training cycles uses as training examples gapped spatial representations , which is an additional element pertaining to judicial exception. Determine pathogenicity of variants using the trained pathogenicity classifier, which is an additional element. (Step 2A, Prong One, Yes). Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a whole integrates the recited judicial exception into a practical application of the exception or whether the claim is “directed to” the judicial exception. Claims 1, 26, 30, and their dependent claims recite additional element the method being “computer-implemented”. Claims also recite additional element “accessing” protein sample data for training pathogenicity classifier. The claimed limitations do not recite a processor, a storage medium, indicating a generic computing system. The above mentioned additional elements of accessing, processing data and generating an output by comparing and analyzing information equate to mere instructions to implement the abstract idea on a generic computer that the courts have stated does not render an abstract idea eligible in Alice Corp., 573 U.S. at 223, 110 USPQ2d at 1983. See also 573 U.S. at USPQ2d at 1984. The independent claims as a whole, including the judicial exception discussed above in Step 2A, Prong One, do not integrate into a practical application.(Step 2A, Prong Two: NO).Therefore the claims are directed to a judicial exception (Step 2A: YES) and thus requires further analysis at Step 2B Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole amount to significantly more than the recited exception i.e., whether any additional element, or combination of additional elements, adds an inventive concept to the claim (MPEP 2106.05). The claims recite an abstract idea with additional elements (Step 2A). The independent claims in combination with additional elements are viewed as merely instructions performed in ordinary computer to implement abstract ideas. The claims, therefore do not include additional elements that are sufficient amount to significantly more than the judicial exception. As a result, the claims as a whole do not provide an inventive concept. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 4-6, 13-21, and 24-29 are rejected under 35U.S.C. 103 as being unpatentable over Sundaram et al. [Nature Genetics, 2018, Vol 50, P1161-1173 and Page S1-S61:Supplementary Information, publication date: Jul23, 2018] in the view of Won et al. [Bioinformatics, 37(24), 2021, 4626–4634, publication date: Jul16, 2021]. Regarding independent claim 1, Sundaram et al. teaches training of a deep learning network for pathogenicity prediction using training data set of protein samples (page 1161, abstract; page 1166, fig.5; page 1167, col 1, para 2 ) as in claim 1 teaches training a pathogenicity predictor using a gapped training set of gapped protein samples, and using a non-gapped training set that includes non-gapped benign protein samples and non-gapped pathogenic protein samples. Regarding claim 1, Sundaram et al. discloses silently a computer-implemented method, by using custom software and publicly available datasets, for training a deep learning network (page S4, sections: Software and code, Data; page 1161, abstract) as claim 1 teaches computer-implemented method of training machine-learning system. Regarding claim 1, Sundaram et al. teaches inputting primary amino acid sequence data and construct progressively higher order representations of protein structure using deep learning network (page 1168, para 1; page S1, col 2, section: Architecture of the deep learning network) as claim 1 recites generating higher order representations for the gapped / non-gapped protein samples Regarding claim 1, Sundaram et al. teaches repeated training of deep learning network (classifier) using training data set (page S2, col 2, para 3) as claim 1 teaches training a pathogenicity predictor over one or more training cycles and generating a trained pathogenicity predictor. Regarding claim 1, Sundaram et al. teaches using a trained deep neural network to identify pathogenic mutations in rare disease patients (page 1161, abstract) as in claim 1 teaches using the trained pathogenicity classifier to determine pathogenicity of variants. Regarding claim 4, Sundaram et al. teaches data set labeled benign and pathogenic variants for training (page 1166-1167, section: A deep learning network for variant pathogenicity classification) as in claim 4 teaches a gapped protein sample labelled pathogenic , which correspond to alternate amino acids at the particular position. Regarding claim 5, Sundaram et al. teaches Primate AI (a network) takes as input the human amino acid (AA) reference and alternate sequence (51 AAs) centered at the variant, the position weight matrix (PWM) conservation profiles calculated from 99 vertebrate species, and the outputs of secondary structure and solvent accessibility prediction deep learning networks, (page 1164, Fig.3) as in claim 5 teaches protein sample includes nucleotide variant alternate amino acid at a particular position substituted by a nucleotide variant. Regarding claim 6, Sundaram et al. teaches inputs to the model were the position-specific frequency matrices (51 AA length, 20 depth) for the flanking amino acid sequence around the variant, the one-hot encoded human reference and alternate sequences (Page S35, Supplementary Table 16; page S53-54, section: Deep learning network architecture) as in claim 6 teaches a non-gapped pathogenic protein sample includes a pathogenic alternate amino acid at a particular position. Regarding claim 13, Sundaram et al. teaches benign protein samples are from human and non-human primate species (page 1161, abstract) as claim 13 teaches non-gapped benign protein samples are derived from common human and non-human primate nucleotide variants. Regarding claim 14, Sundaram et al. teaches simulation of common variants in each primate species (page 1166, Fig. 5). For each primate species, we simulated four times the number of common missense variants observed in human (~83,500 missense variants with allele frequency > 0.1%)-as in claim 14 teaches multiple simulations to obtain the non-gapped pathogenic protein samples. Regarding claim 15, Sundaram et al. teaches using amino acid sequence as its input, the deep learning network outputs high pathogenicity scores to residues at critical protein functional domains (page 1167, col 2, para 2) as in claim 15 teaches using training sample by pathogenicity predictor generates an amino acid class-wise output sequence with pathogenicity scores. Regarding claim 16, Sundaram et al. discloses repeating the training procedure five times and measuring the measuring the performance of trained classifier. Sundaram et al. further discloses validating sample sets of 10,000 primate variants during training (page s1, section: Impact of increasing training data size and using different sources of training Data; page 1166, Fig.5) as in claim 16 teaches measuring performance of pathogenicity predictor during repeated training over a validation set. Regarding claim 17, Sundaram et al. teaches validation set includes two randomly sampled sets of 10, 000 primate variants (page 1170-1, section: Impact of increasing training data size and using different sources of training data) as in claim 17 teaches the validation set includes two representations of protein sample, a gapped and a non- gapped spatial representations. Regarding claims 18, Sundaram et al. teaches average PrimateAI scores for variants at each amino acid position in three different protein domains as in claim 18 teaches a final pathogenicity score for a nucleotide is determined based on a combination of first and second pathogenicity scores for the amino acid substitution in the first and second amino acid output sequences. Regarding claim 19, Sundaram et al. teaches predicted pathogenicity score at each amino acid position in the SCN2A gene, annotated for key functional domains. Plotted along the gene is the average PrimateAI score for missense substitutions at each amino acid position (page 1164, Fig. 3) as claim 19 teaches a final pathogenicity score is based on an average of the first and second pathogenicity scores. Regarding claim 20, Sundaram et al. teaches using same set of same set of labeled benign variants and eight randomly sampled sets of unlabeled variants for training deep learning networks (page S52, section: Semi-supervised learning) as in claim 20 teaches some of the training cycles use a same set of gapped-and non-gapped spatial representations. Regarding claim 21, Sundaram et al. teaches input of identical data during training of deep convolutional neural network models for secondary structure (supplementary information page S51, section: Model architecture and training) as in claim 21 teaches the training cycles use examples that have identical number of gapped spatial representations and non-gapped spatial representations. Regarding claim 24, Sundaram et al. teaches correlation patterns of weights from layers of the secondary-structure DL network show correlations between amino acids that are similar to BLOSUM62 and Grantham score matrices (page S53, section: Matching the labeled benign and unlabeled training sets) as in claim 24 teaches gapped structures are weighted differently from the non-gapped structures such that a contribution of the gapped spatial representations to gradient updates applied to parameters of the pathogenicity predictor in response to the pathogenicity predictor processing the non-gapped spatial representations varies from a contribution of the non-gapped spatial representations to gradient updates applied to the parameters of the pathogenicity predictor in response to the pathogenicity predictor processing the non-gapped spatial representations. Regarding claim 25, Sundaram et al. teaches weighted the probability of sampling a variant based on the sequencing coverage at that position using the regression coefficients (page S53, section: Matching the labeled benign and unlabeled training sets) as in claim 25 teaches variation is determined by pre-defined weights. Regarding independent claim 26, Sundaram et al. teaches repeated training of deep learning network (=further training a trained classifier) (page S2, col 2, para 3) and uses the trained deep neural network to determine pathogenic mutations (page 1161, abstract) as in claim 26 teaches repeated training including training a pathogenicity classifier (model) generating a trained pathogenicity classifier and further training the trained pathogenicity classifier generating a retrained pathogenicity classifier, which is used to determine pathogenicity prediction. Regarding claim 27, Sundaram et al. teaches measuring performance of classifier during repeated training (page 1170-1, Col 2, para 2; page 1169, section: Discussion) and discloses a validation data set (page 1170-1, section: generation of benign and unlabeled variants for model training) as in claim 27 teaches measuring performance of the trained pathogenicity predictor between training cycles over a validation set of non-gapped spatial representations of protein samples. Regarding claim 28, Sundaram et al. discloses for validating and testing of the deep learning network, randomly sampled two sets of 10,000 primate variants for validation and testing, which are withheld from training. The performance of the deep learning network during the course of training is monitored by measuring the ability of the network between variants in the two sets (Supplementary Information, page S48, section: Withheld primate variants for validation and testing, and de novo variants from affected and unaffected individuals; Page S14, section: Performance evaluation of the deep learning network and other classifiers) as in claim 28 teaches measuring performance of the retrained pathogenicity predictor between training cycles over a second validation set that includes gapped spatial representations and non-gapped spatial representations of held-out protein samples . Regarding claim 29, Sundaram et al. teaches impact of data used for training on classification accuracy/performance. Outputs of network trained with different sources of common variation, benchmarked on 10,000 withheld primate variants (page 1166, fig.5) as in claim 29 teaches final pathogenicity score protein sequences sample is determined based on the first amino acid class-wise output sequence. Sundaram et al. does not teach use of gapped protein samples (data) for training pathogenicity predictor. Won et al. teaches use of variant deleted (= gapped) sequence data to train 3Cnet, a neural networks (page 4627, col 2, para 3; page 4630, col 2, para 3). It would have been obvious to a Person Having Ordinary Skill in the Art (PHOSITA) at the time of the claimed invention to combine the teachings of Sundaram et al. with those of Won et al.. Specifically, Sundaram teaches training pathogenicity classifier/predictor with non-gapped protein sequenced data, while Won teaches the use of gapped protein sequenced data for training. A PHOSITA would have been motivated to combine these teachings because training of pathogenicity classifier/predictor with gapped and non-gapped protein data sets is a predictable variation in the field of bioinformatics. Claims 2-3, 7-12, 22-23, and 30 are rejected under 35U.S.C. 103 as being unpatentable over Sundaram et al. in the view of Won et al. as applied to claims 1, 4-6, 13-21, and 24-29 above, further in view of Greiff et al. [Current Opinion in Systems Biology 2020, Vol 24, P109–119, publication date: Oct27, 2020] and Wu et al. [PLoS ONE, 2012, Vol 7(1): e30288, 1-10, online publication date: Jan17, 2012]. Sundaram et al. in the view of Won et al. are applied to claims 1, 4-6, 13-21, and 24-29 as above. Regarding claim 2, Sundaram et al. teaches a truth data set for machine learning training (page 1166-1167, section: A deep learning network for variant pathogenicity classification) as in claim 2 teaches protein samples data are labelled with gapped ground truth sequences. Regarding claim 3, Sundaram et al. teaches benign labeled dataset for machine learning training (page 1166, section: A deep learning network for variant pathogenicity classification) as in claim 3 teaches a gapped protein sample has a benign label for a particular amino acid class. Regarding claim 7, Sundaram et al. teaches generating benign labelled protein sample data sets for model training (page S38:Supplementary Table 17; page 1170-1, section: Generation of benign and unlabeled variants for model training) as in claim 7 teaches for non-gapped benign protein sample is labelled with a benign ground truth sequence that has a benign label for a particular amino acid class. Regarding claim 8, Sundaram et al. teaches labeled benign size truth data set for training machine learning model (page 1166, col 2, para 2) as in claim 8 teaches benign ground truth sequence masked labels for respective remaining amino acid classes that correspond to amino acids different from the benign alternate amino acid. Regarding claim 9, Sundaram et al. teaches a sized truth dataset containing confidently labeled benign and pathogenic variants for training (page 1166, col 2, para 2) as in claim 9 teaches a non-gapped pathogenic protein sample is labelled with a pathogenic ground truth sequence for a particular amino acid class that corresponds to the pathogenic alternate amino acid. Regarding claim 10, Wu at al. teaches masking protein sequence to improve the accuracy of phylogenetic reconstructions (page 2-3, sections : Protein sequence simulation, The sensitivity and specificity of ZORRO masking) as in claim 10 teaches masking labels for remaining amino acid classes that are different from the pathogenic alternate amino acid. Regarding claim 11, Sundaram et al. discloses using pathogenicity prediction network to predict a three-state secondary structure at each amino acid position: alpha helix (H), beta sheet (B), and coils (C) (page S1, section: Architecture of the deep learning network; page 1164, Fig.3) as in claim 11 teaches pathogenicity predictor identifies, using a sample indicator, whether a current training example is a higher order structural representation for gapped/non-gapped protein samples. Regarding claim 12, Wu at al. teaches selecting three genes of different lengths in amino acids for masking (page 4, col 2, para 2) as claim 12 teaches masking the label amino acid class of a reference amino acid at the particular position in the particular gapped protein. Regarding claims 22-23, Sundaram et al. teaches systematic zeroed out the inputs at nearby amino acids (positions –25 to +25) around the variant and measured the change in the neural network’s predicted pathogenicity of the variant, the accuracy of prediction does not improve (Page S11, section: Recognition of protein motifs by the neural network; page S58, section: Interpreting the Deep Learning Models) as in claims 22-23 teaches a masked label does not contribute to training of the pathogenicity predictor, zeroed-out. Regarding independent claim 30, Sundaram et al. teaches training of a deep learning network for pathogenicity prediction using training data set of protein samples (page 1161, abstract; page 1166, fig.5; page 1167, col 1, para 2; page S56, section: ClinVar classification accuracy) as in claim 30 teaches training a pathogenicity predictor using a gapped training set that includes respective gapped protein. Regarding claim 30, Sundaram et al. teaches an adequately sized truth dataset containing confidently labelled benign and pathogenic variants for training (page 1166, section: A deep learning network for variant pathogenicity classification) as in claim 30 teaches a gapped training set of protein samples are labelled with respective gapped ground truth sequences . Regarding claim 30, Sundaram et al. teaches training of a deep learning network with benign labeled dataset (page 1170-1, section: Generation of benign and unlabeled variants for model training) as in claim 30 teaches a benign label for a particular amino acid class that corresponds to a reference amino acid at a particular position in the particular gapped protein. Regarding claim 30, Sundaram et al. teaches input of the human amino acid (AA) reference and alternate sequence data to the network (page 1164, fig.3) as claim 30 teaches a reference amino acid at a particular position in the particular gapped protein, and has respective pathogenic labels for respective remaining amino acid classes correspond to alternate amino acids. Regarding claim 30, Sundaram et al. teaches inputting the flanking amino acid sequence, generating secondary and crystal structure which are used for training deep learning network (page 1170-1: Architecture of the deep learning network ; page S31-S34: Supplementary Table 12-15) as claim 30 teaches generating spatial representations for protein samples. Regarding claim 30, Sundaram et al. teaches repeated training of deep learning network (classifier) (page 1170-2, col 2, para 3) and uses the trained deep neural network to determine pathogenic mutations (page 1161, abstract) as in claim 26 teaches repeated training including training a pathogenicity classifier (model) generating a trained pathogenicity classifier and further training the trained pathogenicity classifier generating a retrained pathogenicity classifier, which is used to determine pathogenicity prediction. Sundaram et al. does not teach: gapped protein samples (data) for training pathogenicity predictor, ground-truth labeled data, and masking protein sample data. Won et al. teaches use of variant deleted (= gapped) sequence data to train 3Cnet, a neural networks (page 4627, col 2, para 3; page 4630, col 2, para 3). Greiff et al. teaches feature space structure and ground truth–labeled data for machine learning training (page 1, abstract). Wu at al. teaches masking on protein sequence to improve the accuracy of phylogenetic reconstructions (page 2-3, sections : Protein sequence simulation, The sensitivity and specificity of ZORRO masking). It would have been obvious to a Person Having Ordinary Skill in the Art (PHOSITA) at the time of the claimed invention to combine the teachings of Sundaram et al. and Won et al. with the teachings of Greiff et al. and Wu et al. Specifically, Sundaram, Won and Greiff teach training a model with protein sequenced data, while Wu teaches masking on protein sequenced data. A PHOSITA would have been motivated to combine all these teachings because training a model with different data sets is a predictable variation in the field of bioinformatics where pathogenicity classifier/predictor can learn from missing-residue (gapped) contexts as well as ordinary variant examples. Conclusion No claims are allowed. Any inquiry concerning this communication or earlier communications from the examiner should be directed to PRAMOD KUMAR MOHANTA whose telephone number is (571)272-8775. The examiner can normally be reached Mon-Fri 9:00am-5:00pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Larry D Riggs can be reached at (571) 270-3062. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /P.K.M./ Examiner, Art Unit 1686 /LARRY D RIGGS II/Supervisory Patent Examiner, Art Unit 1686
Read full office action

Prosecution Timeline

Sep 26, 2022
Application Filed
Apr 29, 2026
Non-Final Rejection (signed) — §101, §103
Jul 17, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month