DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Status of claims
Pending:
1-20
Examined:
1-20
Independent:
1, 9, 15
Priority
As detailed on the 09/15/2023 filing receipt, this application claims domestic priority to as early as 09/06/2022.
Drawings
The drawings filed 09/06/2023 are accepted.
Information Disclosure Statement
No Information Disclosure Statement has been provided.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 7 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 7 recites “…about 15 percent…”. The term “about” in claim 7 is a relative term which renders the claim indefinite. The term “about” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Analysis of claims in Step 1.
Step 1: Are the claims directed to a 101 process, machine, manufacture, or composition of matter (MPEP 2106.03)?
Independent claim 1 is directed to a 101 machine or manufacture, here a "system," with non-transitory elements such as "a computing device comprising a processor and memory."
Independent claim 9 is directed to a 101 process, here a "method," with process steps such as "obtaining…, transforming…"
Independent claim 15 is directed to a 101 machine or manufacture, here a non-transitory "computer-readable medium."
[Step 1: claims 1-20: YES]
In accordance with MPEP § 2106, claims found to recite statutory subject matter (Step 1: YES) are then analyzed to determine if the claims recite any concepts that equate to an abstract idea, law of nature or natural phenomenon (Step 2A, Prong 1). In the instant application, the claims recite the following limitations that equate to an abstract idea:
Mental processes recited include:
Claims 1, 9 and 15 recite: "… predicting phosphorylation associations from the protein sequence…generating the context-aware protein sequence based upon the predicted phosphorylation associations…” The limitation is an act of evaluating, analyzing, observing and judging data that could be practically performed in the human mind and/or with pen and paper (See MPEP 2106.04(a)(2) subsection III).
Claim 5, 12 and 18 recites "wherein the Phosformer based transformer is trained using filtered protein sequencies." Filtering involves selecting protein sequences that could be practically performed in the human mind and/or with pen and paper.
Mathematical concepts recited include:
Claims 1, 9 and 15 recite: “transform the protein sequence to a context-aware protein sequence by a Phosformer based transformer, the transformation comprising: predicting phosphorylation associations from the protein sequence based upon a trained Phosformer model…”
Claim 5, 12 and 18 recites "…Phosformer based transformer…"
Claims 1, 5, 9, 12, 15 and 18, as indicated above, recite mental processes. Predicting phosphorylation associations is involved with determining the likelihood of assocations; generating protein sequences is involved with determining the protein sequence and filtering protein sequences is involved with selecting for certain protein sequences, which are all acts of evaluating, analyzing, observing and judging data. Acts of evaluating and analyzing data could be practically performed in the human mind and/or with pen and paper because they merely require making observations, evaluations, judgments, and opinions (See MPEP 2106.04(a)(2) subsection III). Overall, under the broadest reasonable interpretation, the indicated claims above can be practically carried out in the human mind or with pen and paper as claimed, which falls under the "mental processes" grouping of abstract ideas.
Claims 1, 5, 9, 12, 15 and 18, as indicated above, are mathematical concepts and/or mathematical formulas. Transforming the protein sequence to a context-aware protein sequence by a transformer and predicting phosphorylation associations require performing a series of mathematical operations. Also, the transformer model is a mathematical concept and formula. Therefore, under the broadest reasonable interpretation, the indicated claims above falls under the “mathematical concepts” grouping of abstract ideas.
As such, claims 1-20 recite an abstract idea (Step 2A, Prong 1: YES).
Claims found to recite a judicial exception under Step 2A, Prong 1 are then further analyzed to determine if the claims as a whole integrate the recited judicial exception into a practical application or not (Step 2A, Prong 2). The above indicated judicial exceptions are not integrated into a practical application because the claims do not recite an additional elements that apply, rely on or use the judicial exception in such a manner to amount to integration into a practical application. For example, there are no limitations that reflect an improvement to technology or applies or uses the recited judicial exception in some other meaningful way. Rather, the instant claims recite additional elements that equate to mere instructions to implement an abstract idea or insignificant extra solution activity. Specifically, the instant claims recite the following additional elements:
Claim 1 recites "a computing device comprising a processor and memory; and an application for phosphosite prediction comprising machine readable instructions stored in the memory... obtain a protein sequence…generating the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite; and render the predicted phosphosite for presentation to a user."
Claim 2, 10 and 16 recite "wherein the Phosformer model is pretrained based upon phosphorylation data"
Claim 4, 11 recites "wherein the kinase-substrate pairs are generated from a plurality of experimental databases."
Claim 5, 12 and 18 recites "wherein the Phosformer based transformer is trained using filtered protein sequencies."
Claim 6, 13 and 19 recites "wherein the filtered protein sequencies are generated based upon a random mask."
Claim 9 recites "obtaining, by at least one computing device, a protein sequence…generating the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite; and rendering the predicted phosphosite for presentation."
Claim 15 recites "A non-transitory computer readable medium having a program, that when executed by processing circuitry, causes the processing circuitry to: obtain a protein sequence… generate the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite."
Claims 3, 8, 11, 14, 17 and 20 are providing information on what the data represents and do not recite any additional elements.
As indicated above, claim 1 recites "a computing device comprising a processor and memory… machine readable instructions stored in the memory"; claim 9 recites "computing device" and claim 15 recites "non-transitory computer readable medium having a program," which equate to generic computer components, which are tools used to execute the abstract idea. The use of a computer or other machinery in its ordinary capacity for economic or other tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., a fundamental economic practice or mathematical equation) does not integrate a judicial exception into a practical application or provide significantly more. (see MPEP 2106.05(f)). The elements of claims 1-2, 4-6, 9-13, 15-16 and 18-19, as indicated above, equate to insignificant extra solutional activities of data gathering and outputting. Extra-solution activity includes both pre-solution and post-solution activity. An example of pre-solution activity is a step of gathering data for use in a claimed process, e.g., a step of obtaining information about credit card transactions, which is recited as part of a claimed process of analyzing and manipulating the gathered information by a series of steps in order to detect whether the transactions were fraudulent. An example of post-solution activity is an element that is not integrated into the claim as a whole, e.g., a printer that is used to output a report of fraudulent transactions, which is recited in a claim to a computer programmed to analyze and manipulate information about credit card transactions in order to detect whether the transactions were fraudulent (See MPEP 2106.05(g)). Limitations that add insignificant extra-solution activity to the judicial exception, as discussed in MPEP § 2106.05(g) have been identified by the courts to not integrate a judicial exception into a practical application (MPEP 2106.04(d)(I)). Additionally, the listed additional elements are mere instructions to apply an exception because they recite no more than an idea of a solution or outcome and does not recite a technological solution to a technological problem. (See MPEP 2106.05(f)(1)). As such, as currently recited, the claims do not appear to recite an improvement to technology or apply or use the recited judicial exception in some other meaningful way. Therefore, claims 1-20 are directed to an abstract idea (Step 2A, Prong 2: NO).
Claims found to be directed to a judicial exception are then further evaluated to determine if the claims recite an inventive concept that provides significantly more than the judicial exception itself (Step 2B). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims recite additional elements that equate to well-understood, routine and conventional activities, insignificant extra-solution activity or mere instructions to implement the abstract idea on a generic computer. The instant claims recite the following additional elements:
Claim 1 recites "a computing device comprising a processor and memory; and an application for phosphosite prediction comprising machine readable instructions stored in the memory... obtain a protein sequence…generating the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite; and render the predicted phosphosite for presentation to a user."
Claim 2, 10 and 16 recite "wherein the Phosformer model is pretrained based upon phosphorylation data"
Claim 4, 11 recites "wherein the kinase-substrate pairs are generated from a plurality of experimental databases."
Claim 5, 12 and 18 recites "wherein the Phosformer based transformer is trained using filtered protein sequencies."
Claim 6, 13 and 19 recites "wherein the filtered protein sequencies are generated based upon a random mask."
Claim 9 recites "obtaining, by at least one computing device, a protein sequence…generating the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite; and rendering the predicted phosphosite for presentation."
Claim 15 recites "A non-transitory computer readable medium having a program, that when executed by processing circuitry, causes the processing circuitry to: obtain a protein sequence… generate the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite."
Claims 3, 8, 11, 14, 17 and 20 are providing information on what the data represents and do not recite any additional elements.
The additional elements in claims 1-2, 4-6, 9-13, 15-16 and 18-19 do not comprise an inventive concept when considered individually or as an ordered combination that transforms the claimed judicial exception into a patent-eligible application of the judicial exception. The limitations equate to insignificant extra solutional activities. As explained by the Supreme Court, the addition of insignificant extra-solution activity does not amount to an inventive concept, particularly when the activity is well-understood or conventional. (see MPEP 2106.05(g)). Additionally, obtaining protein sequences and phosphorylation data for training prediction models are known methods as disclosed by Meng ("Mini-review: recent advances in post-translational modification site prediction based on deep learning." (2022): 3522-3532.; published 9 Jul 2022; as cited on the attached 892 form). Meng discloses available PTM databases that can be used to train a model (Page 3523, Section 2 PTM Databases and Table 1, page 3525). Limitations that add insignificant extra-solution activity to the judicial exception, e.g., mere data gathering in conjunction with a law of nature or abstract idea such as a step of obtaining information about credit card transactions so that the information can be analyzed by an abstract mental process, as discussed in CyberSource v. Retail Decisions, Inc., 654 F.3d 1366, 1375, 99 USPQ2d 1690, 1694 (Fed. Cir. 2011) (see MPEP § 2106.05(g)) have been identified by the courts to not to be enough to qualify as "significantly more" when recited in a claim with a judicial exception.
Furthermore, limitations that equate to mere data gathering and outputting via generic computer components, such as receiving data at a computer or outputting data, amount to insignificant extra-solution activity as set forth by the courts in Mayo, 566 U.S. at 79, 101 USPQ2d at 1968 and OIP Techs., Inc, v, Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1092-93 (Fed. Cir. 2015). Storing and retrieving information in memory were identified by the courts as well-understood, routine and conventional in Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93. Also, limitations that equate to mere data gathering and outputting via generic computer components, such as receiving data at a computer or outputting data via a graphic display device, amount to insignificant extra-solution activity as set forth by the courts in Mayo, 566 U.S. at 79, 101 USPQ2d at 1968 and OIP Techs., Inc. v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1092-93 (Fed. Cir. 2015). Overall, the use of a computer or other machinery in its ordinary capacity for economic or other tasks (e.g., to receive, store, or transmit data) or simply adding a general purpose computer or computer components after the fact to an abstract idea (e.g., a fundamental economic practice or mathematical equation) does not integrate a judicial exception into a practical application or provide significantly more as identified by the courts in Affinity Labs v. DirecTV, 838 F.3d 1253, 1262, 120 USPQ2d 1201, 1207 (Fed. Cir. 2016) (cellular telephone); TLI Communications LLC v. AV Auto, LLC, 823 F.3d 607, 613, 118 USPQ2d 1744, 1748 (Fed. Cir. 2016) (computer server and telephone unit).
Therefore, the claims do not amount to significantly more than the judicial exception itself (Step 2B: No). As such, claims 1-20 are not patent eligible.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Jiang ("A Pretrained ELECTRA model for kinase-specific phosphorylation site prediction." Computational Methods for Predicting Post-Translational Modification Sites. New York, NY: Springer US, 2022. 105-124.; published first online 14 June 2022; cited on the attached 892 form).
Regarding independent claim 1, Jiang teaches a computing device comprising a processor and memory; and an application for phosphosite prediction comprising machine readable instructions stored in the memory that, when executed by the processor, cause the computing device to at least: obtain a protein sequence with “…this work used the high-performance computing infra structure provided by Research Computing Support Services at the University of Missouri, as well as the Extreme Science and Engineering Discovery Environment (XSEDE)…” (page 122, last paragraph, Acknowledgements section) and “The input of the tool is a file containing protein sequences in the FASTA format, and the output of the tool is a file that includes the probability scores for each candidate site for each kinase. Residues with probability scores higher than a threshold (0.5 used as the default) are predicted as the substrates for the kinase under consideration. The ELECTRA model was trained on one GPU, and hence it needs one GPU to run the prediction. The software environment requirements are Python 3, TensorFlow-GPU 1.15. CUDA 10, NumPy, Pandas, scikit-learn, and SciPy. The default batch size is 128 for prediction, which can be adjusted according to the VRAM of the user’s local machine.” (page 120, para. 2).
Jiang teaches transform the protein sequence to a context-aware protein sequence by a Phosformer based transformer, the transformation comprising: predicting phosphorylation associations from the protein sequence based upon a trained Phosformer model with “A large-scale task-independent and unlabeled protein fragment dataset is essential for the pretrained model to unsupervised learn a comprehensive context-dependent biological properties representation of protein sequence segments.” (page 120, section 4 Notes, bullet 3); “The generator is trained with maximum likelihood to predict the original value of the masked-out tokens based on the context provided by the other, non masked amino acids in a fragment input.” (page 110, para. 2) and Figure 1 (page 109). Fig. 1 depicts the kinase-specific phosphorylation site prediction model. Fig. 1 caption discloses: The framework of the kinase-specific phosphorylation site prediction model. The left part represents pretraining. The ELECTRA discriminator is trained to model the representations for protein fragments by detecting the replaced tokens from the corrupt input generated by the generator on a sizeable unlabeled dataset. The generator is discarded after pretraining. The right part represents fine-tuning, where the ELECTRA model is trained on the labeled general phosphorylation site dataset, and the learned weights are transferred to the kinase-specific phosphorylation site prediction model for other kinase-specific training tasks. (Fig. 1 caption, page 109).
It is noted that Phosformer is interpreted as a transformer-based model for predicting phosphorylation that corresponds to the kinase-specific phosphorylation site prediction model in Fig. 1 as taught by Jiang.
Jiang teaches generating the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite; and render the predicted phosphosite for presentation to a user with “By visualizing the fragment representations generated by the pretrained ELECTRA and the untrained one-hot encoding method, the pretrained ELECTRA model demonstrated its ability to extract different latent biologically interpretable patterns under lying a specific kinase family to tell where it belongs to.” (page 121, para. 3 to page 122, para. 1).
Regarding claim 2, Jiang teaches wherein the Phosformer model is pretrained based upon phosphorylation data with Table 1 titled Phosphorylation data collected in this study (page 108) and “For pretraining the ELECTRA protein fragment representation model, the protein sequences were collected from the latest 2021 UniProt/Swiss-Prot [30]. We extracted a peptide of 33 residues centered at the potential phosphorylation site S or T together with 16 residues at each side from protein sequences, which is the same window size as MusiteDeep (see Note 1). Both general and kinase specific phosphorylation site protein sequences were obtained from the MusiteDeep paper. Table 1 summarizes the collected phosphorylation data used in this study. For general phosphorylation site prediction, phosphorylation sites on S or T of Homo sapiens annotated by UniProt/Swiss-Prot were used as positive data. In contrast, the sites with the same target amino acids but were not annotated as phosphorylation sites from the same proteins were regarded as the negative data.” (page 107, section 2.1 dataset, paragraph 1)
Regarding claim 3, Jiang teaches wherein the phosphorylation data comprises a plurality of kinase-substrate pairs with Table 1 titled Phosphorylation data collected in this study (page 108) and “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the samesubstrates were used as negative data. Since serine/threonine-specific kinase typi cally can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2).
Regarding claim 4, Jiang teaches wherein the kinase-substrate pairs are generated from a plurality of experimental databases with “For pretraining the ELECTRA protein fragment representation model, the protein sequences were collected from the latest 2021 UniProt/Swiss-Prot. We extracted a peptide of 33 residues centered at the potential phosphorylation site S or T together with 16 residues at each side from protein sequences, which is the same window size as MusiteDeep. Both general and kinase specific phosphorylation site protein sequences were obtained from the MusiteDeep paper. Table 1 summarizes the collected phosphorylation data used in this study. For general phosphorylation site prediction, phosphorylation sites on S or T of Homo sapiens annotated by UniProt/Swiss-Prot were used as positive data. In contrast, the sites with the same target amino acids but were not annotated as phosphorylation sites from the same proteins were regarded as the negative data.” (page 107, section 2.1 dataset, paragraph 1); “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2) and Table 1 titled Phosphorylation data collected in this study (page 108)
Regarding claim 5, Jiang teaches wherein the Phosformer based transformer is trained using filtered protein sequencies with “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2)
Regarding claim 6, Jiang teaches wherein the filtered protein sequencies are generated based upon a random mask with “The major advantage of MLM is that it uses a huge amount of unlabeled fragment data, which is shown in Table 1. After extracting the fragments, we have over 24 million fragment samples. The vocabulary size is 25, including 20 common amino acids, and using a dash token (“-”) as a pseudo amino acid for the padding positions at protein termini. Special tokens [MASK], [CLS], [SEP], and [UNK]are used, where [MASK]is for the masked-out token, [CLS] represents the start token, [SEP] represents the end token, and [UNK] is for unknown amino acids. In the embedding layer, each input fragment has 15% of its amino acids randomly replaced with special [MASK] tokens, and then is encoded into a token embeddings matrix with a size of [35, 128] and position embeddings matrix with a size of [35, 128] (see Note 5). The token embeddings and position embeddings are summed together as input tokens x ¼ [x1,..., xn](n ¼ 35) and fed into the generator.” (page 110, para. 1).
Regarding claim 7, Jiang teaches wherein about 15 percent of domain segments are randomly masked out with “In the embedding layer, each input fragment has 15% of its amino acids randomly replaced with special [MASK] tokens, and then is encoded into a token embeddings matrix with a size of [35, 128] and position embeddings matrix with a size of [35, 128] (see Note 5). The token embeddings and position embeddings are summed together as input tokens x ¼ [x1,..., xn](n ¼ 35) and fed into the generator.” (page 110, para. 1).
Regarding claim 8, Jiang teaches wherein the filtered protein sequencies comprise kinase- substrate sequences with “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2)
Regarding independent claim 9, Jiang teaches obtaining, by at least one computing device, a protein sequence with “…this work used the high-performance computing infra structure provided by Research Computing Support Services at the University of Missouri, as well as the Extreme Science and Engineering Discovery Environment (XSEDE)…” (page 122, last paragraph, Acknowledgements section) and “The input of the tool is a file containing protein sequences in the FASTA format, and the output of the tool is a file that includes the probability scores for each candidate site for each kinase. Residues with probability scores higher than a threshold (0.5 used as the default) are predicted as the substrates for the kinase under consideration. The ELECTRA model was trained on one GPU, and hence it needs one GPU to run the prediction. The software environment requirements are Python 3, TensorFlow-GPU 1.15. CUDA 10, NumPy, Pandas, scikit-learn, and SciPy. The default batch size is 128 for prediction, which can be adjusted according to the VRAM of the user’s local machine.” (page 120, para. 2).
Jiang teaches transforming, by the at least one computing device, the protein sequence to a context-aware protein sequence by a Phosformer based transformer, where the transformation comprises: predicting phosphorylation associations from the protein sequence based upon a trained Phosformer model with “A large-scale task-independent and unlabeled protein fragment dataset is essential for the pretrained model to unsupervised learn a comprehensive context-dependent biological properties representation of protein sequence segments.” (page 120, section 4 Notes, bullet 3); “The generator is trained with maximum likelihood to predict the original value of the masked-out tokens based on the context provided by the other, non masked amino acids in a fragment input.” (page 110, para. 2) and Figure 1 (page 109). Fig. 1 depicts the kinase-specific phosphorylation site prediction model. Fig. 1 caption discloses: The framework of the kinase-specific phosphorylation site prediction model. The left part represents pretraining. The ELECTRA discriminator is trained to model the representations for protein fragments by detecting the replaced tokens from the corrupt input generated by the generator on a sizeable unlabeled dataset. The generator is discarded after pretraining. The right part represents fine-tuning, where the ELECTRA model is trained on the labeled general phosphorylation site dataset, and the learned weights are transferred to the kinase-specific phosphorylation site prediction model for other kinase-specific training tasks. (Fig. 1 caption, page 109).
It is noted that Phosformer is interpreted as a transformer-based model for predicting phosphorylation that corresponds to the kinase-specific phosphorylation site prediction model in Fig. 1 as taught by Jiang.
Jiang teaches generating the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite; and rendering the predicted phosphosite for presentation with “By visualizing the fragment representations generated by the pretrained ELECTRA and the untrained one-hot encoding method, the pretrained ELECTRA model demonstrated its ability to extract different latent biologically interpretable patterns under lying a specific kinase family to tell where it belongs to.” (page 121, para. 3 to page 122, para. 1).
Regarding claim 10, Jiang teaches pretraining the Phosformer model based upon phosphorylation data with Table 1 titled Phosphorylation data collected in this study (page 108) and “For pretraining the ELECTRA protein fragment representation model, the protein sequences were collected from the latest 2021 UniProt/Swiss-Prot [30]. We extracted a peptide of 33 residues centered at the potential phosphorylation site S or T together with 16 residues at each side from protein sequences, which is the same window size as MusiteDeep (see Note 1). Both general and kinase specific phosphorylation site protein sequences were obtained from the MusiteDeep paper. Table 1 summarizes the collected phosphorylation data used in this study. For general phosphorylation site prediction, phosphorylation sites on S or T of Homo sapiens annotated by UniProt/Swiss-Prot were used as positive data. In contrast, the sites with the same target amino acids but were not annotated as phosphorylation sites from the same proteins were regarded as the negative data.” (page 107, section 2.1 dataset, paragraph 1)
Regarding claim 11, Jiang teaches wherein the phosphorylation data comprises a plurality of kinase-substrate pairs with Table 1 titled Phosphorylation data collected in this study (page 108) and “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2)
Regarding claim 12, Jiang teaches wherein the Phosformer based transformer is trained using filtered protein sequencies with “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2)
Regarding claim 13, Jiang teaches wherein the filtered protein sequencies are generated based upon a random mask with “The major advantage of MLM is that it uses a huge amount of unlabeled fragment data, which is shown in Table 1. After extracting the fragments, we have over 24 million fragment samples. The vocabulary size is 25, including 20 common amino acids, and using a dash token (“-”) as a pseudo amino acid for the padding positions at protein termini. Special tokens [MASK], [CLS], [SEP], and [UNK]are used, where [MASK]is for the masked-out token, [CLS] represents the start token, [SEP] represents the end token, and [UNK] is for unknown amino acids. In the embedding layer, each input fragment has 15% of its amino acids randomly replaced with special [MASK] tokens, and then is encoded into a token embeddings matrix with a size of [35, 128] and position embeddings matrix with a size of [35, 128] (see Note 5). The token embeddings and position embeddings are summed together as input tokens x ¼ [x1,..., xn](n ¼ 35) and fed into the generator.” (page 110, para. 1).
Regarding claim 14, Jiang teaches wherein the filtered protein sequencies comprise kinase- substrate sequences with “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2)
Regarding independent claim 15, Jiang teaches A non-transitory computer readable medium having a program, that when executed by processing circuitry, causes the processing circuitry to: obtain a protein sequence with “…this work used the high-performance computing infra structure provided by Research Computing Support Services at the University of Missouri, as well as the Extreme Science and Engineering Discovery Environment (XSEDE)…” (page 122, last paragraph, Acknowledgements section) and “The input of the tool is a file containing protein sequences in the FASTA format, and the output of the tool is a file that includes the probability scores for each candidate site for each kinase. Residues with probability scores higher than a threshold (0.5 used as the default) are predicted as the substrates for the kinase under consideration. The ELECTRA model was trained on one GPU, and hence it needs one GPU to run the prediction. The software environment requirements are Python 3, TensorFlow-GPU 1.15. CUDA 10, NumPy, Pandas, scikit-learn, and SciPy. The default batch size is 128 for prediction, which can be adjusted according to the VRAM of the user’s local machine.” (page 120, para. 2).
Jiang teaches transform the protein sequence to a context-aware protein sequence by a Phosformer based transformer, where the transformation comprises: predict phosphorylation associations from the protein sequence based upon a trained Phosformer model with “A large-scale task-independent and unlabeled protein fragment dataset is essential for the pretrained model to unsupervised learn a comprehensive context-dependent biological properties representation of protein sequence segments.” (page 120, section 4 Notes, bullet 3); “The generator is trained with maximum likelihood to predict the original value of the masked-out tokens based on the context provided by the other, non masked amino acids in a fragment input.” (page 110, para. 2) and Figure 1 (page 109). Fig. 1 depicts the kinase-specific phosphorylation site prediction model. Fig. 1 caption discloses: The framework of the kinase-specific phosphorylation site prediction model. The left part represents pretraining. The ELECTRA discriminator is trained to model the representations for protein fragments by detecting the replaced tokens from the corrupt input generated by the generator on a sizeable unlabeled dataset. The generator is discarded after pretraining. The right part represents fine-tuning, where the ELECTRA model is trained on the labeled general phosphorylation site dataset, and the learned weights are transferred to the kinase-specific phosphorylation site prediction model for other kinase-specific training tasks. (Fig. 1 caption, page 109).
It is noted that Phosformer is interpreted as a transformer-based model for predicting phosphorylation that corresponds to the kinase-specific phosphorylation site prediction model in Fig. 1 as taught by Jiang.
Jiang teaches generate the context-aware protein sequence based upon the predicted phosphorylation associations, the context-aware protein sequence comprising a predicted phosphosite with “By visualizing the fragment representations generated by the pretrained ELECTRA and the untrained one-hot encoding method, the pretrained ELECTRA model demonstrated its ability to extract different latent biologically interpretable patterns under lying a specific kinase family to tell where it belongs to.” (page 121, para. 3 to page 122, para. 1).
Regarding claim 16, Jiang teaches pretraining the Phosformer model based upon phosphorylation data with Table 1 titled Phosphorylation data collected in this study (page 108) and “For pretraining the ELECTRA protein fragment representation model, the protein sequences were collected from the latest 2021 UniProt/Swiss-Prot [30]. We extracted a peptide of 33 residues centered at the potential phosphorylation site S or T together with 16 residues at each side from protein sequences, which is the same window size as MusiteDeep (see Note 1). Both general and kinase specific phosphorylation site protein sequences were obtained from the MusiteDeep paper. Table 1 summarizes the collected phosphorylation data used in this study. For general phosphorylation site prediction, phosphorylation sites on S or T of Homo sapiens annotated by UniProt/Swiss-Prot were used as positive data. In contrast, the sites with the same target amino acids but were not annotated as phosphorylation sites from the same proteins were regarded as the negative data.” (page 107, section 2.1 dataset, paragraph 1).
Regarding claim 17, Jiang teaches wherein the phosphorylation data comprises a plurality of kinase-substrate pairs with Table 1 titled Phosphorylation data collected in this study (page 108) and “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2)
Regarding claim 18, Jiang teaches wherein the Phosformer based transformer is trained using filtered protein sequencies with “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2).
Regarding claim 19, Jiang teaches wherein the filtered protein sequencies are generated based upon a random mask with “The major advantage of MLM is that it uses a huge amount of unlabeled fragment data, which is shown in Table 1. After extracting the fragments, we have over 24 million fragment samples. The vocabulary size is 25, including 20 common amino acids, and using a dash token (“-”) as a pseudo amino acid for the padding positions at protein termini. Special tokens [MASK], [CLS], [SEP], and [UNK]are used, where [MASK]is for the masked-out token, [CLS] represents the start token, [SEP] represents the end token, and [UNK] is for unknown amino acids. In the embedding layer, each input fragment has 15% of its amino acids randomly replaced with special [MASK] tokens, and then is encoded into a token embeddings matrix with a size of [35, 128] and position embeddings matrix with a size of [35, 128] (see Note 5). The token embeddings and position embeddings are summed together as input tokens x ¼ [x1,..., xn](n ¼ 35) and fed into the generator.” (page 110, para. 1).
Regarding claim 20, Jiang teaches wherein the filtered protein sequencies comprise kinase- substrate sequences with “For kinase-specific phosphorylation site prediction, we only consider kinase families CDK, CK2, MAPK, PKA, and PKC, each of which has more than 100 known substrate phosphorylation sites in our collected dataset. For each kinase family, only sites annotated by the specific kinase family were used as positive data, whereas all other residues of the same types (S or T) in the same substrates were used as negative data. Since serine/threonine-specific kinase typically can phosphorylate both S and T residues, we combined phosphoserine and phosphothreonine sites in the data collection…” (page 107, section 2.1 dataset, paragraph 2).
Conclusion
No claims are allowed.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KETTIP KRIANGCHAIVECH whose telephone number is (571)272-1735. The examiner can normally be reached 8:30am-5:00pm EDT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Larry D. Riggs can be reached at (571) 270-3062. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/K.K./Examiner, Art Unit 1686
/Karlheinz R. Skowronek/Supervisory Patent Examiner, Art Unit 1687