DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d) regarding application JP2020-202081 filed December 4th, 2020. The certified copy has been filed in this application. Acknowledgement is made of 371 of PCT/JP2021/040948 filed November 8th, 2021. The effective filing date is December 4th, 2020.
Status of Claims
Claims 1-20 are currently pending and examined on the merits.
Information Disclosure Statement
The Information Disclosure Statement(s) is/are acknowledged and the references contained therein have been considered by the Examiner. This includes the Information Disclosure Statements(s) filed on: May 30th, 2023.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
“acquisition unit that acquires”, in claim 1
“inversion unit that generates” , in claim 1
“generation unit that generates” , in claims 1, 3, and 11
“integration unit that integrates”, in claim 3
“prediction unit that predicts”, in claims 3, 7, 8, 10, 12-16
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 3, 7, 8, and 11-16 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 states “…an acquisition unit that acquires sequence information relating to a genome sequence”. One of ordinary skill in the art of bioinformatics and genomics would not understand the mechanism by which the acquisition unit acquires sequence information, and neither the specification nor the claims state such a mechanism. To further prosecution, the Examiner interprets the acquisition unit as any computational method of acquiring genomic sequences.
Claim 1 further states “an inversion unit that generates, on a basis of the sequence information, inversion information in which the sequence is inverted”. One of ordinary skill in the art of bioinformatics and genomics would not understand the mechanism by which the inversion unit generates inversion information, and neither the specification nor the claims state such a mechanism. To further prosecution, the Examiner interprets the inversion unit as any computational method of generating an inverted sequence or reverse compliment.
Claims 1, 3, and 11 state “…a generation unit that generates, on a basis of the inversion information, protein information relating to a protein.” One of ordinary skill in the art of bioinformatics and genomics would not understand the mechanism by which the generation unit generates protein information, and neither the specification nor the claims state such a mechanism. To further prosecution, the Examiner interprets the mechanism the generation unit uses as any computational method of generating protein information, such as a search.
Claims 3, 7, 8, 10, and 12-16 state “…prediction unit that predicts first protein information…”. Claim 7 states that “…the first prediction unit executes machine learning…”; however, one of ordinary skill in the art of bioinformatics and genomics would not understand the mechanism by which the prediction unit executes machine learning, nor which machine learning method. Neither the specification nor the claims explicitly state such a mechanism. To further prosecution, the Examiner interprets the prediction mechanism associated with the prediction unit as any machine learning method applied to predict protein information.
Claim 3, 6, and 8 state “…integration unit that integrates…” and “…integration unit executes machine learning…”. The applicant’s specification states embodiments of a function for the integration unit, stating “The integration unit may execute machine learning using the first protein information and the second protein information as inputs to predict the protein information…” on paragraphs 0015-0017. However, one of ordinary skill in the art of bioinformatics and genomics would not understand the specific mechanism by which the integration unit integrates protein information or executes machine learning. The specification and claims lack such a definition. To further prosecution, the Examiner interprets integration unit as any machine learning method applied to combine two or more pieces of protein information for use in predicting protein structure and/or function.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of mental steps, mathematic concepts, organizing human activity, or a natural law without significantly more.
Regarding:
“acquisition unit that acquires”, in claim 1, the Examiner interprets the acquisition unit as any computational method of acquiring genomic sequences.
“inversion unit that generates” , in claim 1, the Examiner interprets the inversion unit as any computational method of generating an inverted sequence or reverse compliment.
“generation unit that generates” , in claims 1, 3, and 11, the Examiner interprets the mechanism the generation unit uses as any computational method of generating protein information, such as a search.
“integration unit that integrates”, and “…integration unit executes machine learning…” in claims 3, 6, and 8 , the Examiner interprets integration unit as any machine learning method applied to combine two or more pieces of protein information for use in predicting protein structure and/or function.
“prediction unit that predicts”, in claims 3, 7, 8, 10, 12-16, the Examiner interprets the prediction mechanism associated with the prediction unit as any machine learning method applied to predict protein information.
Step 2A, Prong 1
In accordance with MPEP § 2106, claims found to recite statutory subject matter (claim 1-18 are drawn to an apparatus, claim s 19-20 are drawn to a method) (Step 1: YES) are then analyzed to determine if the claims recite any concepts that equate to an abstract idea, law of nature or natural phenomenon (Step 2A, Prong 1). In the instant application, the claims (listed numerically) recite the following limitations that equate to an abstract idea (reasonings in parentheses):
Claim 1 states:
…an inversion unit that generates, on a basis of the sequence information, inversion information in which the sequence is inverted; and (which is a mental step, i.e. can be performed with pen and paper)
…a generation unit that generates, on a basis of the inversion information, protein information relating to a protein. (mental step)
Claim 3 states:
the generation unit includes a first prediction unit that predicts first protein information on a basis of the sequence information, (mental step)
a second prediction unit that predicts second protein information on a basis of the inversion information, and (mental step)
an integration unit that integrates the first protein information and the second protein information to generate the protein information (mental step)
Claim 4 states:
…wherein the protein information includes at least one of a structure of the protein or a function of the protein (mental step)
Claim 5 states:
…the protein information includes at least one of a contact map indicating a bond between amino acid residues forming the protein, (which is a mathematical concept of a mathematical calculation)
…a distance map indicating a distance between amino acid residues forming the protein, (mathematical calculation)
or a tertiary structure of the protein. (mathematical calculation)
Claim 6 states: …the integration unit executes machine learning using the first protein information and the second protein information as inputs to predict the protein information (mental step)
Claim 7 states:
the first prediction unit executes machine learning using the sequence information as an input to predict the first protein information, and (mental step)
the second prediction unit executes machine learning using the inversion information as an input to predict the second protein information. (mental step)
Claim 11 states:
a feature amount calculation unit that calculates a feature amount on a basis of the sequence information, (mathematical calculation)
wherein the generation unit generates the protein information on a basis of the feature amount. (mental step)
Claim 12 states:
the feature amount calculation unit calculates a first feature amount on a basis of the sequence information, (mathematical calculation)
the first prediction unit predicts the first protein information on a basis of the sequence information and the first feature amount, and (mental step)
the second prediction unit predicts the second protein information on a basis of the inversion information and the first feature amount (mental step)
Claim 13 states:
the feature amount calculation unit calculates a first feature amount on a basis of the sequence information and calculates a second feature amount on a basis of the inversion information, (mathematical calculation)
the first prediction unit predicts the first protein information on a basis of the sequence information and the first feature amount, and (mental step)
the second prediction unit predicts the second protein information on a basis of the inversion information and the second feature amount. (mental step)
Claim 17 states:
…the feature amount includes at least one of a secondary structure of the protein, annotation information relating to the protein, the degree of catalyst contact of the protein, or a mutual potential between amino acid residues forming the protein. (mathematical calculation)
Claim 18 states:
…the inversion information is information indicating a bonding order from a C-terminal side of amino acid residues forming the protein. (mental step)
Claims 19 and 20 state:
generating, on a basis of the sequence information, inversion information in which the sequence is inverted; and (mental step)
predicting, on a basis of the inversion information, first protein information relating to a protein. (mental step)
The claims recite an abstract idea of analyzing genomic sequencing data (See MPEP 2106.07(a)).
These recitations are similar to the concepts of collecting information, analyzing it and displaying certain results of the collection and analysis in Electric Power Group, LLC, v. Alstom (830 F.3d 1350, 119 USPQ2d 1739 (Fed. Cir. 2016)), organizing and manipulating information through mathematical correlations in Digitech Image Techs., LLC v Electronics for Imaging, Inc. (758 F.3d 1344, 111 U.S.P.Q.2d 1717 (Fed. Cir. 2014)) and comparing information regarding a sample or test to a control or target data in Univ. of Utah Research Found. v. Ambry Genetics Corp. (774 F.3d 755, 113 U.S.P.Q.2d 1241 (Fed. Cir. 2014)) and Association for Molecular Pathology v. USPTO (689 F.3d 1303, 103 U.S.P.Q.2d 1681 (Fed. Cir. 2012)) that the courts have identified as concepts that can be practically performed in the human mind or mathematical relationships. Therefore, these limitations fall under the “Mental process” and “Mathematical concepts” groupings of abstract ideas.
There are no additional limitations that indicate that the claims require anything other than carrying out the recited mental process or mathematical concept in a generic computer environment. Merely reciting that a mental process is being performed in a generic computer environment does not preclude the steps from being performed practically in the human mind or with pen and paper as claimed. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then if falls within the “Mental processes” grouping of abstract ideas. As such, claim(s) 1-16 recite(s) an abstract idea/law of nature/natural phenomenon (Step 2A, Prong 1: YES).
Step 2A, Prong 2
Claims found to recite a judicial exception under Step 2A, Prong 1 are then further analyzed to determine if the claims as a whole integrate the recited judicial exception into a practical application or not (Step 2A, Prong 2). This judicial exception is not integrated into a practical application because the claims do not recite an additional element that reflects an improvement to technology or applies or uses the recited judicial exception to affect a particular treatment for a condition. Rather, the instant claims recite additional elements that amount to mere instructions to implement the abstract idea in a generic computing environment or mere instructions to apply the recited judicial exception via a generic treatment. Specifically, the claims recite the following additional elements:
[Claim 1] … an acquisition unit that acquires sequence information relating to a genome sequence
[Claim 2] …the sequence information is information relating to at least one of a sequence of amino acids, a sequence of DNA, or a sequence of RNA.
[Claim 6] …the integration unit executes machine learning using the first protein information and the second protein information as inputs to predict the protein information…
[Claim 8]
…the integration unit includes a machine learning model for integration trained on a basis of an error between the protein information predicted using the first protein information for learning predicted using the sequence information for learning associated with correct answer data as a input and
…the second protein information for learning predicted using the inversion information generated on a basis of the sequence information for learning as an input as inputs and the correct answer data.
[Claim 9]
… the first prediction unit includes a first machine learning model trained on a basis of an error between the first protein information for learning and the correct answer data, and
…the first machine learning model is re-trained on a basis of an error between the protein information predicted using the first protein information for learning and the second protein information for learning as inputs and the correct answer data.
[Claim 10]
…the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information for learning and the correct answer data, and
…the second machine learning model is re-trained on a basis of an error between the protein information predicted using the first protein information for learning and the second protein information for learning as inputs and the correct answer data.
[Claim 14]
…The information processing apparatus according to claim 12, wherein the first prediction unit includes a first machine learning model trained on a basis of an error between the first protein information predicted using the sequence information for learning, which is associated with correct answer data, and
…the first feature amount for learning, which is calculated on a basis of the sequence information for learning, as inputs and the correct answer data.
[Claim 15]
…the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information predicted using the inversion information generated on a basis of the sequence information for learning and the first feature amount for learning, which is calculated on a basis of the sequence information for learning, as inputs and the correct answer data.
[Claim 16]
…the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information predicted using the inversion information, which is generated on a basis of the sequence information for learning, and
…the second feature amount for learning calculated on a basis of the inversion information as inputs and the correct answer data.
[Claim 18]
… the sequence information is information indicating a bonding order from an N-terminal side of amino acid residues forming the protein, and
[Claim 18] …acquiring sequence information relating to a genome sequence;
There are no limitations that indicate that the claimed analysis engine or the formats of the provided data require anything other than generic computing systems. As such, these limitations equate to mere instructions to implement the abstract idea on a generic computer that the courts have stated does not render an abstract idea eligible in Alice Corp., 573 U.S. at 223, 110 USPQ2d at 1983. As such, claims 1-20 are directed to an abstract idea (Step 2A, Prong 2: NO).
Step 2B
Claims found to be directed to a judicial exception are then further evaluated to determine if the claims recite an inventive concept that provides significantly more than the judicial exception itself (Step 2B). The claims do not include additional elements that are sufficient to amount to significantly
more than the judicial exception because the claims recite additional elements that equate to mere
instructions to apply the recited exception in a generic way or in a generic computing environment. The
instant claims recite the following additional elements:
[Claim 2] …the sequence information is information relating to at least one of a sequence of amino acids, a sequence of DNA, or a sequence of RNA.
[Claim 4] …wherein the protein information includes at least one of a structure of the protein or a function of the protein.
[Claim 8]
…the integration unit includes a machine learning model for integration trained on a basis of an error between the protein information predicted using the first protein information for learning predicted using the sequence information for learning associated with correct answer data as a input and
…the second protein information for learning predicted using the inversion information generated on a basis of the sequence information for learning as an input as inputs and the correct answer data.
[Claim 10]
…the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information for learning and the correct answer data, and
…the second machine learning model is re-trained on a basis of an error between the protein information predicted using the first protein information for learning and the second protein information for learning as inputs and the correct answer data.
[Claim 14]
…The information processing apparatus according to claim 12, wherein the first prediction unit includes a first machine learning model trained on a basis of an error between the first protein information predicted using the sequence information for learning, which is associated with correct answer data, and
…the first feature amount for learning, which is calculated on a basis of the sequence information for learning, as inputs and the correct answer data.
[Claim 15]
…the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information predicted using the inversion information generated on a basis of the sequence information for learning and the first feature amount for learning, which is calculated on a basis of the sequence information for learning, as inputs and the correct answer data.
[Claim 16]
…the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information predicted using the inversion information, which is generated on a basis of the sequence information for learning, and
…the second feature amount for learning calculated on a basis of the inversion information as inputs and the correct answer data.
[Claim 18]
… the sequence information is information indicating a bonding order from an N-terminal side of amino acid residues forming the protein...
Regarding claims 3-5, 10, 13-15 and 20, the steps of analyzing sequencing data does not integrate the abstract idea into a practical application and constitutes an insignificant extra-solution activity (i.e., data gathering and presentation), which does not impose a meaningful limit on the abstract idea (see MPEP 2106.05 (g)).
As discussed above, there are no additional limitations to indicate that the claimed
analysis requires anything other than generic computer components in order to carry out the recited abstract idea in the claims. Claims that amount to nothing more than an instruction to apply the abstract idea using a generic computer do not render an abstract idea eligible. Alice Corp., 573 U.S. at 223, 110 USPQ2d at 1983. See also 573 U.S. at 224, 110 USPQ2d at 1984. MPEP 2106.05(f) discloses that mere instructions to apply the judicial exception cannot provide an inventive concept to the claims.
Additionally, the claims are directed to well-understood, routine, and conventional activity as evidenced by Wu et al. (Bioinformatics. 2019 Jun 7;36(1):41–48), who teaches protein contact prediction using metagenome sequence data residual neural networks, and Torrisi et al. (Computational and Structural Biotechnology Journal 18 (2020) 1301–1310.) who teaches deep learning methods in protein structure prediction (specifically listing bidirectional RNN architecture capable of locating protein sequences by terminus on pg. 1304)
The additional elements do not comprise an inventive concept when considered individually or
as an ordered combination that transforms the claimed judicial exception into a patent-eligible application of the judicial exception. Therefore, the claims do not amount to significantly more than the
judicial exception itself (Step 2B: No). As such, claims 1-20 is/are not patent eligible.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-8, 11, 13, 16, and 18-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Xueliang Leon Liu et al. (arXiv:1701.08318v1 [q-bio.QM], Jan 2017, pg. 1-38).
Regarding:
“acquisition unit that acquires”, in claim 1, the Examiner interprets the acquisition unit as any computational method of acquiring genomic sequences.
“inversion unit that generates” , in claim 1, the Examiner interprets the inversion unit as any computational method of generating an inverted sequence or reverse compliment.
“generation unit that generates” , in claims 1, 3, and 11, the Examiner interprets the mechanism the generation unit uses as any computational method of generating protein information, such as a search.
Regarding claims 1, 19, and 20, Liu et al teaches a machine learning method used to predict protein function from its sequence (Title, Abstract, pg. 1-2, Results, pg. 7, Methods [Computational Modeling], pg. 25-26).
Liu et al. further teaches protein amino acid acquisition in the “In-class Model Training and Validation” section on pg. 7-8 (“For training, protein amino acid sequences were obtained from the UniProt database and directly used as inputs into the neural network without any feature extraction.”, re: clm. 1, 19, 20 … An information processing apparatus [method], comprising: an acquisition unit that acquires sequence information relating to a genome sequence…)
Liu et al. further discloses in the model architecture on pg. 5-7 that “The forward layer scans the protein sequences from the N towards the C terminus and reversed for the backward layer, allowing the network to make use of context on both sides of each position rather than just what was seen before in a single direction…” and additionally discloses this detail in figure 1b, stating on pg. 32: “The RNN model consists of arbitrary sets of forward and reverse layers of long short term memory (LSTM) neurons taking only the amino acid letters from the sequence as input (red).” (re: clm. 1, … an inversion unit that generates, on a basis of the sequence information, inversion information in which the sequence is inverted…)
Liu et al. further teaches prediction and generation, using inverted sequences, of protein information in Table 1, pg. 38, and figures 3 and 4 (re: clm. 1, … a generation unit that generates, on a basis of the inversion information, protein information relating to a protein.) Liu et al. teaches a LTSM-based machine learning method that predicts proteins in anticipation of claim 1.
Regarding claim 2, Liu et al. teaches processing information related to amino acids (Methods, Results, pgs. 1-5, re: clm. 2, … [Claim 2] The information processing apparatus according to claim 1, wherein the sequence information is information relating to at least one of a sequence of amino acids, a sequence of DNA, or a sequence of RNA.). Liu et al. anticipates claim 2.
Regarding claims 3 and 4, Liu et al. teaches prediction of protein information from sequence information using inversion information on pg. 16: “Lastly “out-of-class” prediction performance was tested, whereby the RNN models were trained on sequences from certain protein families and tested on other functionally homologous but phylogenetically distinct families.”. Results in Table 1 on pg. 38 details the classification performance, including an average between the N terminal and C terminal ordered processing, which reads on integrating two individual sources of protein information to generate a separate series of protein information (prediction) (re: clm. 3, … wherein the generation unit includes a first prediction unit that predicts first protein information on a basis of the sequence information, a second prediction unit that predicts second protein information on a basis of the inversion information, and an integration unit that integrates the first protein information and the second protein information to generate the protein information.) Liu et al. further teaches that the RNN models were trained using the Genome-Edit (CRISPR) functions and details predictions and annotations of such functions in Table 1 (pg. 38) and Fig. 4 (pg. 35) (re: clm. 4, … wherein the protein information includes at least one of a structure of the protein or a function of the protein.). Liu et al. anticipates claims 3 and 4.
Regarding claim 6, To further prosecution, the Examiner interprets integration unit as any machine learning method applied combine two or more pieces of protein information for use in predicting protein structure and/or function.
Regarding claim 6, Liu et al. teaches a machine learning method applied to protein information in which amino acid sequences (the first piece of protein information) are combined with additional features as listed on pg. 5 (“The recurrent neural network (RNN) model contains one or more sets of “bidirectional” recurrent layers with long short term memory (LSTM) neurons processing the input sequence one residue or character at a time (Figure 1b).” and on pg. 15 (“These features include simple amino acid composition and length as well as biochemically relevant properties such as isoelectric point, molecular weight, stability index, hydrophobicity and grand average of hydropathicity (gravy)) to predict protein function (re: clm. 6, … the integration unit executes machine learning using the first protein information and the second protein information as inputs to predict the protein information.) Liu et al. anticipates claim 6.
Regarding claim 7, Liu et al. teaches the performance of the classification, including an average between the N terminal and C terminal ordered processing, and states “The average of the predictions on up to 800 amino acids in the N and C termini significantly increased precision and F1, suggesting several important features throughout the entire sequence that are necessary for function.” As previously disclosed, Liu et al. states that “The forward layer scans the protein sequences from the N towards the C terminus and reversed for the backward layer, allowing the network to make use of context on both sides of each position rather than just what was seen before in a single direction…” on pg. 5-7, assigning N as forward and C as reverse. Therefore, Liu et al. executes machine learning using N terminus sequence information and C terminus inversion information to predict protein information (re: clm. 7, … the first prediction unit executes machine learning using the sequence information as an input to predict the first protein information, and the second prediction unit executes machine learning using the inversion information as an input to predict the second protein information.). Liu et al. anticipates claim 7.
Regarding claim 8, Liu et al. discloses the LTSM architecture accounts for error on pages 5-7, stating:
“These features of the LSTM architecture allow the RNN to maintain, over many recurrent
iterations, the magnitudes of both the relevant signals in feed forward propagation as well as the error gradients in backpropagation, thereby resolving the issues of loss of contextual memory and vanishing/exploding gradients that have limited the usefulness of traditional RNNs in processing long sequences (e.g. hundreds of units/iterations).” Liu et al. further discusses validation metrics on pg. 10 through ROC plots for sensitivity as plotted in Fig. 2b and on pg. 26 as formulas to calculate accuracy (re: clm. 8, … wherein the integration unit includes a machine learning model for integration trained on a basis of an error between the protein information predicted using the first protein information for learning predicted using the sequence information for learning associated with correct answer data as a input and the second protein information for learning predicted using the inversion information generated on a basis of the sequence information for learning as an input as inputs and the correct answer data.). Liu teaches machine learning models trained on a basis of an error of predicted vs. actual in anticipation of claim 8.
Regarding claim 11, Liu et al. teaches predicted functions from protein sequences without an assigned function within a database (pg. 11, Fig. 3). Figure 3 (pg. 34), depicts the number (amount) of unique sequences identified as a result results of training the recurrent neural network (RNN) model using sequences from four protein sources ( pg. 11, “Separately, the trained RNN models also predicted thousands of new hits from the UniRef100 database for each function…, re: clm. 11, … a feature amount calculation unit that calculates a feature amount on a basis of the sequence information, wherein the generation unit generates the protein information on a basis of the feature amount.). Liu et al. teaches the output unique sequence annotations as calculated from a trained RNN model in anticipation of claim 11.
Regarding claim 13, Liu et al., in Table 1, details the Out-of-class RNN classification performance , whereby the RNN models were trained on sequences from certain protein families and tested on other
functionally homologous but phylogenetically distinct families to predict protein function. Liu et al. discloses in the legend on pg. 38 that “The average of the predictions on up to 800 amino acids in the N and C termini significantly increased precision and F1, suggesting several important features throughout the entire sequence that are necessary for function.” Liu et al. further discloses on pg. 17 that “As the “GenomeEdit” Cas9 or Cpf1 enzyme sequences are typically over 1000 amino acids long, the RNN was trained scanning over up to the first 800 amino acids from the N-terminus and subsequently from the C-terminus.” Therefore, Liu et al. calculates (predicts) a number of protein functions (pg. 16, “GenomeEdit”,“Ferritin”, and “P450”. ) given protein sequence information (N-terminal side of amino acid) and inversion information (C-terminal side amino acid) and protein function information
Regarding claim 16, Liu et al. teaches the inclusion of additional machine learning models to test against the RNN model as stated on pg. 14 (“For machine learning benchmark, the performance of the RNN model was compared against other popular machine learning classification models, particularly logistic regression and random forest which are known for speed…”
The applicant’s specification states: “Note that in order to execute each of the prediction of the first contact map 21 by the first prediction unit 18 and the prediction of the second contact map 22 by the second prediction unit 19, the same algorithm may be used or different algorithms may be used…”, indicating that the same machine learning model may be applied for different use cases.
Liu et al. discloses in Fig. 5 that these models were trained on the same input (protein sequence data inclusive of N and C-terminus order protein information), and further discloses accuracy, precision and recall, reading on training said models on a basis of error (re: clm. 16, … the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information predicted using the inversion information, which is generated on a basis of the sequence information for learning, and the second feature amount for learning calculated on a basis of the inversion information as inputs and the correct answer data.). Liu et al. teaches additional machine learning algorithm use in anticipation of claim 16.
Regarding claim 18 Liu et al. states on pg. 5-7 in their Model Architecture section that “The recurrent neural network (RNN) model contains one or more sets of “bidirectional” recurrent layers with long-short term memory (LSTM) neurons processing the input sequence one residue or character at a time (Figure 1b). The forward layer scans the protein sequences from the N- towards the C-terminus and reversed for the backward layer, allowing the network to make use of context on both sides of each position rather than just what was seen before in a single direction.”. This reads on sequence information indicating a bonding order from an N-terminal side, and inversion information indicating a bonding order from a C-terminal side (regarding amino acid residues; re: clm. 18, … The information processing apparatus according to claim 2, wherein the sequence information is information indicating a bonding order from an N-terminal side of amino acid residues forming the protein, and the inversion information is information indicating a bonding order from a C-terminal side of amino acid residues forming the protein.). Liu et al. anticipates claim 18.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 12, 14, 15 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. as applied to claims 1-8, 11, 13, 16, and 18-20 above.
Liu et al. is applied to claims 1-8, 11, 13, 16, and 18-20 above.
Regarding claim 12, Liu et al. teaches calculating protein function from protein sequences and database annotations, but utilizes one-hot encoding of input protein sequences from N and C terminus residues (as stated on pg. 5), and does not share an initial (first) protein feature amount such that it may form the prediction for another feature of protein information (Fig. 1a and 1b, re: clm. 12, … the feature amount calculation unit calculates a first feature amount on a basis of the sequence information, the first prediction unit predicts the first protein information on a basis of the sequence information and the first feature amount, and the second prediction unit predicts the second protein information on a basis of the inversion information and the first feature amount.)
Applying the KSR standard of obviousness to Liu et al., the Examiner concluded that the combination of sharing protein sequences from both orders (N and C-terminus) as input into a prediction pipeline represents an illustration of the reasoning that it would have been “obvious to try" choosing from a finite number of identified, predictable solutions to achieve the claimed process. The rationale to support a conclusion that the claim would have been obvious is that "a person of ordinary skill has good reason to pursue the known options within his or her technical grasp (see MPEP 2143).
Liu et al. state in their Table 1 legend and pg. 16-18 that the average of predictions utilizing both N and C terminal residues increased prediction, and suggests “that multiple features along the entire sequence length (e.g. the binding and nuclease domains) may be required toward accomplishing the “Genome-Edit” function and that many other proteins may exist with only a subset of those features…” (pg. 18). This suggestion implies that longer length could improve RNN performance; a suggestion re-asserted by Liu et al. in their discussion, stating:
“Long sequences have been particularly challenging for RNN training due to the exploding or vanishing gradient issue with back-propagation… Given significantly more computational resources and time, results here have shown that deeper RNN models could be trained on the currently available dataset to make reasonable predictions (Table 1).” (pg. 22)
When there is motivation to solve a problem and there are a finite number of identified, predictable solutions, a person of ordinary skill has good reason to pursue the known options within his or her technical grasp. If this leads to anticipated success, it is likely the product not of innovation but of ordinary skill and common sense." The skilled artisan would have had reason to try these methods with the reasonable expectation that at least one would be successful. Thus, utilizing both N and C termini protein information as input claimed is merely a "predictable use of prior art elements according to their established functions." KSR Int’l 7, 127 S. Ct. at 1740. (paraphrased from Example 3 of MPEP 2143 E.). Therefore, claim 12 would have been prima facie obvious absent evidence to the contrary.
Regarding claims 14 and 15, Liu et al. teaches a machine learning model trained on a basis of error (regarding protein sequences and related functions to said protein sequences as input to a RNN model) as Liu et al. states in the “Model Architecture” section of the results that “The training process [for the RNN] tunes the parameters of the network by minimizing prediction errors (categorical entropy).” (re: clm. 14, … wherein the first prediction unit includes a first machine learning model trained on a basis of an error between the first protein information predicted using the sequence information for learning, which is associated with correct answer data, and the first feature amount for learning, which is calculated on a basis of the sequence information for learning, as inputs and the correct answer data.)
Liu et al. applies the same error minimization to the inversion information as it does the sequence information (i.e., C- and N-termini are processed in the same manner, re: clm. 15, … the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information predicted using the inversion information generated on a basis of the sequence information for learning and the first feature amount for learning, which is calculated on a basis of the sequence information for learning, as inputs and the correct answer data. ).
Again, the applicant’s specification states “…in order to execute each of the prediction of the first contact map… the same algorithm may be used or different algorithms may be used…” allowing for reuse of the RNN model for the first and second prediction units. Therefore, Liu et al. teaches the limitations of claims 14 and 15.
Regarding claim 17, Liu et al. teaches annotation information relating to protein as stated in the methods (“Using the same dataset for each of the four functional classes, 51 ProtParam features (Table S1) were extracted or calculated for each sequence and vectorized. These features include simple amino acid composition and length as well as biochemically relevant properties such as isoelectric point, molecular weight…”), which the applicant’s specification also states may be information, stating:
“As the information relating to a structure, for example, the name of the functional group of the protein is given. In addition, a molecular weight or the like of the protein may be given as the annotation information.” (re: clm. 17, … wherein the feature amount includes at least one of a secondary structure of the protein, annotation information relating to the protein, the degree of catalyst contact of the protein, or a mutual potential between amino acid residues forming the protein.) Liu et al. teaches protein annotation in address of the limitations of claim 17.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. as applied to claims 1-8, 11-17, and 18-20 above and in view of Hanson et al. (Bioinformatics, 34(23), 2018, pg. 4039–4045).
Liu et al. is applied to claims 1-8, 11-16, and 18-20 above.
Regarding claim 5, Liu et al. does not teach a structure map (re: clm. 5, …the protein information includes at least one of a contact map indicating a bond between amino acid residues forming the protein, a distance map indicating a distance between amino acid residues forming the protein, or a tertiary structure of the protein.)
Hanson et al. teaches a protein contact map prediction method by stacking residual convolutional networks with two-dimensional residual bidirectional recurrent LSTM networks, and using both one-dimensional sequence-based and two-dimensional evolutionary coupling-based information (Abstract, pg. 2, Fig. 1, pg. 4041). Hanson et al. includes predicted structures and protein sequences as input to their model as disclosed in sec. 2.2-2.3, pg. 4-5.
Hanson et al. does not explicitly teach inversion information (re: clm. 5, … The information processing apparatus according to claim 4…)
Applying the KSR standard of obviousness to Liu et al. and Hanson et al. the examiner concludes that there was a finding that there was some teaching, suggestion, or motivation, either in the references themselves or in the knowledge generally available to one of ordinary skill in the art, to modify the reference or to combine reference teachings to predictably lead to a method to generate a predicted contact map derived from sequence information.
One of ordinary skill in the art of bioinformatics would be motivated to apply Hanson et al.’s protein contact map prediction method because the application would lead to a stronger method of generating protein information from a given sequence.
There would have been a reasonable expectation of success because both Liu et al. and Hanson et al. teach methods within the same field of invention. In support of this expectation, both Liu et al. and in, Hanson et al. apply bidirectional LTSM to protein sequences (Liu et al. Results pg. 5, and Hanson et al. Abstract and Title), with Liu et al. stating a need for structural output on pg. 19 (“Despite much advances in recent years, the folded structure of proteins still cannot be reliably predicted from their primary
amino acid sequences…), which Hanson et al.’s method could provide. Therefore, claim 5 of the applicant’s invention would have been prima facie obvious to one of skill in the art at the time of filing of the application, absent evidence to the contrary.
Claims 9 and 10 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. in view of Hanson et al. as applied to claims 1-8, 11-17, and 18-20 above and in view of James K. Baker (US11270188B2).
Liu et al. in view of Hanson et al. is applied to claims 1-8, 11-17 and 18-20 above.
Regarding claim 10, Liu et al. in view of Hanson et al. teach a machine learning method of predicting protein structures with methods of integrating protein sequence information and inversion information and with the machine learning model being trained and validated with error mitigations (re: clm. 9, … the first prediction unit includes a first machine learning model trained on a basis of an error between the first protein information for learning and the correct answer data…clm. 10, … The information processing apparatus according to claim 8…, the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information for learning and the correct answer data…)
Liu et al. in view of Hanson et al. does not teach a second machine learning model re-trained on a basis of an error (re: clm. 9, … the first machine learning model is re-trained on a basis of an error between the protein information predicted using the first protein information for learning and the second protein information for learning as inputs and the correct answer data…clm. 10, … the second prediction unit includes a second machine learning model trained on a basis of an error between the second protein information for learning and the correct answer data, and the second machine learning model is re-trained on a basis of an error between the protein information predicted using the first protein information for learning and the second protein information for learning as inputs and the correct answer data.)
Baker teaches:
a combination machine learning system with a joint optimization objective,
producing joint optimization outputs from a number (N) of neural networks (clm. 1),
modifying, the joint optimization network after training the combination machine-learning system; and
re-training the combination machine-learning system with the modified joint network (clm. 15, re: clm. 9, … the first machine learning model is re-trained on a basis of an error…clm. 10, … the second machine learning model is re-trained on a basis of an error…)
Applying the KSR standard of obviousness to Liu et al., Hanson et al., and Baker, the examiner concludes that there was a finding that there was some teaching, suggestion, or motivation, either in the references themselves or in the knowledge generally available to one of ordinary skill in the art, to modify the reference or to combine reference teachings to predictably lead to a method of predicting protein contact maps and protein structure with error testing to validate the protein information predicted from sequence and inversion information with a re-training method built in.
One of ordinary skill in the art of bioinformatics would be motivated to apply Baker’s re-training method to the protein structure prediction method of Liu et al. and Hanson et al. because the application would lead to a stronger method of generating protein information from a given sequence.
There would have been a reasonable expectation of success because both Liu et al. in view of Hanson et al., and Baker utilize machine learning methods with optimization techniques. In support of this expectation, Liu et al., Hanson et al. and Baker each state a desire to minimize prediction errors in their text (Liu et al., pg. 5; Hanson et al., pg. 4 “…a bottleneck layer…”; Baker, clm. 24). Therefore, claims 9 and 10 of the applicant’s invention would have been prima facie obvious to one of skill in the art at the time of filing of the application, absent evidence to the contrary.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOHN T STUBBS whose telephone number is (571)272-0340. The examiner can normally be reached M-F 8-5 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Larry Riggs can be reached at 571-270-3062. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.T.S./Examiner, Art Unit 1686
/LARRY D RIGGS II/Supervisory Patent Examiner, Art Unit 1686