DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
2. Claims 1-13 are currently pending and under exam herein.
Claims 1-13 are rejected.
Priority
3. Claimed benefit of foreign priority to the prior Chinese Patent Application 202310210018.3, filed on March 7, 2023, and Chinese Patent Application 202310714738.3, filed on June 15, 2023 is acknowledged. Certified copies of the priority documents have been received. In this action, all claims are examined as though they had an effective filing date 7 March 2023. In future actions, the effective filing date of one or more claims may change, due to amendments to the claims, or further analysis of the disclosure(s) of the priority application(s).
Information Disclosure Statement
4. No information disclosure statement has been filed herein.
Drawings
5. The drawing submitted 11 July 2023 are accepted by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
6. Claims 8-11 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The term “large” in claims 8 and 9 is a relative term which renders the claims indefinite. The term “large” is not defined by the claims, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. It is unclear how many species would need to have the phenotype-genotype relationship in the invention. For the purpose of examination and with the broadest reasonable interpretation, a ‘large number of species’ will be considered to be more than one. Claims 10 and 11 are similarly rejected because they depend from claim 8 and fail to resolve the indefiniteness issue.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
7. Claims 1-8 and 12-13 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 2A, Prong 1
In accordance with MPEP § 2106, claims found to recite statutory subject matter (Step 1: YES) are then analyzed to determine if the claims recite any concepts that equate to an abstract idea, law of nature or natural phenomenon (Step 2A, Prong 1). In the instant application, the claims recite the following limitations that equate to an abstract idea:
Claim 1 recites: performing protein domain analysis on the reference genome data of bacteria
Claim 1 recites: determining a feature value dataset based on all protein domains obtained by analysis for each bacterium
Claim 1 recites: training a bacterial shape prediction model based on the shape information of each bacterium and the feature value dataset
Claim 1 recites: determining influence weights of each protein domain on bacterial shape according to the bacterial shape prediction model
Claim 1 recites: determining candidate genes which regulate the shape of bacteria based on the influence weights
Claim 2 recites: the method according to claim 1, wherein the step of determining the feature value dataset based on all protein domains obtained by the analysis for each bacterium comprises: constructing a protein domain frequency matrix based on all protein domains obtained by the analysis for each bacterium
Claim 2 recites: obtaining the feature value dataset
Claim 3 recites: the method according to claim 1, wherein the step of training the bacterial shape prediction model based on the shape information of each bacterium and the feature value dataset comprises: determining a grouping list based on the shape information of each 32 bacterium, wherein the grouping list comprises a test group and a training group;
Claim 3 recites: performing multiple trainings using the shape information of each bacterium in the training group and the feature value dataset
Claim 3 recites: adjusting a proportion of bacterial species corresponding to various shapes in the test group and the training group based on the prediction indicator values to obtain an adjusted grouping list
Claim 3 recites: performing next training based on the adjusted grouping list; and obtaining the bacterial shape prediction model in response to the prediction indicator value reaching a preset threshold.
Claim 3 recites: obtaining prediction indicator values for the test group predicted by models trained in each round
Claim 4 recites: determining the influence weight of a protein domain corresponding to the current node on bacterial shape based on the degree of purity reduction of the current node when performing splitting in each of the decision trees
Claim 5 recites: the method according to claim 1, wherein the step of determining candidate genes which regulate the shape of bacteria, based on the influence weights comprises: performing cross-validation on the bacterial shape prediction model to determine a relationship between a number of protein domains and an error rate of the bacterial shape prediction model
Claim 4 recites: the method according to claim 1, wherein the step of determining the influence weights of each protein domain on bacterial shape based on the bacterial shape prediction model comprises: obtaining each decision tree in the bacterial shape prediction model, with each protein domain used as a classification node when performing feature classification
Claim 4 recites: obtaining a degree of purity reduction of current node when performing splitting in each of the decision trees
Claim 5 recites: determining a number of key protein domains based on the relationship between the number of protein domains and the error rate of the bacterial shape prediction model
Claim 5 recites: determining the candidate genes based on the number of key protein domains and the respective influence weights.
Claim 6 recites: determining that the candidate genes are not key genes for regulating a shape of target bacteria in response to the shape information after knocking out the candidate genes being identical to before
Claim 6 recites: determining that the candidate genes are key genes for regulating the shape of the target bacteria in response to the shape information after knocking out the candidate genes being different from before
Claim 7 recites: the method according to claim 6, wherein the method further comprises: in response to a number of candidate genes which are not key genes for regulating the shape of the target bacteria exceeding a preset number, returning to the step of determining the number of key protein domains based on the relationship between the number of protein domains and the error rate of the bacterial shape prediction model
Claim 7 recites: determining a new number of key protein domains
Claim 7 recites: determining new candidate genes based on the new number of key protein domains and respective influence weights of the key protein domains
Claim 8 recites: performing protein domain analysis on the reference genome data of target species
Claim 8 recites: determining a feature value dataset based on protein domains obtained from the protein domain analysis
Claim 8 recites: training a phenotype prediction model for the large number of species based on the phenotype information and the feature value dataset
Claim 8 recites: determining influence weights of each protein domain on the phenotype of the target species according to the phenotype prediction mode
Claim 8 recites: determining the candidate genes which regulate the phenotype of the target species based on the influence weights of each protein domain
The limitations regarding 'performing protein domain analysis' (which is known in the art to involve math such as hidden Markov models and statistical significance scorings), 'determining a feature value dataset' (which involves counting the features of each type in each bacteria type), 'adjusting a proportion' of samples between the training and test sets (which involves calculating simple proportions), and 'determining a [new] number of key protein domains' based on a relationship (which involves counting), are verbal equivalents that describe a mathematical calculation that is performed as the limitation and are so simple that they could be performed in the human mind or with pen and paper. Therefore, these limitations fall under the "Mathematical concepts" and "Mental processes" groupings of abstract ideas.
The limitations regarding 'training a bacterial shape prediction model', 'training a phenotype prediction model', 'performing multiple trainings', 'determining influence weights' according to the model, 'performing cross-validation', ‘obtaining each decision tree’, 'obtaining a degree of purity reduction' and 'performing a next training'; are directed to training and making predictions with a mathematical model, which involves mathematical concepts such as probability, statistics, information theory. Therefore, these limitations fall under the "Mathematical concepts" grouping of abstract ideas.
The remaining limitations for 'determining [new] candidate genes which regulate the shape', 'determining the candidate genes which regulate the phenotype (which involves interpreting the weights output from the model), 'constructing a protein domain frequency matrix', 'determining a grouping list' (which involves grouping bacteria based on the data), ‘obtaining prediction indicator values’, ‘obtaining the feature value dataset’ and 'determining the candidate genes are or are not key genes' (by interpreting results of knockout data) are generically recited data analysis steps that can be practically performed in the human mind because the human mind is capable of identifying relevant information, making mental evaluations or judgments, comparing values, and determining information from other values. Therefore, these limitations fall under the "Mental processes" groupings of abstract ideas.
While claims 12 and 13 recite performing the analysis with a processor, there are no additional limitations that indicate that this processor requires anything other than carrying out the recited mental process or mathematical concept in a generic computer environment. Merely reciting that a mental process is being performed in a generic computer environment does not preclude the steps from being performed practically in the human mind or with pen and paper as claimed. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then if falls within the "Mental processes" grouping of abstract ideas. As such, claims 1-8 and 12-13 recite an abstract idea (Step 2A, Prong 1: YES).
Whereas claims 1-8 and 12-13 are directed to an abstract idea, the aspects of the invention that are directed to identifying correlations between presence of protein domains and bacterial shape or other phenotypes are also directed to natural phenomena, and thus are also subject to the “natural phenomenon” judicial exception (see MPEP 2106.04(b)).
Step 2A, Prong 2
Claims found to recite a judicial exception under Step 2A, Prong 1 are then further analyzed to determine if the claims as a whole integrate the recited judicial exception into a practical application or not (Step 2A, Prong 2). This judicial exception is not integrated into a practical application because the claims do not recite an additional element that reflects an improvement to technology or applies or uses the recited judicial exception in some other meaningful way. Rather, the instant claims recite additional elements that amount to mere instructions to implement the abstract idea in a generic way and in a generic computing environment or insignificant extra-solution activity. Specifically, the claims recite the following additional elements:
Claim 1 recites: obtaining reference genome data of bacteria
Claim 1 recites: obtaining shape information of each bacterium
Claim 8 recites: a method for identifying candidate genes which regulate a phenotype of a large number of species, comprising: obtaining reference genome data of a large number of species
Claim 6 recites: the method according to claim 5, wherein the method further comprises: obtaining shape information of a target bacterium after knocking out the candidate genes
Claim 7 recites: obtaining the shape information of the target bacteria after knocking out the new candidate gene
Claim 8 recites: obtaining phenotype information of the large number of species
claim 8, to obtain a species with altered phenotype
Claim 12 recites: a computer device comprising a memory
Claim 12 recites: a processor
Claim 12 recites: a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to claim 1 when executing the computer program
Claim 13 recites: a non-transitory computer-readable storage medium
Claim 13 recites: storing a computer program, wherein the computer program is executed by a processor to implement the method according to claim 1.
The limitations for 'obtaining reference genome data', 'obtaining shape information', 'obtaining phenotype information' merely serve to gather data that is used an input for the judicial exception. Therefore, these limitations are mere data gathering activities. As set forth in MPEP 2106.05(g), mere data gathering activity has been identified by the courts as insignificant extra-solution activity that does not provide a practical application.
There are no limitations that indicate that the processor requires anything other than a generic computing system. As such, these limitations equate to mere instructions to implement the abstract idea on a generic computer that the courts have stated does not render an abstract idea eligible in Alice Corp., 573 U.S. at 223, 110 USPQ2d at 1983. See also 573 U.S. at 224, 110 USPQ2d at 1984 (see MPEP 2106.05(f)). The limitations to store a computer program are insignificant extra-solution activities because the storage does not add a meaningful limitation to the computer implemented process.
The above recited additional elements do not provide a practical application of the recited judicial exception. As such, claims 1-8 and 12-13 are directed to an abstract idea (Step 2A, Prong 2: NO).
Step 2B
Claims found to be directed to a judicial exception are then further evaluated to determine if the claims recite an inventive concept that provides significantly more than the judicial exception itself (Step 2B).
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims recite additional elements that equate to mere instructions to apply the recited exception in a generic computing environment or well-understood, routine and conventional activity.
The limitations directed to gathering data do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As set forth in MPEP section 2106.05(d), the courts have recognized that limitations directed to data gathering that are claimed as insignificant extra-solution activity are routine, well understood and conventional (Mayo Collaborative servs. V. Prometheus Labs., Inc., 566 U.S. at 79, 101 USPQ2d at 1968). In addition, the specification indicates that gene and protein sequences from reference genomes can be obtained from public databases (p. 6, para. 8-13), and that phenotypic information can be obtained by querying public bacterial information databases (such as BacDive website)(p. 7, para. 12-18). Therefore, the insignificant extra-solution activities recited in the claims are well-understood, routine and conventional.
The limitations of claims 12 and 13, pertaining to the computer system used to execute the method, are directed to performing judicial exceptions with a generic computing system on a generic computer. These limitations are not sufficient to amount to significantly more than the judicial exception because, as set forth in the MPEP section 2106.05(d)(II)), using a generic computing environment or generic computer to perform the judicial exception, has been deemed well-understood, routine and conventional activity including receiving or transmitting data over a network (Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362), performing repetitive calculations (Bancorp Services v. Sun Life, 687 F.3d 1266, 1278, 103 USPQ2d 1425, 1433 (Fed. Cir. 2012)), and storing and retrieving information in memory (Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015)). Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception are insufficient to provide significantly more.
The additional elements do not comprise an inventive concept when considered individually or as an ordered combination that transforms the claimed judicial exception into a patent-eligible application of the judicial exception. Therefore, the claims do not amount to significantly more than the judicial exception itself (Step 2B: No). As such, claims 1-8 and 12-13 are not patent eligible.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
8. Claims 1, 3, 8, and 12-13 are rejected under 35 U.S.C. 102 (a)(1) as being unpatentable over Weimann et al. (mSystems, Methods and Protocols Molecular Biology and Physiology, Vol. 1, p. 1-19). The italicized text corresponds to the instant claim limitations.
Pertaining to claims 1 and 8, Weimann et al. teaches development of Traitar, a microbial trait analyzer for prediction of 67 phenotypes directly from genome sequences. Weimann et al. further discloses that Traitar provides classification models based on protein family annotations (i.e. pfam domain annotations) for a wide variety of different phenotypes (including the use of various substrates as source of carbon and energy for growth, oxygen requirement, morphology, antibiotic susceptibility, and enzymatic activity) across 234 bacterial species. Weimann et al. further discloses that morphology includes bacillus or coccobacillus shape, coccus shape, coccus with predominant clusters or groups, coccus with predominate pairs or chains, motile, spore formation and other morphologies. Weimann et al. further teaches that by the method, the phenotype predictions are mapped to predicted protein coding genes (p. 2, para. 4; Fig. 1; Fig. 2; Table 1; Table 3; p. 3, para. 2; a method of identifying candidate genes, which regulate shape of bacteria (claim 1); a method for identifying candidate genes which regulate a phenotype of a large number of species (claim 8).
Pertaining to claims 1 and 8, Weimann et al. discloses that the input to Traitar is either a nucleotide sequence FASTA file for every sample from each bacterial species, which is run through gene prediction software, or a protein sequence FASTA file. Weimann et al. further discloses obtaining genomic data of microbial communities by single cell genome and metagenomic sequencing and obtaining genomic data from microbial isolates by shotgun genome sequencing (p. 3, para. 3; Fig. 1; Fig. 5; obtaining reference genome data of bacteria (claim 1); obtaining reference genome data of a large number of species (claim 8).
Pertaining to claims 1 and 8, Weimann et al. discloses the Traitar workflow, whereby protein-coding genes from various different bacterial species are predicted from input genomes and then annotated with Pfam protein families (p. 3, para. 3; Fig. 2; performing protein domain analysis on the reference genome data of bacteria; determining a feature value dataset based on all protein domains obtained by analysis for each bacterium (claim 1); performing protein domain analysis on the reference genome data of target species; determining a feature value dataset based on protein domains obtained from the protein domain analysis (claim 8).
Pertaining to claims 1 and 8, Weimann et al. discloses obtaining phenotype data from the GIDEON database, including traits that can be grouped into categories, including the use of various substrates as sources of carbon and energy for growth, oxygen requirement, morphology (i.e. shape), antibiotic susceptibility, and enzymatic activity (p. 15, para. 4; obtaining shape information of each bacterium (claim 1); obtaining phenotype information of the large number of species (claim 8).
Pertaining to claims 1 and 8, Weimann et al. teaches using both the phenotypic data and the Pfam annotations from a large number of microbial genomes for training ‘phenotype classification models’ to predict the presence or absence of each trait for every input sequence (p. 2, para. 4; p. 3, para. 3; Fig. 2; training a bacterial shape prediction model based on the shape information of each bacterium and the feature value dataset (claim 1); training a phenotype prediction model for the large number of species based on the phenotype information and the feature value dataset (claim 8).
With respect to claims 1 and 8, Weimann et al. teaches that the Traitar software classifier associates the predicted phenotypes (including morphology) with the protein families that contributed to these predictions. Weimann teaches determining influence weights of protein domains on a phenotype according to a phenotype prediction model by determining weights associated with protein-family features in linear support vector machine phenotype prediction models. Weimann et al. expressly teaches that the weighs of the linear SVMs can be directly linked to features relevant to the classification and uses the SVM model weights to identify important protein families associates with the predicted phenotypes. Weimann et al. discloses further ranking the protein family features by their correlation with the phenotype using Pearson’s correlation coefficient (Fig. 2; p. 3, para. 3; p. 17, para. 4; determining influence weights of each protein domain on bacterial shape according to the bacterial shape prediction model (claim 1); determining influence weights of each protein domain on the phenotype of the target species according to the phenotype prediction model (claim 8).
Pertaining to claims 1 and 8, Weimann et al. teaches that the SVM outputs are used to predict protein family features (domains) that influence phenotypes. Weimann further discloses that protein families that are important for the phenotype predictions are further mapped to the predicted protein-coding genes. Therefore, in identifying the important protein domains for cell shape, they also identify the proteins containing those domains and the genes encoding them (p. 17, para. 4; p. 3; para. 3; Fig. 2; determining candidate genes which regulate the shape of bacteria based on the influence weights (claim 1); and determining candidate genes which regulate the phenotype of the target species based on the influence weights of each protein domain (claim 8).
Pertaining to claim 3, Weimann et al. discloses using cross validation to assess the performance of the classifiers individually for each phenotype as described below: for a given phenotype, bacterial samples that were annotated with that phenotype were divided into 10 folds. Each fold was selected once for testing the model which was trained on the remaining folds. Weimann et al. further discloses that after models were trained by cross validation, they were tested on test datasets. Weimann et al. discloses that accuracies were calculated from the cross validation and performance was calculated based on a consensus vote. The performance for each phenotype was reported as an accuracy, which indicates that the performance was calculated at a specific threshold. Finally Weismann et al. discloses that the models were tested on an independent test set (i.e. another 42 bacteria from GIDEON and 296 bacteria from Bergey’s Manual of Systematic Bacteriology) predictions were combined into a consensus vote (p. 16, para. 7 – p. 17, para. 2 (i.e. “cross validation section”); p. 2, para. 4; Table 2; the method according to claim 1, wherein the step of training the bacterial shape prediction model based on the shape information of each bacterium and the feature value dataset comprises: determining a grouping list based on the shape information of each bacterium, wherein the grouping list comprises a test group and a training group; performing multiple trainings using the shape information of each bacterium in the training group and the feature value dataset; obtaining prediction indicator values for the test group predicted by models trained in each round, adjusting a proportion of bacterial species corresponding to various shapes in the test group and the training group based on the prediction indicator values to obtain an adjusted grouping list, and performing next training based on the adjusted grouping list; and obtaining the bacterial shape prediction model in response to the prediction indicator value reaching a preset threshold).
With respect to claim 12, Weimann et al. discloses that their software for executing the method (Traitar) can be run on a standard laptop with Linux/Unix. The run time (wall clock time) for annotating and phenotyping a typical microbial genome with 3 Mbp is 9 min (3 min/Mbp) on an Intel Core i5-2410M dual-core processor with 2.30 GHz, requiring only a few megabytes of memory (p. 16, para. 5 section “software requirements”; a computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to claim 1 when executing the computer program).
Regarding claim 13, Weimann et al. discloses that their software for executing the method (Traitar) can be run on a standard laptop with Linux/Unix. The run time (wall clock time) for annotating and phenotyping a typical microbial genome with 3 Mbp is 9 min (3 min/Mbp) on an Intel Core i5-2410M dual-core processor with 2.30 GHz, requiring only a few megabytes of memory. As the method is performed using software, the method inherently includes the use of a processor and memory (p. 16, para. 5 section “software requirements”; a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the method according to claim 1).
9. Claims 9-11 are rejected under 35 U.S.C. 102 (a)(1) as being unpatentable over Ursell et al. (BMC Biology, Vol. 15, 2017, p. 1-15). The italicized text corresponds to the instant claim limitations.
Regarding claim 9, Ursell et al. teaches the quantification of bacterial cell shape and size across a genomic-scale knockout library. Ursell et al. further teaches that some gene deletions had an effect on the cell phenotype (p. 3, col. 1, para. 2; Figure 5a; a method of regulating a phenotype of a large number of species, comprising: knocking out one or more of the candidate genes of the species according to claim 8, to obtain a species with altered phenotype).
Pertaining to claim 10, Ursell et al. teaches the quantification of bacterial cell shape and size across a genomic-scale knockout library. Ursell et al. further teaches that some gene deletions had an effect on the cell phenotype (p. 3, col. 1, para. 2; Figure 5a; the method according to claim 9, wherein the species comprise a bacterium, fungus, virus, plant, or animal).
Concerning claim 11, Ursell et al. teaches the quantification of bacterial cell shape and size across a genomic-scale knockout library. Ursell et al. further teaches that some gene deletions had an effect on the cell shape (p. 3, col. 1, para. 2; Figure 5a; the method according to claim 9, wherein the phenotype comprise shape, temperature, metabolic products, height, stress resistance, or mode of locomotion).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
10. Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Weimann et al. (mSystems, Methods and Protocols Molecular Biology and Physiology, Vol. 1, p. 1-19), as applied to claims 1, 3, 8 and 12-13 above, further in view of Davis et al. (Scientific Reports, 2016, p. 1-12). The italicized text corresponds to the instant claim limitations.
The limitations of claims 1, 3, 8 and 12-13 were taught by Weimann et al. above.
The limitations of claims 9-11 were taught by Ursell et al. above.
Pertaining to claim 2, Weimann et al. is silent to: the method according to claim 1, wherein the step of determining the feature value dataset based on all protein domains obtained by the analysis for each bacterium comprises: constructing a protein domain frequency matrix based on all protein domains obtained by the analysis for each bacterium, and obtaining the feature value dataset. However, this limitation was known in the art at the time of the effective filing date of the invention as taught by Davis et al.
Pertaining to claim 2. Davis et al. teaches generation of a matrix that is used for machine learning to predict association of k-mer genomic sequences with resistance to antibiotics (phenotype), the matrix comprising each gene sequence (analogous to each protein domain) with each phenotype (resistance to antibiotics). (Fig. 1 “matrix of merged k-mer counts”; the method according to claim 1, wherein the step of determining the feature value dataset based on all protein domains obtained by the analysis for each bacterium comprises: constructing a protein domain frequency matrix based on all protein domains obtained by the analysis for each bacterium, and obtaining the feature value dataset).
An invention would have been prima facie obvious to one of ordinary skill in the art at the effective filing date of the invention if some motivation in the prior art would have led that person to combine the prior art teachings to arrive at the claimed invention. Davis et al. taught that this k-mer approach to identifying genome associations with antimicrobial resistance is an advantage because the classifier can identify entire gene regions as well as SNP-level variation that can affect resistance, which is a more detailed resolution than other gene-level approaches (p. 10, para. 5). Therefore, one of ordinary skill in the art would have been motivated to utilize the antimicrobial resistance prediction taught by Davis et al. in the microbial trait analyzer taught by Weimann et al. in order to improve resolution of associations. Furthermore, one of ordinary skill in the art would predict that methods taught by Davis et al. could be readily added to the microbial trait analyzer taught by Weimann et al. with a reasonable expectation of success because they both pertain to predicting phenotypes from genome-wide data across various species of bacteria. The invention is therefore prima facie obvious.
11. Claims 4-5 are rejected under 35 U.S.C. 103 as being unpatentable over Weimann et al. (mSystems, Methods and Protocols Molecular Biology and Physiology, Vol. 1, p. 1-19), as applied to claims 1, 3, 8 and 12-13 above, in view of Simon et al. (PeerJ 2022, p. 1-20), as evidenced by Goldstein et al. (Statistical Applications in Genetics and Molecular Biology: Vol. 10, 2011, p. 1-36). The italicized text corresponds to the instant claim limitations.
The limitations of claims 1, 3, 8 and 12-13 were taught by Weimann et al. above.
The limitations of claim 2 were taught by Weimann et al. and Davis et al. above.
Regarding claim 5, Weimann et al. discloses that Traitar software annotates the proteins with protein families and that protein families that were found to be important for the phenotype predictions were further mapped to the predicted protein-coding genes. Therefore, in identifying the important protein domains for cell shape, they also identify the proteins containing those domains and the corresponding protein-coding genes (Fig. 2; p. 3; para. 3; determining the candidate genes based on the number of key protein domains and the respective influence weights).
Pertaining to claim 4, Weimann et al. is silent to: the method according to claim 1, wherein the step of determining the influence weights of each protein domain on bacterial shape based on the bacterial shape prediction model comprises: obtaining each decision tree in the bacterial shape prediction model, with each protein domain used as a classification node when performing feature classification; obtaining a degree of purity reduction of current node when performing splitting in each of the decision trees; and determining the influence weight of a protein domain corresponding to the current node on bacterial shape based on the degree of purity reduction of the current node when performing splitting in each of the decision trees (claim 4); the method according to claim 1, wherein the step of determining candidate genes which regulate the shape of bacteria, based on the influence weights comprises: performing cross-validation on the bacterial shape prediction model to determine a relationship between a number of protein domains and an error rate of the bacterial shape prediction model; determining a number of key protein domains based on the relationship between the number of protein domains and the error rate of the bacterial shape prediction model (claim 5). However, these limitations were known in the art at the time of the effective filing date of the invention as taught by Simon et al.
Pertaining to claim 4, Simon et al. teaches RFPDR, a method of predicting protein-phenotype interactions in plants using a random forest modeling approach. Simon et al. further disclose that features in the model were 153 disease resistance protein predictors across 79 species (including protein length and sequence-based composition estimates (including amino acid, dipeptide and tripeptide composition features), autocorrelation, composition/transition/distribution, and conjoint triad descriptors), and that that RF models are built on decision trees and that feature weights/importance scores (Gini importance values) were used to select important features to include in the model. As evidenced by Goldstein et al., in random forest-based classification, samples are classified by a decision tree wherein at each node, the algorithm recursively searches for a split that partitions the data in a way to minimize a splitting criterion (gini). After a stopping criterion is met, the final splits partition the predictor space into hyper rectangles. These regions are referred to as nodes of the tree. The variable at the top of the tree represents the “strongest” splitting variables and subsequent variables are conditional on those variables above it. Goldstein et al. further disclose that the splitting criterion is the gini-index and that trees are grown to maximal depth (i.e. when nodes are pure of class) (Simon et al. p. 3, para. 2-3; p. 11, para. 5; p. 12, para. 2-3; p. 5, para. 2-3; Goldstein et al. p. 7, para. 2-3; p. 10, para. 2; Fig. 3; p. 21, para. 3-p. 22, para. 1; the method according to claim 1, wherein the step of determining the influence weights of each protein domain on bacterial shape based on the bacterial shape prediction model comprises: obtaining each decision tree in the bacterial shape prediction model, with each protein domain used as a classification node when performing feature classification; obtaining a degree of purity reduction of current node when performing splitting in each of the decision trees; and determining the influence weight of a protein domain corresponding to the current node on bacterial shape based on the degree of purity reduction of the current node when performing splitting in each of the decision trees).
Pertaining to claim 5, Simon et al. discloses performing 10-fold cross-validation for two types of models, full dimension models (FD-RFPDR) and reduced dimension models (RD-RFPDR). Simon et al. further discloses that performance metrics were defined from the confusion matrices, which summarize the results of the RF classification. Specificity, accuracy, precision, sensitivity and F1-score were estimated as well as area under the Receiver Operating Curve. The number of sequence-derived features used in each model were counted with FD-RFPDR using 9,631 features and RD-RFPDR using 1,133 features (p. 4, para. 5-6; Table 1; Fig. 3; the method according to claim 1, wherein the step of determining candidate genes which regulate the shape of bacteria, based on the influence weights comprises: performing cross-validation on the bacterial shape prediction model to determine a relationship between a number of protein domains and an error rate of the bacterial shape prediction model).
Pertaining to claim 5, Simon et al. discloses using the Gini-based importance for the determination of the most discriminant features for automatic disease resistance protein prediction: after ranking features by importance scores, the 1% and 5%-top features and the first features (which accounted for 50% of the importance), were further used for modeling. (p. 5, para. 2; determining a number of key protein domains based on the relationship between the number of protein domains and the error rate of the bacterial shape prediction model).
An invention would have been prima facie obvious to one of ordinary skill in the art at the effective filing date of the invention if some motivation in the prior art would have led that person to combine the prior art teachings to arrive at the claimed invention. Simon et al. taught that RD-RFPDR showed to be sensitive and specific in identifying disease resistance proteins, while robust to data imbalance (p. 14, para. 2). Therefore, one of ordinary skill in the art would have been motivated to utilize the random forest-based modeling of protein-phenotype interactions taught by Simon et al. in the microbial trait analyzer taught by Weimann et al. in order to improve sensitivity and specificity of the model and improve robustness to data imbalances. Furthermore, one of ordinary skill in the art would predict that the random forest-based modeling taught by Simon et al. could be readily added to the microbial trait analyzer taught by Weimann et al. with a reasonable expectation of success because they both pertain to predicting phenotype from genomic/proteomic data across various species. Additionally, both methods use specific aspects of proteins as features rather than the proteins themselves (i.e. Weimann et al. uses protein domains and Simon et al. uses protein regions (for example hydrophobic regions)). The invention is therefore prima facie obvious.
12. Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Weimann et al. (mSystems, Methods and Protocols Molecular Biology and Physiology, Vol. 1, p. 1-19), as applied to claims 1, 3, 8 and 12-13 above, in view of Simon et al. (PeerJ 2022, p. 1-20), as evidenced by Goldstein et al. (Statistical Applications in Genetics and Molecular Biology: Vol. 10, 2011, p. 1-36), as applied to claims 4-5 above, and further in view of Ursell et al. (BMC Biology, Vol. 15, 2017, p. 1-15). The italicized text corresponds to the instant claim limitations.
The limitations of claims 1, 3, 8, and 12-13 were taught by Weimann et al. above.
The limitations of claim 2 were taught by Weimann et al. and Davis et al. above.
The limitations of claims 4 and 5 were taught by Weimann et al. and Simon et al. above.
Pertaining to claim 6, Weimann et al. and Simon et al. are silent to the method according to claim 5, wherein the method further comprises: obtaining shape information of a target bacterium after knocking out the candidate genes; determining that the candidate genes are not key genes for regulating a shape of target bacteria in response to the shape information after knocking out the candidate genes being identical to before; and determining that the candidate genes are key genes for regulating the shape of the target bacteria in response to the shape information after knocking out the candidate genes being different from before. However, this limitation was known in the art at the time of the effective filing date of the invention as taught by Ursell et al.
Pertaining to claim 6, the claim is interpreted to mean that given experimental data of morphology of candidate gene deletion strains, some predictions are correct and some are incorrect.
Pertaining to claim 6, Ursell et al. teaches the quantification of bacterial cell shape and size across a genomic-scale knockout library. Ursell et al. further teaches that some gene deletions had no effect on cell phenotype, while others did (p. 3, col. 1, para. 2; Figure 5a; the method according to claim 5, wherein the method further comprises: obtaining shape information of a target bacterium after knocking out the candidate genes; determining that the candidate genes are not key genes for regulating a shape of target bacteria in response to the shape information after knocking out the candidate genes being identical to before; and determining that the candidate genes are key genes for regulating the shape of the target bacteria in response to the shape information after knocking out the candidate genes being different from before).
An invention would have been prima facie obvious to one of ordinary skill in the art at the effective filing date of the invention if some motivation in the prior art would have led that person to combine the prior art teachings to arrive at the claimed invention. Ursell et al. taught that their study provides reproducible quantification of cell shape as well as the ability to test quantitative models, (p. 3, col. 1, para. 2). Therefore, one of ordinary skill in the art would have been motivated to utilize the knockout data taught by Ursell et al. in the microbial trait analyzer taught by Weimann et al. and Simon et al. in order to test the predictions of the morphology prediction model. Furthermore, one of ordinary skill in the art would predict that the morphological analysis of the knockout library taught by Ursell et al. could be readily added to the microbial trait analyzer taught by Weimann et al. and Simon et al. with a reasonable expectation of success because both methods and data pertain to correlating genotype of bacteria with morphology on a genomic-scale. Furthermore, the system and data of Ursell et al. were intended for use in validation studies; Ursell et al. discloses that precise quantification of cell morphology in bacteria and eukaryotes will undoubtedly be a valuable tool for mapping genotype-phenotype relationships and validating models (p. 12, col. 2, para. 3; p. 3, col. 1, para. 2). The invention is therefore prima facie obvious.
13. Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Weimann et al. (mSystems, Methods and Protocols Molecular Biology and Physiology, Vol. 1, p. 1-19, as applied to claims 1, 3, 8 and 12-13 above, in view of Simon et al. (PeerJ 2022, p. 1-20), as evidenced by Goldstein et al. (Statistical Applications in Genetics and Molecular Biology: Vol. 10, 2011, p. 1-36), as applied to claims 4-5 above, in view of Ursell et al. (BMC Biology, Vol. 15, 2017, p. 1-15), as applied to claim 6 above, and further in view of King et al. (2004, Nature, Vol. 427, p. 247-252). The italicized text corresponds to the instant claim limitations.
The limitations of claims 1, 3, 8 and 12-13 were taught by Weimann et al. above.
The limitations of claim 2 were taught by Weimann et al. and Davis et al. above.
The limitations of claims 4-5 were taught by Weimann et al. and Simon et al. above.
The limitations of claim 6 were taught by Weimann et al., Simon et al. and Ursell et al. above.
Regarding claim 7, Weimann et al., Simon et al. and Ursell et al. are silent to: the method according to claim 6, wherein the method further comprises: in response to a number of candidate genes which are not key genes for regulating the shape of the target bacteria exceeding a preset number, returning to the step of determining the number of key protein domains based on the relationship between the number of protein domains and the error rate of the bacterial shape prediction model, and determining a new number of key protein domains; determining new candidate genes based on the new number of key protein domains and respective influence weights of the key protein domains; and obtaining the shape information of the target bacteria after knocking out the new candidate genes. However, these limitations were known in the art at the time of the effective filing date of the invention as taught by King et al.
Regarding claim 7, the claim is interpreted to mean that upon validation of the predictions by knockout analysis, if the rate of successful validation is low (i.e. the rate of correct predictions is low), the training is repeated.
Regarding claim 7, King et al. teaches an automated looped method of functional genomic hypothesis generation and testing. King et al. discloses that this involves iterative cycles of modeling (to generate hypotheses of gene function), experimental validation (wherein phenotypes of gene deletions are measured) and refining the model with results of the validation experiment. King et al. validates this approach by interrogating the aromatic amino acid synthesis (AAA) pathway in yeast by deletion of genes followed conducting auxotrophic growth experiments on the knock out strains to assess the behaviour (phenotype) of the mutants. King et al. teaches that in using this approach, model performance (i.e. prediction accuracy) improves with each iteration of the modelling/validation. Although King et al. does not disclose deciding to continue the cycle in response to the model performance being below a threshold, given their demonstration of an upwards trend in performance with progressive iterations, it would be obvious to try. (Fig. 1; Fig. 2; Fig. 3; p. 248, col. 1, para. 3 - col. 2, para. 2; p. 248, col. 2, para. 5 – p. 249, col. 1, para. 1; the method according to claim 6, wherein the method further comprises: in response to a number of candidate genes which are not key genes for regulating the shape of the target bacteria exceeding a preset number, returning to the step of determining the number of key protein domains based on the relationship between the number of protein domains and the error rate of the bacterial shape prediction model, and determining a new number of key protein domains; determining new candidate genes based on the new number of key protein domains and respective influence weights of the key protein domains; and obtaining the shape information of the target bacteria after knocking out the new candidate genes.
An invention would have been prima facie obvious to one of ordinary skill in the art at the effective filing date of the invention if some motivation in the prior art would have led that person to combine the prior art teachings to arrive at the claimed invention. King et al. taught that their method of automated cycling of model-based hypothesis generation, experimental validation, and model refinement, significantly outperforms other approaches at a reduced cost (p. 248, para. 1). Therefore, one of ordinary skill in the art would have been motivated to utilize the iterative modeling approach taught by King et al. in the microbial trait analyzer taught by Weimann et al., Simon et al. and Ursell et al. in order to improve performance and reduce cost. Furthermore, one of ordinary skill in the art would predict that the iterative modeling approach taught by King et al. could be readily added to the microbial trait analyzer taught by Weimann et al., Simon et al. and Ursell et al. with a reasonable expectation of success because both methods and data pertain to functional genomic hypothesis generation and experimentation in microbial organisms by machine learning modeling combined with phenotypic analyses of gene deletion strains. The invention is therefore prima facie obvious.
Conclusion
14. No claims are allowed.
Claims 9-11 were analyzed for subject matter eligibility under 35 U.S.C. 101 and were found to contain eligible subject matter at step 2A prong 1 because the limitations of the claims do not contain a judicial exception.
E-mail Communications Authorization
15. Per updated USPTO Internet usage policies, Applicant and/or applicant's representative is encouraged to authorize the USPTO examiner to discuss any subject matter concerning the above application via Internet e-mail communications. See MPEP 502.03. To approve such communications, Applicant must provide written authorization for e-mail communication by submitting the following statement via EFS-Web (using PTO/SB/439) or Central Fax (571-273-8300): "Recognizing that Internet communications are not secure, / hereby authorize the USPTO to communicate with the undersigned and practitioners in accordance with 37 CFR 1.33 and 37 CFR 1.34 concerning any subject matter of this application by video conferencing, instant messaging, or electronic mail. / understand that a copy of these communications will be made of record in the application file."
Written authorizations submitted to the Examiner via e-mail are NOT proper. Written authorizations must be submitted via EFS-Web (using PTO/SB/439) or Central Fax (571-273- 8300). A paper copy of e-mail correspondence will be placed in the patent application when appropriate. E-mails from the USPTO are for the sole use of the intended recipient, and may contain information subject to the confidentiality requirement set forth in 35 USC § 122. See also MPEP 502.03.
Inquiries
16. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JENNIFER J SMITH whose telephone number is (571)272-7801. The examiner can normally be reached Monday-Friday 7:00 AM - 3:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Olivia Wise can be reached at (571) 272-2249. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.J.S./Examiner, Art Unit 1685
/OLIVIA M. WISE/Supervisory Patent Examiner, Art Unit 1685