DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Claims 1-5 are currently pending and under examination herein.
Claims 1-5 are rejected.
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been filed in parent Application No. CN202210535743.3, filed on 5/18/2022. As such, the effective filing date of the claims 1-5 is 2/24/2021.
Drawings
Nucleotide and/or Amino Acid Sequence Disclosures
REQUIREMENTS FOR PATENT APPLICATIONS CONTAINING NUCLEOTIDE AND/OR AMINO ACID SEQUENCE DISCLOSURES
Items 1) and 2) provide general guidance related to requirements for sequence disclosures.
37 CFR 1.821(c) requires that patent applications which contain disclosures of nucleotide and/or amino acid sequences that fall within the definitions of 37 CFR 1.821(a) must contain a "Sequence Listing," as a separate part of the disclosure, which presents the nucleotide and/or amino acid sequences and associated information using the symbols and format in accordance with the requirements of 37 CFR 1.821 - 1.825. This "Sequence Listing" part of the disclosure may be submitted:
In accordance with 37 CFR 1.821(c)(1) via the USPTO patent electronic filing system (see Section I.1 of the Legal Framework for Patent Electronic System (https://www.uspto.gov/PatentLegalFramework), hereinafter "Legal Framework") as an ASCII text file, together with an incorporation-by-reference of the material in the ASCII text file in a separate paragraph of the specification as required by 37 CFR 1.823(b)(1) identifying:
the name of the ASCII text file;
ii) the date of creation; and
iii) the size of the ASCII text file in bytes;
In accordance with 37 CFR 1.821(c)(1) on read-only optical disc(s) as permitted by 37 CFR 1.52(e)(1)(ii), labeled according to 37 CFR 1.52(e)(5), with an incorporation-by-reference of the material in the ASCII text file according to 37 CFR 1.52(e)(8) and 37 CFR 1.823(b)(1) in a separate paragraph of the specification identifying:
the name of the ASCII text file;
the date of creation; and
the size of the ASCII text file in bytes;
In accordance with 37 CFR 1.821(c)(2) via the USPTO patent electronic filing system as a PDF file (not recommended); or
In accordance with 37 CFR 1.821(c)(3) on physical sheets of paper (not recommended).
When a “Sequence Listing” has been submitted as a PDF file as in 1(c) above (37 CFR 1.821(c)(2)) or on physical sheets of paper as in 1(d) above (37 CFR 1.821(c)(3)), 37 CFR 1.821(e)(1) requires a computer readable form (CRF) of the “Sequence Listing” in accordance with the requirements of 37 CFR 1.824.
If the "Sequence Listing" required by 37 CFR 1.821(c) is filed via the USPTO patent electronic filing system as a PDF, then 37 CFR 1.821(e)(1)(ii) or 1.821(e)(2)(ii) requires submission of a statement that the "Sequence Listing" content of the PDF copy and the CRF copy (the ASCII text file copy) are identical.
If the "Sequence Listing" required by 37 CFR 1.821(c) is filed on paper or read-only optical disc, then 37 CFR 1.821(e)(1)(ii) or 1.821(e)(2)(ii) requires submission of a statement that the "Sequence Listing" content of the paper or read-only optical disc copy and the CRF are identical.
Specific deficiencies and the required response to this Office Action are as follows:
Specific deficiency – Nucleotide and/or amino acid sequences appearing in the drawings are not identified by sequence identifiers in accordance with 37 CFR 1.821(d). Sequence identifiers for nucleotide and/or amino acid sequences must appear either in the drawings or in the Brief Description of the Drawings.
Required response – Applicant must provide:
Replacement and annotated drawings in accordance with 37 CFR 1.121(d) inserting the required sequence identifiers;
AND/OR
A substitute specification in compliance with 37 CFR 1.52, 1.121(b)(3) and 1.125 inserting the required sequence identifiers into the Brief Description of the Drawings, consisting of:
A copy of the previously-submitted specification, with deletions shown with strikethrough or brackets and insertions shown with underlining (marked-up version);
A copy of the amended specification without markings (clean version); and
A statement that the substitute specification contains no new matter.
Specification
The specification filed on 7/4/2023 is accepted.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-5 are rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea and a natural phenomenon without significantly more.
In accordance with MPEP 2106, claims found to recite statutory subject matter (Step 1: YES) are then analyzed to determine if the claims recite any concepts that equate to an abstract idea, law of nature, or natural phenomenon (Step 2A, Prong 1).
Claim 1 recites a predicting method of transcription factor binding sites based on weighted multi-granularity scanning, comprising the following steps: carrying out data augmentation on an initial data set D= {D1,D2,…,Dn}, Di= {xk,yk} (1 ≤ k ≤ n) of transcription factor binding sites, where xk represents a DNA sequence fragment and yk represents whether the DNA sequence fragment is a binding site, taking the value as a binding site or a non-binding site, calculating an inverse sequence, a complementary sequence and a complementary inverse sequence of each piece of data, expanding the number of data sets to four times the original number to obtain a data sets D*=D1,D2,…,D4n,Di,=xk',yk' (1≤k'≤4n), and randomly mixing positive and negative samples in the data set D*; carrying out one-hot coding on each piece of DNA sequence data in the data set D* with the formula {1000,A; 0001,T; 0100,C; 0010,G} to obtain a feature vector F1, and then combining with multi-base feature coding for feature representation to obtain a feature vector F2, splicing the feature vectors F1 and F2 to obtain a combined feature representation F, and encoding the result category with the formula: { 1, binding site; 0, non-binding site}; dividing the data set D* after feature representation in step (2) according to the ratio Q:R of the number of samples in the training set to the number of samples in the test set to obtain a training set Dtrain and a test set Dtest, where Q is the number of samples in the training set in the data set D* and R is the number of samples in the test set in the data set D*; Q has a value in the range of 2-5, and R has a value of 1; using T decision trees to calculate a weight vector W=(W1、W2…Wi…Wd) (1 ≤I ≤d) for the training set Dtrain, where d is the length of the feature, and the specific calculation formula is as follows:
PNG
media_image1.png
81
155
media_image1.png
Greyscale
where d is the total number of features, Scorei is the importance score of the i-th column of features in the weight vector W, and the specific calculation formula is as follows:
PNG
media_image2.png
84
209
media_image2.png
Greyscale
where Scorenode(t) is the importance score of the t-th decision tree node, and the specific calculation formula is as follows:
PNG
media_image3.png
47
286
media_image3.png
Greyscale
where Gnode,0 and Gnode,1 represent the Gini index of the nodes belonging to category 0 under the node branch and the Gini index of the nodes belonging to category 1 under the node branch, respectively; Gnode is the Gini index of each node, and the specific formula is as follows:
PNG
media_image4.png
68
275
media_image4.png
Greyscale
where N is the number of samples in the training set Dtrain, Nnode,0 is the number of nodes belonging to category 0, and Nnode,1 is the number of nodes belonging to category 1; carrying out weighted multi-granularity scanning on the feature F of each sample in the training set Dtrain, in which the specific steps are as follows: using a sliding window with a length of μ to slide with a step length of L on the feature vector F and the weight vector W with a length of d, respectively, and extracting the feature vectors in the window separately to obtain fu and wu with a length of μ, where u is the number of times the sliding window slides, and u has a value in the range of 1≤u≤d-μ+1; according to the formula
PNG
media_image5.png
31
94
media_image5.png
Greyscale
. calculating the features of the weighted multi-granularity scanning, where wuT is the transposition of vector wu; sending the feature Fu' into a completely random forest A and an ordinary random forest B to obtain F’Au and F'Bu, respectively, and finally, splicing F’Au and F'Bu to obtain feature F*; and inputting F* into the cascade forest, carrying out the model training to obtain a classification prediction model of transcription factor binding sites, inputting the test set Dtest into the classification prediction model, and outputting a result of 1 or 0; in which 1 indicates that the DNA sequence is a transcription factor binding site, and 0 indicates that the DNA sequence is a non-transcription factor binding site.
Claim 2 recites the predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in the multi-base feature coding method, the length L of the feature column is obtained according to the formula L=4m, where m is the length of the base in the multi-base, m has a value of 3, bases A, T, C and G form a sequence set C with a length of 3bp: {'AAA', 'AAT', 'AAG', 'AAC', 'ATA', 'ATT', 'ATG', 'ATC', 'AGA', 'AGT', 'AGG', 'AGC', 'ACA', 'ACT', 'ACG', 'ACC', 'TAA', 'TAT', 'TAG', 'TAC', 'TTA', 'TTT', 'TTG', 'TTC', 'TGA', 'TGT', 'TGG', 'TGC', 'TCA', 'TCT', 'TCG', 'TCC', 'GAA', 'GAT', 'GAG', 'GAC', 'GTA', 'GTT', 'GTG', 'GTC', 'GGA', 'GGT', 'GGG', 'GGC', 'GCA', 'GCT', 'GCG', 'GCC', 'CAA', 'CAT', 'CAG', 'CAC', 'CTA', 'CTT', 'CTG', 'CTC', 'CGA', 'CGT', 'CGG', 'CGC', 'CCA', 'CCT', 'CCG', 'CCC'}, each element in set C is set as a feature column, there are 64 feature columns in total, and its element is the feature name of the feature column; the calculation method of the feature vector F2 is as follows: from the beginning end of the DNA sequence sample, a window with a step size of 1 and a length of 3bp slides on the DNA sequence sample and extracts features, and the feature column corresponding to the sequence in the window has a value of 1 until the end of the DNA sequence sample, that is, the length of the feature vector F2 is 64.
Claim 3 recites the predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in step (3), Q has a value of 4, and R has a value of 1.
Claim 4 recites the predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in step (4), T has a value of 462, and the maximum depth of the tree is 11.
Claim 5 recites the predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in step (5), μ has a value of 50, and L has a value of 1.
The limitations of using T decision trees to calculate a weight vector W=(W1、W2…Wi…Wd) (1 ≤I ≤d) for the training set Dtrain, where d is the length of the feature, and the specific calculation formula is as follows:
PNG
media_image1.png
81
155
media_image1.png
Greyscale
where d is the total number of features, calculating the importance score of the i-th column of features in the weight vector W, represented as Scorei ; the calculations of the importance score of the t-th decision tree node represented as Scorenode(t) where Gnode,0 and Gnode,1 represent the Gini index of the nodes belonging to category 0 under the node branch and the Gini index of the nodes belonging to category 1 under the node branch, respectively; calculation of the Gini index of each node where N is the number of samples in the training set Dtrain, Nnode,0 is the number of nodes belonging to category 0, and Nnode,1 is the number of nodes belonging to category 1; calculating the features of the weighted multi-granularity scanning, where wuT is the transposition of vector wu; sending the feature Fu' into a completely random forest A and an ordinary random forest B to obtain F’Au and F'Bu, respectively, and finally, splicing F’Au and F'Bu to obtain feature F*; and inputting F* into the cascade forest, carrying out the model training to obtain a classification prediction model of transcription factor binding sites, inputting the test set Dtest into the classification prediction model are all verbal equivalents of mathematical calculations and fall under the “mathematical concept” grouping of ideas.
The limitations reciting a predicting method of transcription factor binding sites based on weighted multi-granularity scanning, comprising the following steps: carrying out data augmentation on an initial data set D= {D1,D2,…,Dn}, Di= {xk,yk} (1 ≤ k ≤ n) of transcription factor binding sites, where xk represents a DNA sequence fragment and yk represents whether the DNA sequence fragment is a binding site, taking the value as a binding site or a non-binding site, calculating an inverse sequence, a complementary sequence and a complementary inverse sequence of each piece of data, expanding the number of data sets to four times the original number to obtain a data sets D*=D1,D2,…,D4n,Di,=xk',yk' (1≤k'≤4n), and randomly mixing positive and negative samples in the data set D*; carrying out one-hot coding on each piece of DNA sequence data in the data set D* with the formula {1000,A; 0001,T; 0100,C; 0010,G} to obtain a feature vector F1, and then combining with multi-base feature coding for feature representation to obtain a feature vector F2, splicing the feature vectors F1 and F2 to obtain a combined feature representation F, and encoding the result category with the formula: { 1, binding site; 0, non-binding site}; dividing the data set D* after feature representation in step (2) according to the ratio Q:R of the number of samples in the training set to the number of samples in the test set to obtain a training set Dtrain and a test set Dtest, where Q is the number of samples in the training set in the data set D* and R is the number of samples in the test set in the data set D*; Q has a value in the range of 2-5, and R has a value of 1; carrying out weighted multi-granularity scanning on the feature F of each sample in the training set Dtrain, in which the specific steps are as follows: using a sliding window with a length of μ to slide with a step length of L on the feature vector F and the weight vector W with a length of d, respectively, and extracting the feature vectors in the window separately to obtain fu and wu with a length of μ, where u is the number of times the sliding window slides, and u has a value in the range of 1≤u≤d-μ+1; according to the formula
PNG
media_image5.png
31
94
media_image5.png
Greyscale
; outputting a result of 1 or 0; in which 1 indicates that the DNA sequence is a transcription factor binding site, and 0 indicates that the DNA sequence is a non-transcription factor binding site; predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in the multi-base feature coding method, the length L of the feature column is obtained according to the formula L=4m, where m is the length of the base in the multi-base, m has a value of 3, bases A, T, C and G form a sequence set C with a length of 3bp: {'AAA', 'AAT', 'AAG', 'AAC', 'ATA', 'ATT', 'ATG', 'ATC', 'AGA', 'AGT', 'AGG', 'AGC', 'ACA', 'ACT', 'ACG', 'ACC', 'TAA', 'TAT', 'TAG', 'TAC', 'TTA', 'TTT', 'TTG', 'TTC', 'TGA', 'TGT', 'TGG', 'TGC', 'TCA', 'TCT', 'TCG', 'TCC', 'GAA', 'GAT', 'GAG', 'GAC', 'GTA', 'GTT', 'GTG', 'GTC', 'GGA', 'GGT', 'GGG', 'GGC', 'GCA', 'GCT', 'GCG', 'GCC', 'CAA', 'CAT', 'CAG', 'CAC', 'CTA', 'CTT', 'CTG', 'CTC', 'CGA', 'CGT', 'CGG', 'CGC', 'CCA', 'CCT', 'CCG', 'CCC'}, each element in set C is set as a feature column, there are 64 feature columns in total, and its element is the feature name of the feature column; the calculation method of the feature vector F2 is as follows: from the beginning end of the DNA sequence sample, a window with a step size of 1 and a length of 3bp slides on the DNA sequence sample and extracts features, and the feature column corresponding to the sequence in the window has a value of 1 until the end of the DNA sequence sample, that is, the length of the feature vector F2 is 64 falls under the mental process grouping of ideas. Evaluating a result and classifying it as binding or non-binding, predicting method of transcription factor binding sites based on weighted multi-granularity scanning, calculating variations of sequences (e.g. inverse, complementary, inverse complementary), carrying out one-hot feature encoding, carrying out multi-granularity weighted scanning via a sliding window can be practically performed in the human mind or with a pen and paper and therefore, is a mental process. Although it may take a long amount of time, these limitations still amount to a mental process. The courts do not distinguish between mental processes that are performed entirely in the human mind and mental processes that require a human to use a physical aid (See, e.g., Benson, 409 U.S. at 67, 65, 175 USPQ at 674-75, 674 and Synopsys, Inc. v. Mentor Graphics Corp., 839 F.3d 1138, 1139, 120 USPQ2d 1473, 1474 (Fed. Cir. 2016).
The limitations of the predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in step (4), T has a value of 462, and the maximum depth of the tree is 11 merely serve to further limit the random forest construction model which is a verbal equivalent of a mathematical process.
The limitations reciting in step (3), Q has a value of 4, and R has a value of 1 and in step (5), μ has a value of 50, and L has a value of 1 merely serve to further limit the recited mental process. As such, claims 1-5 recite an abstract idea.
Claims found to recite a judicial exception under Step 2A, Prong 1 are then further analyzed to determine if the claims as a whole integrate the recited judicial exception into a practical application or not (Step 2A, Prong 2). This judicial exception is not integrated into a practical application because the claims do not recite additional elements that reflects an improvement to technology or applies or uses the recited judicial exception in some other meaningful way. Specifically, the claims recite no additional elements. As such, claims 1-5 do not integrate into a practical application.
Claims found to be directed to a judicial exception are then further evaluated to determine if the claims recite an inventive concept that provides significantly more than the judicial exception itself (Step 2B). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the instant claims recite no additional elements.
There are no additional elements that comprise an inventive concept when considered individually or as an ordered combination that transforms the claimed judicial exception into a patent-eligible application of the judicial exception. Therefore, the claims do not amount to significantly more than the judicial exception itself (Step 2B: No). As such, claims 1-5 are not patent eligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
The present rejection(s) reference specific passages from cited prior art. However,
Applicant is advised that the rejections are based on the entirety of each cited prior art. That is,
each cited prior art reference “must be considered in its entirety”. (See MPEP 2141.02(VI))
Therefore, Applicant is advised to review all portions of the cited prior art if traversing a
rejection based on the cited prior art.
Claim(s) 1 and 3-5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhou et al. ((2017). Deep Forest: Towards An Alternative to Deep Neural Networks. arXiv.Org. https://arxiv.org/abs/1702.08835v2) in view of Cao et al. (Simple tricks of convolutional neural network architectures improve DNA–protein binding prediction, Bioinformatics, Volume 35, Issue 11, June 2019, Pages 1837–1843,), Zhou et al. ("Prediction of TF-Binding Site by Inclusion of Higher Order Position Dependencies," in IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 17, no. 4, pp. 1383-1393, 1 July-Aug. 2020, doi: 10.1109/TCBB.2019.2892124.), Louppe et al. (Understanding variable importances in forests of randomized trees. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 1 (NIPS'13), Vol. 1. Curran Associates Inc., Red Hook, NY, USA, 431–439), Salisman et al. (Weight Normalization: A simple reparameterization to accelerate training of deep neural networks. https://arxiv.org/abs/1602.07868) and Wang et al. (A New Method Combining DNA Shape Features to Improve the Prediction Accuracy of Transcription Factor Binding Sites. In: Huang, DS., Jo, KH. (eds) Intelligent Computing Theories and Application. ICIC 2020).
Regarding claim 1,
Zhou (2017) teaches a method of predicting transcription factor binding sites using deep learning methods comprising a complete prediction model containing five basic components: a validation datasets, an effective feature extraction procedure, an efficient predicting algorithm, a set of fair evaluation criteria and a comparison step (see “Methods and Materials”; page 1384). Although Zhou (2017) teaches the general pipeline of the disclosed invention, he does not specifically teach the data preparation steps, feature extraction, and inputting the features into a prediction/evaluation forest model.
Regarding step (1),
Cao et al. teaches the expansion of a dataset by using the reverse complement sequence as another training sample to double the size of the training set which is known as the reverse complement (RC) augmentation trick to aid in DNA-protein binding prediction pipelines (see “2.4 The reverse complement augmentation trick”; page 1839). Although Cao does not teach increasing the data set four times using the reverse sequence and complement sequence, it is would be an inherent property to have both of these missing sequences as the reverse string of sequences can be derived from the reverse complement sequence and the complement sequence can be derived from the original sequence and be compiled into a dataset for a dataset expansion step which is a known strategy in bioinformatics pipelines according to Cao. Accordingly, expanding the dataset using the reverse complement sequence necessarily involve having the reverse sequence and complement sequence. It is elementary that the mere recitation of a newly discovered function or property, inherently possessed by things in the prior art, does not cause a claim drawn to distinguish of the prior art. Under the principles of inherency, if a prior art device, in its normal and usual operation, would necessarily perform the method claimed, then the method claimed will be considered to be anticipated by the prior art device. Additionally, where the Patent Office has reason to believe that a functional limitation asserted to be critical for establishing novelty in the claimed subject matter may, in fact, be an inherent characteristic of the prior art, it possesses the authority to require the applicant to prove that the subject matter shown to be in the prior art does not possess the characteristic relied on (see MPEP § 2112). Therefore, it would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Cao’s dataset expansion into Zhou (2017)’s transcription factor binding prediction pipeline in order to promote the using of deeper CNN models and improving prediction performance as recited by Cao. The dataset expansion incorporation could have been accomplished with reasonable expectation of success as the dataset expansion is directed to the same problem of binding prediction in the same field of endeavor.
Regarding step (2),
Zhou (2017) and Cao do not teach carrying out one-hot coding on each piece of DNA sequence data in the data set D* with the formula {1000,A; 0001,T; 0100,C; 0010,G} to obtain a feature vector F1, and then combining with multi-base feature coding for feature representation to obtain a feature vector F2, splicing the feature vectors F1 and F2 to obtain a combined feature representation F, and encoding the result category with the formula: {1, binding site; 0, non-binding site}. Wang teaches the composition of a matrix wherein four bases are expressed as A = [1,0,0,0], T = [0,1,0,0], C = [0,0,1,0] and G = [0,0,0,1] by one-hot encoding to be used in subsequent feature extraction and machine learning pipelines (see “2.1 Data Processing” on page 81). Although the classification of nucleotides differs between Wang and Zhou (2017) has modified, it does not change the subsequent binary classification for feature prediction, therefore, the selection of specific numerical values in the claimed invention is merely a design choice. In addition, the limitations of combining feature representations to arrive at a binary classification is routine optimization as one of ordinary skill of the art would have motivated to combine multiple desired feature representations as neural networks will mine feature information from different types of data to find correlations between different features as evidenced by Wang (see second paragraph of “Introduction” on page 80). Of note, Zhou (2020) already discloses the concatenation of raw input features and subsequent feature re-representation (see page 3; Fig. 3) with the only difference being what feature representations are desired from the user and modifying the existing pipeline to extract features desired by the user would amount to routine optimization. Therefore, it would have been obvious to one of ordinary skill in the art to incorporate Wang’s feature extraction step into Zhou (2017) as modified’s pipeline in order to improve accuracy over existing models (see “Abstract” on page 79). The incorporation would have been accomplished with reasonable expectation of success as Wang already recites the combination of feature vectors into a machine learning pipeline (see first paragraph on “2.3 Output Stage” on page 82).
Regarding step (3),
Although Zhou (2017) as modified does not explicitly teach dividing the data set D* after feature representation in step (2) according to the ratio Q:R of the number of samples in the training set to the number of samples in the test set to obtain a training set Dtrain and a test set Dtest, where Q is the number of samples in the training set in the data set D* and R is the number of samples in the test set in the data set D*; Q has a value in the range of 2-5, and R has a value of 1 (explicitly recited as a growing set and an estimating set in an 80/20 split in “3.1 Configuration” on page 3). Of note, an estimating set is described as Zhou (2017) to be used to evaluate the performance of the model, which is functionally the same as a test set which is also used the measure performance. Furthermore, if Q had a value of 5 and R had a value of 1, the resulting ratio is the same, therefore, the limitations of the claim are met.
Regarding step (4),
Zhou (2017) as modified does not explicitly teach using T decision trees to calculate a weight vector W=(W1、W2…Wi…Wd) (1 ≤I ≤d) for the training set Dtrain, where d is the length of the feature, and the specific calculation formula is as follows:
PNG
media_image1.png
81
155
media_image1.png
Greyscale
where d is the total number of features, Scorei is the importance score of the i-th column of features in the weight vector W, and the specific calculation formula is as follows:
PNG
media_image2.png
84
209
media_image2.png
Greyscale
where Scorenode(t) is the importance score of the t-th decision tree node, and the specific calculation formula is as follows:
PNG
media_image3.png
47
286
media_image3.png
Greyscale
where Gnode,0 and Gnode,1 represent the Gini index of the nodes belonging to category 0 under the node branch and the Gini index of the nodes belonging to category 1 under the node branch, respectively; Gnode is the Gini index of each node, and the specific formula is as follows:
PNG
media_image4.png
68
275
media_image4.png
Greyscale
where N is the number of samples in the training set Dtrain, Nnode,0 is the number of nodes belonging to category 0, and Nnode,1 is the number of nodes belonging to category 1. Louppe discloses an ensemble of T-independently trained decision trees with classification driven nodes split based on the Gini Index (see “2.1 Single classification and regression trees and random forests”). Namely, the importance score computation of the t-th decision tree node is disclosed as the impurity measure i(t) (see Equation 1 under “2.1 Single classification and regression trees and random forests”) and the importance score of the i-th column of features in the weight vector W, is disclosed in description of the Gini Index as an impurity function (see “2.2. Variable importances” for Gini importance disclosure and Equation 4 as the explicit disclosure of importance score). Therefore, it would have been obvious to one of ordinary skill in the art to incorporate Louppe’s binary classification tree model into Zhou (2017) as modified’s existing pipeline in order to identify which predictor variables are the most important to make predictions to understand the underlying process (see “Motivation” on page 1). This would have been accomplished with reasonable expectation of success as Zhou (2017) as modified’s existing pipeline already includes T decision trees in the same field of endeavor.
Louppe is silent as to the weight vector computation. Salisman teaches a weight normalization step as an integral part of the artificial neural network pipeline (see page 2; Equation 2). Additionally, it is noted that this is merely a normalization step that is known in the art as evidenced by Salimans and although the equations may use different variables, the objective for each of the equations remains the same (e.g. normalization and deriving the Gini importance); thus, one of ordinary skill in the art would be able to make the appropriate modifications on the base equation. Therefore, it would have been obvious to one of ordinary skill in the art to modify Zhou (2017) as modified’s existing pipeline with Salisman’s normalization equation in order to improve the convergence of the optimization procedure (see under “2. Weight Normalization” on page 2). The modification would be accomplished with reasonable expectation of success as both operate in the same field of endeavor using deep neural networks.
Regarding step (5),
Zhou (2017) as modified does not explicitly teach carrying out weighted multi-granularity scanning on the feature F of each sample in the training set Dtrain, in which the specific steps are as follows: using a sliding window with a length of μ to slide with a step length of L on the feature vector F and the weight vector W with a length of d, respectively, and extracting the feature vectors in the window separately to obtain fu and wu with a length of μ, where u is the number of times the sliding window slides, and u has a value in the range of 1≤u≤d-μ+1; according to the formula
PNG
media_image5.png
31
94
media_image5.png
Greyscale
, calculating the features of the weighted multi-granularity scanning, where wuT is the transposition of vector wu; sending the feature Fu' into a completely random forest A and an ordinary random forest B to obtain F’Au and F'Bu, respectively, and finally, splicing F’Au and F'Bu to obtain feature F*, however, he does teach a multi-grained scanning procedure with different sizes of sliding windows for data examination (pages 2-3 in “2.2 Multi-Grained Scanning” and “2.3 Overall Procedure and Hyper-Parameters”). Although the Zhou (2017) as modified does not teach the explicit formula, Zhou (2020) teaches the process of feature extraction via a sliding window as described is functionally equivalent to the claimed invention (see Fig. 4 wherein sliding windows will generate feature vectors and the transformed feature vectors are used to generate subsequent forests) and the recitation of the formula is merely the mathematical equivalent of the disclosed process.
Regarding step (6),
Zhou (2020) as modified teaches inputting F* into the cascade forest, carrying out the model training to obtain a classification prediction model of transcription factor binding sites, inputting the test set Dtest into the classification prediction model (see second paragraph in “2.3 Overall Procedure and Hyper-Parameters” and Fig. 4) and outputting a result of 1 or 0; in which 1 indicates that the DNA sequence is a transcription factor binding site, and 0 indicates that the DNA sequence is a non-transcription factor binding site (Wang: see first paragraph of “Results” on page 2 discussing binary classification in relation to binding site prediction).
Regarding claim 3,
The predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in step (3), Q has a value of 4, and R has a value of 1. Of note, the ratio amounts to a 75/25 split rather than the Zhou (2020) as modified’s disclosed 80/20 (explicitly recited as a growing set and an estimating set in an 80/20 split in “3.1 Configuration” on page 3). However, in MPEP 2144.05, the courts have ruled that in the case where the claimed ranges "overlap or lie inside ranges disclosed by the prior art" a prima facie case of obviousness exists (see In re Wertheim, 541 F.2d 257, 191 USPQ 90 (CCPA 1976); In re Woodruff, 919 F.2d 1575, 16 USPQ2d 1934 (Fed. Cir. 1990)). Therefore, the limitations of the claim are sufficiently met by Zhou (2020) as modified.
Regarding claim 4,
Although Zhou (2020) as modified does not explicitly recite the predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in step (4), T has a value of 462, and the maximum depth of the tree is 11, he does disclose that the complete random forest in the multi-grained scanning step contains 30 trees in each forest with a tree growth of less than 20 (see Table 1 on page 4). Similarly, to above, in MPEP 2144.05, the courts have ruled that in the case where the claimed ranges "overlap or lie inside ranges disclosed by the prior art" a prima facie case of obviousness exists (see In re Wertheim, 541 F.2d 257, 191 USPQ 90 (CCPA 1976); In re Woodruff, 919 F.2d 1575, 16 USPQ2d 1934 (Fed. Cir. 1990)). Therefore, the limitations of the claim are sufficiently met by Zhou (2020) as modified. Concerning the value of the 462 trees, the courts have held In re Harza, 274 F.2d 669, 124 USPQ 378 (CCPA 1960), that mere duplication of parts has no patentable significance unless a new and unexpected result is produced (see MPEP 2144.04). Additionally, it is explicitly shown in Zhou (2020)’s disclosure that increasing the number of trees while lowering the growth is reasonable via the cascade forest section of the gcForest algorithm (see Table 1 on page 4), and one of ordinary skill in the art would recognize that the modifications could be made to arrive at the claimed invention.
Regarding claim 5,
Zhou (2020) as modified explicitly recites the predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in step (5), μ has a value of 50, and L has a value of 1 (adjustable window sliding sizes as seen in Figure 4 of page 4). The Examiner notes that the iteration of times a window slides is directly related to the size of the window and the size of the matrix (which has been established as being adjustable via the 4*L disclosure above), therefore, if the sliding window is adjustable and the size of the matrix is adjustable, the value of μ (which is the number of times the window slides) must also be inherently adjustable. Therefore, Zhou (2020) as modified meets the limitations of this claim. Accordingly, the number of times a window slides necessarily involve the size of the sliding window and the size of the matrix. It is elementary that the mere recitation of a newly discovered function or property, inherently possessed by things in the prior art, does not cause a claim drawn to distinguish of the prior art. Under the principles of inherency, if a prior art device, in its normal and usual operation, would necessarily perform the method claimed, then the method claimed will be considered to be anticipated by the prior art device. Additionally, where the Patent Office has
reason to believe that a functional limitation asserted to be critical for establishing novelty in the
claimed subject matter may, in fact, be an inherent characteristic of the prior art, it possesses
the authority to require the applicant to prove that the subject matter shown to be in the prior art
does not possess the characteristic relied on (see MPEP § 2112).
Claims 2 is rejected under 35 U.S.C. 103 as being unpatentable over Zhou et al. ((2017). Deep Forest: Towards An Alternative to Deep Neural Networks. arXiv.Org. https://arxiv.org/abs/1702.08835v2), as applied in claim 1, in view of Cao et al. (Simple tricks of convolutional neural network architectures improve DNA–protein binding prediction, Bioinformatics, Volume 35, Issue 11, June 2019, Pages 1837–1843,), Zhou et al. ("Prediction of TF-Binding Site by Inclusion of Higher Order Position Dependencies," in IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 17, no. 4, pp. 1383-1393, 1 July-Aug. 2020, doi: 10.1109/TCBB.2019.2892124.), Louppe et al. (Understanding variable importances in forests of randomized trees. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 1 (NIPS'13), Vol. 1. Curran Associates Inc., Red Hook, NY, USA, 431–439), Salisman et al. (Weight Normalization: A simple reparameterization to accelerate training of deep neural networks. https://arxiv.org/abs/1602.07868) and Wang et al. (A New Method Combining DNA Shape Features to Improve the Prediction Accuracy of Transcription Factor Binding Sites. In: Huang, DS., Jo, KH. (eds) Intelligent Computing Theories and Application. ICIC 2020) further in view of Zhang et al. (CRIP: predicting circRNA-RBP-binding sites using a codon-based encoding and hybrid deep neural networks. RNA. 2019 Dec;25(12):1604-1615).
Regarding claim 2,
The combination of Zhou (2020) and Wang teach the predicting method of transcription factor binding sites based on weighted multi-granularity scanning according to claim 1, wherein in the multi-base feature coding method, the length L of the feature column is obtained according to the formula L=4m, where m is the length of the base in the multi-base, wherein each element in set C is set as a feature column, there are 64 feature columns in total, and its element is the feature name of the feature column; the calculation method of the feature vector F2 is as follows: from the beginning end of the DNA sequence sample, a window with a step size of 1 and a length of 3bp slides on the DNA sequence sample and extracts features, and the feature column corresponding to the sequence in the window has a value of 1 until the end of the DNA sequence sample, that is, the length of the feature vector F2 is 64. In particular, Zhou (2020) teaches the pipeline including the sliding window for feature extraction which generates forests to eventually feeds into the cascade forest model (see Fig. 4 wherein sliding windows will generate feature vectors and the transformed feature vectors are used to generate subsequent forests) while Wang teaches the composition of a 4*L matrix where L is the length of the sequence (35bp in the disclosure) for input into a subsequent model (see “2.1 Data Processing” on page 81).
However, they are both silent as to the limitation wherein m has a value of 3, bases A, T, C and G form a sequence set C with a length of 3bp: {'AAA', 'AAT', 'AAG', 'AAC', 'ATA', 'ATT', 'ATG', 'ATC', 'AGA', 'AGT', 'AGG', 'AGC', 'ACA', 'ACT', 'ACG', 'ACC', 'TAA', 'TAT', 'TAG', 'TAC', 'TTA', 'TTT', 'TTG', 'TTC', 'TGA', 'TGT', 'TGG', 'TGC', 'TCA', 'TCT', 'TCG', 'TCC', 'GAA', 'GAT', 'GAG', 'GAC', 'GTA', 'GTT', 'GTG', 'GTC', 'GGA', 'GGT', 'GGG', 'GGC', 'GCA', 'GCT', 'GCG', 'GCC', 'CAA', 'CAT', 'CAG', 'CAC', 'CTA', 'CTT', 'CTG', 'CTC', 'CGA', 'CGT', 'CGG', 'CGC', 'CCA', 'CCT', 'CCG', 'CCC'}. Zhang teaches the extraction of 3-mers from RNA sequences using a sliding window with a step-size of 1 (see 4th paragraph in “Investigation on feature encoding”). It would have been obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Zhang’s 3-mer extraction into Zhou (2017) as modified by Cao, Wang, Louppe, Zhou (2020), and Salisman (as recited in claim 1)’s pipeline in order to map similarly to the translation of codons which will reduce the dimensionality of a classic k-mer method and group them according to common biological properties (see 4th paragraph in “Investigation on feature encoding”). This could have been accomplished with reasonable expectation of success as the sliding window for 35bps already exists in Zhou (2020) as modified’s existing pipeline and the incorporation of Zhang would amount to a simple substitution in the same field of endeavor.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Wang et al. (Specific and intrinsic sequence patterns extracted by deep learning from intra-protein binding and non-binding peptide fragments. Sci Rep 7, 14916 (2017)) discloses sequence patterns extracted from deep learning models. Scikit-learn (scikit-learn 1.9.0 documentation. Sci-kit Learn. (2016, September 28)) discloses common techniques in bioinformatics and machine learning. Rosenthal (Time series for scikit-learn people (part I): Where’s the X matrix? | Ethan Rosenthal)teaches that scikit-learn is standard within the art.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PETER NGUYEN whose telephone number is (571)272-0127. The examiner can normally be reached Monday - Friday 7:30am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Olivia M. Wise can be reached at (571) 272-2249. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/P.N./Examiner, Art Unit 1685
/OLIVIA M. WISE/Supervisory Patent Examiner, Art Unit 1685