DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Notice to Applicant
2. This communication is in response to the communication filed 10/31/2023. Claims 1-20 are currently pending.
Claim Rejections - 35 USC § 101
3. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
3.1. Claims 1-18 are rejected under 35 U.S.C. § 101 because while the claims (1) are to a statutory category (i.e., process, machine, manufacture or composition of matter, the claims (2A1) recite an abstract idea (i.e., a law of nature, a natural phenomenon); (2A2) do not recite additional elements that integrate the abstract idea into a practical application; and (2B) are not directed to significantly more than the abstract idea itself.
In regard to (1), the claims are to a statutory category (i.e., statutory categories including a process, machine, manufacture or composition of matter). In particular, independent claims 1 and 10, and their respective dependent claims are directed, in part, to systems and methods for identifying novel pore-forming toxins (PFTs) based on sequence and structure data.
In regard to (2A1), the claims, as a whole, recite and are directed to an abstract idea because the claims include one or more limitations that correspond to an abstract idea including mental processes and mathematical concepts. For example, the claims are directed to an abstract idea because the claims, except for certain limitations (* identified below in bold), under the broadest reasonable interpretation, can be reasonably and practically performed in the human mind and/or with pen and paper using observation, evaluation, judgment and/or opinion. That is, other than reciting the certain additional elements, nothing in the claims precludes the limitations from being practically performed in the mind and/or with pen and paper. For example, a person is capable of reasonably and practically identify a plurality of proteins with known sequences and structures, determining a plurality of protein clusters based on pairwise structural similarity values of the plurality of proteins, identifying a respective group of proteins, generating a graphical model, calculating a segment interaction score for the generated structural segmentation of the input protein, classifying the input protein as a potential novel PFT, etc. using observation, evaluation, judgment and/or opinion.
Moreover, the claims are also abstract because the claims are directed to a mathematical concept. For example, independent claims 1 and 10 recite “calculating a segment interaction score.” This is a mathematical calculation and thus, and abstract idea under mathematical concepts.
CLAIM 1:
A method for identifying novel pore-forming toxins (PFTs) based on sequence and structure data, comprising:
identifying, in a dataset comprising PFT information, a plurality of proteins with known sequences and structures;
determining a plurality of protein clusters based on pairwise structural similarity values of the plurality of proteins;
for each respective protein cluster of the plurality of protein clusters, identifying a respective group of proteins that have a pairwise sequence identity a) lower than a threshold pairwise sequence identity, b) above a threshold pairwise sequence identity, or c) within a predetermined pairwise sequence identity range;
generating a graphical model trained using sequence and structure data of proteins from each respective group of proteins, wherein the graphical model is configured to generate a structural segmentation of an input protein based on a sequence of the input protein;
calculating a segment interaction score for the generated structural segmentation of the input protein, wherein the segment interaction score compares the generated structural segmentation of the input protein with structures of the proteins from each respective group of proteins; and
in response to determining that the segment interaction score is greater than a threshold segment interaction score, classifying the input protein as a potential novel PFT.
CLAIM 2
The method of claim 1, further comprising:
determining whether the sequence of the input protein is classified as a PFT sequence using a machine learning model configured to classify sequences as a PFT sequence or a non-PFT sequence; and
in response to determining that the sequence is classified as a PFT sequence, identifying the input protein as a novel PFT.
CLAIM 3
The method of claim 2, further comprising: receiving confirmation that the input protein is not a novel PFT; and re-training the machine learning model such that the sequence of the input protein is identified as a non-PFT sequence.
CLAIM 4
The method of claim 1, further comprising generating, for output on a computing device, an indication that the input protein is classified as a potential novel PFT and that the input protein shares a functionality of a particular protein cluster from the plurality of protein clusters.
CLAIM 5
The method of claim 1, wherein determining the plurality of protein clusters further comprises:
mapping a structural representation of each protein from the plurality of proteins from a high-dimensional space to a two-dimensional space that preserves structural correlations among the plurality of proteins; and
executing a clustering algorithm on the two-dimensional space to determine the plurality of protein clusters.
CLAIM 6
The method of claim 5, wherein the clustering algorithm is a K-means clustering algorithm.
CLAIM 7
The method of claim 1, wherein generating the graphical model further comprises: aligning structures of the proteins from each respective group of proteins to identify common structural regions using an iterative pairwise alignment algorithm; and
identifying consensus and non-consensus secondary structure segments in the aligned structures, wherein training the graphical model comprises maximizing a probability of the identified consensus and non-consensus secondary structure segments in the structural segmentation of the input protein.
CLAIM 8
The method of claim 1, wherein the graphical model is a semi-Markov conditional fields (semi-CRFs) model.
CLAIM 9
The method of claim 1, wherein proteins in a respective group of proteins share functionality and have a low sequence identity, optionally less than 60, 50, 40, 30, 20, or 10% full length sequence identity.
CLAIM 10
A system for identifying novel pore-forming toxins (PFTs) based on sequence and structure data, comprising:
a processor, and
memory having stored therein instructions that when executed by the processor cause the processor to:
identify, in a dataset comprising PFT information, a plurality of proteins with known sequences and structures;
determine a plurality of protein clusters based on pairwise structural similarity values of the plurality of proteins;
for each respective protein cluster of the plurality of protein clusters, identify a respective group of proteins that have a pairwise sequence identity a) lower than a threshold pairwise sequence identity, b) above a threshold pairwise sequence identity, or c) within a predetermined pairwise sequence identity range;
generate a graphical model trained using sequence and structure data of proteins from each respective group of proteins, wherein the graphical model is configured to generate a structural segmentation of an input protein based on a sequence of the input protein;
calculate a segment interaction score for the generated structural segmentation of the input protein, wherein the segment interaction score compares the generated structural segmentation of the input protein with structures of the proteins from each respective group of proteins; and
in response to determining that the segment interaction score is greater than a threshold segment interaction score, classify the input protein as a potential novel PFT.
CLAIM 11
The system of claim 10, wherein the memory further comprises instructions for: determining whether the sequence of the input protein is classified as a PFT sequence using a machine learning model configured to classify sequences as a PFT sequence or a non- PFT sequence; and
in response to determining that the sequence is classified as a PFT sequence, identifying the input protein as a novel PFT.
CLAIM 12
The system of claim 11, wherein the memory further comprises instructions for: receiving confirmation that the input protein is not a novel PFT; and re-training the machine learning model such that the sequence of the input protein is identified as a non-PFT sequence.
CLAIM 13
The system of claim 10, wherein the memory further comprises instructions for: generating, for output on a computing device, an indication that the input protein is classified as a potential novel PFT and that the input protein shares a functionality of a particular protein cluster from the plurality of protein clusters.
CLAIM 14
The system of claim 10, wherein determining the plurality of protein clusters further comprises:
mapping a structural representation of each protein from the plurality of proteins from a high-dimensional space to a two-dimensional space that preserves structural correlations among the plurality of proteins; and
executing a clustering algorithm on the two-dimensional space to determine the plurality of protein clusters.
CLAIM 15
The system of claim 14, wherein the clustering algorithm is a K-means clustering algorithm.
CLAIM 16
The system of claim 10, wherein generating the graphical model further comprises: aligning structures of the proteins from each respective group of proteins to identify common structural regions using an iterative pairwise alignment algorithm; and
identifying consensus and non-consensus secondary structure segments in the aligned structures, wherein training the graphical model comprises maximizing a probability of the identified consensus and non-consensus secondary structure segments in the structural segmentation of the input protein.
CLAIM 17
The system of claim 10, wherein the graphical model is a semi-Markov conditional fields (semi-CRFs) model.
CLAIM 18
The system of claim 10, wherein proteins in a respective group of proteins share functionality and have low sequence identity.
* The limitations that are in bold are considered “additional elements” that are further analyzed below in subsequent steps of the 101 analysis. The limitations that are not in bold are abstract and/or can be reasonably and practically performed in the human mind and/or with pen paper.
In regard to (2A2), the claims do not recite additional elements that integrate the abstract idea into a practical application. The additional elements in the claims (i.e., * identified above in bold) do not integrate the abstract idea into a practical application because the additional elements merely add insignificant extra-solution activity to the abstract idea; merely link the use of the judicial exception to a particular technological environment or field of use; and/or simply append technologies and functions, specified at a high level of generality, to the abstract idea (i.e., the additional elements do not amount to more than a recitation of the words “apply it” (or an equivalent) or are more than mere instructions to implement an abstract idea or other exception on a computer).
Here, the additional elements (e.g., graphical model trained, machine learning model, computing device, processor, memory, etc.) are recited at a high-level of generality such that it amounts to no more than mere instructions to apply the abstract idea using generic computer technologies. Moreover, the claims recite “configured to generate,” and “cause the processor to”, etc. devoid of any meaningful technological improvement details and thus, further evidence the additional elements are merely being used to leverage generic technologies to automate what otherwise could be done manually. Accordingly, the additional elements do not integrate the abstract idea into a practical application because they do not impose any meaningful limits on practicing the abstract idea.
Furthermore, the additional elements do not recite improvements to the functioning of a computer, or to any other technology or technical field—the additional elements merely recite general purpose computer technology; the additional elements do not recite applying or using a judicial exception to effect a particular treatment or prophylaxis for disease or medical condition—there is no actual administration of a particular treatment; the additional elements do not recite applying the judicial exception with, or by use of, a particular machine—the additional elements merely recite general purpose computer technology; the additional elements do not recite limitations effecting a transformation or reduction of a particular article to a different state or thing—the additional elements do not recite transformation such as a rubber mold process; the additional elements do not recite applying or using the judicial exception in some other meaningful way beyond generally linking the use of the judicial exception to a particular technological environment—the additional elements merely leverage general purpose computer technology to link the abstract idea to a technological environment.
In regard to (2B), the claims, individually, as a whole and in combination with one another, do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional elements or combination of elements in the claims, other than the abstract idea per se, amount to no more than a recitation of (A) a generic computer structure(s) that serves to perform computer functions that serve to merely link the abstract idea to a particular technological environment (i.e., computers); and/or (B) functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Here, as discussed above with respect to integration of the abstract idea into a practical application, the additional elements fall under (A), that is, the additional elements amount to no more than mere instructions to apply the exception using generic computer technologies. Mere instructions to apply an exception using generic computer technologies cannot provide an inventive concept.
Moreover, paragraphs [0047]-[0048] of applicant's specification (US 2024/0395356) recites that the system/method is implemented using computing devices, such as a desktop computer, a notebook computer, a laptop computer, a mobile computing device, a smart phone, a tablet computer, a server, a mainframe, an embedded device, and the like which are well-known general purpose or generic-type computers and/or technologies. The use of generic computer components recited at a high level of generality to process information through an unspecified processor/computer does not impose any meaningful limit on the computer implementation of the abstract idea. Thus, taken alone, the additional elements do not amount to significantly more than the above-identified judicial exception (the abstract idea). Looking at the limitations as an ordered combination adds nothing that is not already present when looking at the elements taken individually. There is no indication that the combination of elements improves the functioning of a computer or improves any other technology. Their collective functions merely provide conventional computer implementation.
Furthermore, the additional elements fall under (B) because the additional elements are merely well-known general purpose computers, components and/or technologies that receive, transmit, store, display, generate and otherwise process information which are akin to functions that courts consider well-understood, routine, and conventional activities previously known to the pertinent industry, such as, performing repetitive calculations; receiving or transmitting data over a network; electronic recordkeeping; retrieving and storing information in memory; and sorting information (See, for example, MPEP § 2106). For example, identifying proteins, groups of proteins are a kin to sorting information; calculating a segment interaction score is akin to performing repetitive calculations.
Therefore, the claims are not patent-eligible under 35 U.S.C. § 101.
Conclusion
4. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Tomaszewski whose telephone number is (313)446-4863. The examiner can normally be reached M-F 5:30 am - 2:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Peter H Choi can be reached at (469) 295-9171. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL TOMASZEWSKI/Primary Examiner, Art Unit 3681