Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Claims 1-20 are pending.
Claims 1-20 are rejected.
Claims 1 and 16 are objected to.
Priority
The instant application, filed on 31 October 2022, is a CON of a PCT filed on 24 January 2022. The instant application claims priority to a foreign application filed on 28 January 2021. As such, claims 1-20 have an effective filing date of 28 January 2021.
Information Disclosure Statement
The IDS filed on 31 October 2022 was considered by the examiner.
Drawings
The drawings filed on 31 October 2022 are accepted.
Specification
The use of the term BLUETOOTH ™, which is a trade name or a mark used in commerce, has been noted in this application. The term should be accompanied by the generic terminology; furthermore the term should be capitalized wherever it appears or, where appropriate, include a proper symbol indicating use in commerce such as ™, SM , or ® following the term.
Although the use of trade names and marks used in commerce (i.e., trademarks, service marks, certification marks, and collective marks) are permissible in patent applications, the proprietary nature of the marks should be respected and every effort made to prevent their use in any manner which might adversely affect their validity as commercial marks.
Claim Objections
Claims 1 and 16 objected to because of the following informalities:
Claim 1 states: “An artificial intelligence(A)-based…”
Claim 16 states: “An artificial intelligence (A()-based…”
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed inventions are directed to abstract ideas without significantly more.
Step 2A, Prong 1
In accordance with MPEP § 2106, claims found to recite statutory subject matter (Step
1: YES) are then analyzed to determine if the claims recite any concepts that equate to an
abstract idea, law of nature or natural phenomenon (Step 2A, Prong 1). In the instant application,
the claims recite the following limitations that equate to abstract ideas:
Claim 1 and dependent claims 2-15 and 19, claim 16 and dependent claims 17 and 18, and claim 20 recite:
Determining a plurality of candidate drug molecules
Performing activity prediction
Screening the plurality of drug molecules
Claim 2 and dependent claims 3 and 4 recite:
Screening compounds in a compound library
Pre-processing the plurality of screened compounds
Claim 3 recites:
Deduplicating the plurality of compounds
Claim 4 recites:
Removing an enantiomer from the plurality of filtered compounds
Claim 5 and dependent claims 6-10 recite:
Encoding a molecular structure
Encoding a protein structure
Fusing embedding features
Mapping the activity fusion feature
Claim 6 recites:
Determining a molecular graph
Performing image encoding
Claim 7 recites:
Determining a protein sequence based on protein structure
Performing text transformation on the protein sequence
Claim 8 recites:
Summing the embedding features
Concatenating the embedding features
Claim 9 recites:
Mapping the embedding features
Performing affine transformation
Claim 14 and dependent claim 15 recite:
Clustering the plurality of drug molecules
Selecting candidate drug molecules
Claim 15 recites:
Determining a candidate drug molecule with highest activity information
Performing weighted summation on activity information, molecular docking information, and a drug property
Ranking the candidate drug molecules
Selecting target drug molecules
The limitations for the listed claims are evaluations or judgements that can be made through mental observations or mathematical calculations which fall under the “mental processes” and “mathematical concepts” groupings of abstract ideas. Under the broadest reasonable interpretation, the abstract ideas recited in the claims are determined to cover performance either in the mind (calculations by hand or pen and paper) or by mathematical operation (calculations/algorithms). See MPEP § 2106.04(a)(2), subsection III. The courts do not distinguish between mental processes that are performed entirely in the human mind and mental processes that require a human to use a physical aid (e.g., pen and paper or a slide rule) to perform the claim limitation (see, e.g., Benson, 409 U.S. at 67, 65, 175 USPQ at 674-75, 674: noting that the claimed "conversion of [binary-coded decimal] numerals to pure binary numerals can be done mentally," i.e., "as a person would do it by head and hand."); Synopsys, Inc. V. Mentor Graphics Corp., 839 F.3d 1138, 1139, 120 USPQ2d 1473, 1474 (Fed. Cir. 2016): holding that claims to a mental process of "translating a functional description of a logic circuit into a hardware component description of the logic circuit" are directed to an abstract idea, because the claims "read on an individual performing the claimed steps mentally or with pencil and paper"). Nor do the courts distinguish between claims that recite mental processes performed by humans and claims that recite mental processes performed on a computer. As the Federal Circuit has explained, "[c]ourts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person's mind" (see Versata Dev. Group V. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015); Mortgage Grader, Inc. v. First Choice Loan Servs. Inc., 811 F.3d 1314, 1324, 117 USPQ2d 1693, 1699 (Fed. Cir. 2016): holding that computer-implemented method for "anonymous loan shopping" was an abstract idea because it could be "performed by
humans without a computer").
While claims 16-19 and claim 20 recite performing aspects of the methods with a “memory”, “processor”, and/or “non-transitory computer readable medium”, there are no additional limitations that indicate that the memory, processor, or non-transitory computer readable medium would require anything other than carrying out the recited mental process or mathematical concept in a generic computer environment. Merely reciting that a mental process is being performed in a generic computer environment does not preclude the steps from being performed practically in the human mind or with pen and paper as claimed. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation on generic computer components, then it falls into the “mental processes” grouping of abstract ideas. As such, claims 1-20 recite abstract ideas (Step 2A, Prong 1: YES).
Step 2A, Prong 2
Claims found to recite a judicial exception under Step 2A, Prong 1 are then further
analyzed to determine if the claims as a whole integrate the recited judicial exception into a
practical application or not (Step 2A, Prong 2). This judicial exception is not integrated into a
practical application because the claims do not recite an additional element that reflects an
improvement to technology or applies or uses the recited judicial exception in some other
meaningful way. Rather, the instant claims recite additional elements that amount to mere
instructions to implement the abstract idea or insignificant extra-solution activity. Specifically,
the claims recite the following additional elements:
Claim 1 and dependent claims 2-15 and 19, claim 16 and dependent claims 17 and 18, and claim 20 recite:
Preforming homology modeling
Performing molecular docking
Claim 3 recites:
Performing Lipinski’s Rule of Five-based screening on the compounds
Claim 4 recites:
Chemically filtering the plurality of screened compounds
Claims 10 recites:
Mapping the fusion activity feature to latent vector space
Performing nonlinear mapping on the latent vector
Claim 11 recites:
Performing similarity processing on a protein sequence
Performing structure optimization based on a three-dimensional structure of a protein
Claim 12 and dependent claim 13 recite:
Performing molecular docking simulations
Preprocessing drug molecules
Molecular docking scoring
Claim 13 recites:
Performing format transformation
Constructing a three-dimensional conformation of each candidate drug molecule
Determining hydrogen addible position of drug molecule
Adding a hydrogen atom to a hydrogen atom addible position of drug molecule
Claim 16 and dependent claims 17 and 18 recite:
At least one memory
At least one processor
Claim 19 recites:
At least one memory
At least one processor
Claim 20 recites:
A non-transitory computer readable medium.
The limitations for defining terms describe mental processes with additional elements. This judicial exception is not integrated into a practical application because these additional
elements do not add any meaningful limitations. The claims do not include additional elements
that are sufficient to amount to significantly more than the judicial exception because they only
describe more specificity to the types of variables. As such, these limitations equate to mere
instructions to implement the abstract ideas.
There are no limitations that indicate that the memory, processor, or non-transitory computer readable medium would require anything other than a generic computing system. As such, these limitations equate to mere instructions to implement the abstract ideas on a generic computer that the courts have stated do not render an abstract idea eligible in Alice Corp., 573 U.S. at 223, 110 USPQ2d at 1983. See also 573 U.S. at 224, 110 USPQ2d at 1984.
The above recited additional elements do not provide a practical application of the recited
judicial exception. As such, claims 1-20 are directed to an abstract idea (Step 2A, Prong 2:NO).
Step 2B
Claims found to be directed to a judicial exception are then further evaluated to determine if the claims recite an inventive concept that provides significantly more than the judicial exception itself (Step 2B). The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims recite additional elements that are known and commonly used techniques in the art and are mere instructions to apply the recited exception in a generic computing environment.
As discussed above, there are no additional limitations to indicate that the claimed
method requires more than routine use of technology. Additionally, there are no additional limitations to indicate that the program requires anything other than generic computer components in order to carry out the recited abstract ideas in the claims. Claims that amount to nothing more than an instruction to apply the abstract idea using a generic computer do not render an abstract idea eligible. Alice Corp., 573 U.S. at 223, 110 USPQ2d at 1983. See also 573 U.S. at 224, 110 USPQ2d at 1984. In addition, mere display of collected and analyzed information that could be performed by the human mind do not render an abstract idea eligible. See Electric Power Group v. Alstom, S.A., 830 F.3d 1350, 1353-54, 119 USPQ2d 1739, 1741-42 (Fed. Cir. 2016)
The additional elements do not comprise an inventive concept when considered individually or as an ordered combination that transforms the claimed judicial exception into a patent-eligible application of the judicial exception. Therefore, the claims do not amount to significantly more than the judicial exception itself (Step 2B: No). As such, claims 1-20 are not patent eligible.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 5, 7, 11-12, 14-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by CN112201313A. (Beijing Jingtai Technology Co.,Ltd,.EFD 15 September 2020) (Herein referred to as BJT.)
With respect to independent claims 1, 16, and 20, and dependent claims 2, 5, 7, 11-12, 14-15 and 17-19, of the instant application, BJT teaches “[a]n automated small molecule drug screening method, suitable for execution in a computing device, the method comprising the steps of: Structural and activity data of multiple molecules targeting the target were collected, and a first candidate molecule library targeting the target was constructed based on the structural and activity data.” (Claim 1 of BJT machine translation) BJT further teaches to “[g]enerate vector features corresponding to each structural data, and use the vector features as sample input, the activity value as sample output, and the corresponding activity data as sample label to train the first prediction model; Each molecule in the first candidate molecule library is input into the first prediction model, and several molecules with the highest output activity values are selected to form the second candidate molecule library.” (Claim 2 of BJT) BJT additionally teaches “[f]or target proteins whose three-dimensional structures have not been resolved, their three-dimensional structures can be obtained through homology modeling for subsequent screening and analysis.”([0040]) BJT discloses “[e]ach molecule in the second candidate molecule library was docked with the target site, and several molecules with excellent docking conformations were selected to form a third candidate molecule library; and Multiple molecules in the third candidate molecular library are clustered, and multiple molecules with excellent performance are selected from each cluster to form a fourth candidate molecular library.” (Claim 3of BJT) BJT further discloses “the small molecule drug screening method according to the present invention further includes the step of: docking each molecule in the second candidate molecule library with the target site, and selecting multiple molecules with excellent docking conformations to form a third candidate molecule library.” ([0010]) BJT also teaches “the small molecule drug screening method according to the present invention further includes the step of: outputting molecular information of each candidate molecule library, the molecular information including molecular structure data, activity data, docking conformation, docking score and clustering status.” ([0016]) In regard to claim 16 and dependent claims 17 and 18 of the instant application, BJT teaches the use of a memory and one or processors in claim 9. In regard to claim 20 of the instant application, BJT teaches the use of a computer readable storage medium in claim 10.
With respect to claim 2, BJT further teaches “the construction process of the first candidate molecular library also includes at least one of the following processes: substructure or similarity matching or chemical property-based filtering. Substructure matching uses the RDKit package to search a compound database based on the structure data of active molecules using a vectorization method. It determines whether molecules in the candidate molecule library contain the target substructure and outputs the matching compounds. Filtering based on chemical properties, such as filtering according to drug class rules, or selecting based on other chemical properties, can screen out molecules with superior properties.”([0044])
With respect to claim 5, BJT further teaches “[t]he structural data and active data are stored in at least one of the following: smiles file, sdf file, mol file, mol2 file, and csv file; The structural data are represented using chemical terminology, and the activity data include enzyme activity and/or cell activity.” (Claim 7) BJT additionally teaches “structural and activity data of multiple molecules targeting the target are collected, and a candidate molecule library targeting the target, namely the first candidate molecule library, is constructed based on the structural and activity data… structural and activity data of these molecules can be obtained from patents, articles, or existing databases (e.g., PubChem, ChEMBL, PDBbind, etc.) targeting specific molecules. These data are presented as data files, including but not limited to SMILES files, SDF files, MOOL files, MOOL2 files, CSV files, etc. In these types of files, the structure of small molecules is represented by Chemical Markup Language (CML), including but not limited to SMILES (Simplified Molecular Input Line Entry Specification) and Chemical Table file (CT file). Activity (property) data includes, but is not limited to, molecular enzyme activity and cellular activity, such as IC50, Ki, Kd, and other information. This type of file also includes molecular identification information such as molecular numbers.” ([0038], [0039]) BJT teaches “a vector feature corresponding to each structural data is generated, and the vector feature is used as the sample input, the predicted activity value is used as the sample output, and the corresponding activity data is used as the sample label to train the first prediction model of the small molecule. Generally, the first prediction model can predict the activity value based on the structural features of the molecule, which converts the structural data of the active molecule into vector features represented by numbers.” ([0045])
With respect to claim 7, BJT further teaches “structural and activity data of these molecules can be obtained from patents, articles, or existing databases (e.g., PubChem, ChEMBL, PDBbind, etc.) targeting specific molecules. These data are presented as data files, including but not limited to SMILES files, SDF files, MOOL files, MOOL2 files, CSV files, etc. In these types of files, the structure of small molecules is represented by Chemical Markup Language (CML), including but not limited to SMILES (Simplified Molecular Input Line Entry Specification) and Chemical Table file (CT file). Activity (property) data includes, but is not limited to, molecular enzyme activity and cellular activity, such as IC50, Ki, Kd, and other information. This type of file also includes molecular identification information such as molecular numbers. The activity data obtained in this invention are structural and activity data of active molecules targeting the target, which can be used to train a molecular library and model for that specific target, thereby improving the efficiency and specificity of subsequent molecular screening. Target sites can be enzyme proteins, G-protein coupled receptors (GPCRs), ion channels, nuclear receptors, structural proteins, carrier proteins, etc. In the drug screening process, the three-dimensional structure of the target protein and the binding mode of the protein to small molecules are very important. Therefore, in step S210, the protein structure of the target site can also be obtained at the same time. The protein structure information mainly comes from the PDB (protein data bank) database and is a PDB format file. A PDB file is a standard file format that contains information such as the coordinates of atoms.” ([0039], [0040])
With respect to claim 11, BJT further teaches “[f]or target proteins whose three-dimensional structures have not been resolved, their three-dimensional structures can be obtained through homology modeling for subsequent screening and analysis.” ([0040]) BJT additionally teaches “the construction process of the first candidate molecular library also includes at least one of the following processes: substructure or similarity matching or chemical property-based filtering. Substructure matching uses the RDKit package to search a compound database based on the structure data of active molecules using a vectorization method. It determines whether
molecules in the candidate molecule library contain the target substructure and outputs the matching compounds.” ([0044]) BJT discloses “The computing device performs automated feature engineering processing and hyperparameter-optimized machine learning model training according to the data type, outputs several high-performing models for voting and scoring, and selects the best ensemble model based on the voting results.” ([0048])
With respect to claim 12, BJT further teaches “[m]olecular docking involves docking an active molecule with a protein pocket. After docking, multiple conformations can be generated, and the optimal conformation is automatically selected to calculate the affinity or binding activity between the small molecule and the protein, which is then used as a docking score. Before molecular docking, elements such as water molecules, ions, metals, ligands, and cofactors of the target protein crystal can be removed to facilitate clearer molecular docking. Molecular docking can be achieved by calling molecular docking software, or the execution logic of the molecular docking software can be modularly deployed in a computing device. In this way, as long as the second candidate molecular library is output on the page, the docking scoring of each molecule in the molecular library can be automatically continued, and multiple molecules with excellent docking conformations can be selected to form a third candidate molecular library. Among them, excellent docking conformation can mean a high docking score, such as ranking in the top 5% of docking scores, but it is not limited to this.” ([0060]) BJT additionally teaches “excellent performance can be reflected in a high activity value in the first
prediction model, a high docking score, or a high ranking in other physicochemical properties.” ([0066])
With respect to claim 14, BJT further teaches “[m]ultiple molecules in the third candidate molecular library are clustered, and multiple molecules with excellent performance are selected from each cluster to form a fourth candidate molecular library.” (Claim 3) BJT additionally teaches “selecting multiple molecules with excellent docking conformations to form a third candidate molecule library.” ([0010]) BJT discloses that “each molecule in the second candidate molecule library is docked with the target site, and several third candidate molecules with excellent docking conformations are selected…[and] the docking scoring of each molecule in the molecular library can be automatically continued, and multiple molecules with excellent docking conformations can be selected to form a third candidate molecular library. Among them, excellent docking conformation can mean a high docking score, such as ranking in the top 5% of docking scores…” ([0059], [0060])
With respect to claim 15, BJT further teaches “selecting multiple molecules with the highest output activity values to form a sixth candidate molecule library” and “the output activity value is ranked first, for example, the top 5% of the activity values… Here, for each molecule in the first candidate molecule library, based on the activity value output by the model, several molecules with the highest activity values are selected to form the second candidate molecule library.” ([0013], [0058]) BJT additionally teaches “a method that extracts information about the three-dimensional spatial structure and pharmacophore properties of molecules, vectorizes this information, and then uses the mean clustering algorithm for clustering.” ([0065]) BJT discloses “several molecules with excellent performance are selected to form a fourth candidate molecule library…[and] excellent performance can be reflected in a high activity value in the first prediction model, a high docking score, or a high ranking in other physicochemical properties.” ([0066]) BJT further discloses “based on the activity value output by the model, several molecules with the highest activity values are selected to form the second candidate molecule library.” ([0058])
With respect to claim 16, BJT further teaches “a computing device is provided, comprising: a memory; one or more processors; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the small molecule drug screening method.” (Claim 9, [0019])
With respect to claim 17, BJT further teaches “a computing device is provided, comprising: a memory; one or more processors; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the small molecule drug screening method.” (Claim 9, [0019])
With respect to claim 18, BJT further teaches “a computing device is provided, comprising: a memory; one or more processors; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the small molecule drug screening method.” (Claim 9, [0019])
With respect to claim 19, BJT further teaches “a computing device is provided, comprising: a memory; one or more processors; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the small molecule drug screening method.” (Claim 9, [0019])
With respect to claim 20, BJT further teaches “a computer-readable storage medium is provided for storing one or more programs, said one or more programs including
instructions that, when executed by a computing device, cause the computing device to perform the small molecule drug screening method as described above.” (Claim 10, [0020])
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over BJT in view of Srivastava et al (US2013/0184462A1, 22 May 2011) and Rocca et al (Molecules, 24 December 2014, pages 206 - 223). (Herein referred to as Srivastava and Rocca, respectively.)
BJT teaches the methods of claims 1 and 2 as described above.
BJT does not teach ” performing Lipinski's Rule of Five-based screening on the compounds in the compound library based on the target protein, to obtain a plurality of compounds obeying the Lipinski's Rule of Five; and deduplicating the plurality of compounds obeying the Lipinski's Rule of Five, to obtain the plurality of screened compounds.”
Srivastava teaches “[a]ll the active derivatives showed compliance with Lipinski's rule of five for oral bioavailability and toxicity risk assessment parameters.” ([0047]) Srivastava further teaches “we also checked the compliance of compounds to Lipinski's rule-of-five for drug likeness.” ([0070])
Rocca teaches “[t]he database used, containing 26,191 molecules, was built combining libraries of alkaloids and berberine analogues. Subsequently, the filter for drug-like properties, the Lipinski’s rule of five and the deduplication led us to globally consider 14,175 compounds.” (Page 210, Section 2.1, paragraph 2) Rocca further teaches “they were filtered basing on their drug-like properties as it has been addressed by the Lipinski’s rule of 5 [97] and the duplicated structures were removed.” (Page 215, Section 3.1, paragraph 1)
It would have been prima facie obvious to one of ordinary skill in the art at the effective
filing date of the invention to have applied the compound screening methods of Srivastava and Rocca to the drug screening method of BJT. Lipinski’s Rule of Five is a well-known set of criteria utilized in the art to screen compounds. Srivastava discloses “[m]olecules violating more than one of these rules may have problems with bioavailability.” [0067] Pollastri discloses that “[i]n the mid- to late 1990s, because of the drug discovery paradigm shift from phenotypic screens to combinatorial chemistry and high-throughput screening, the physicochemical properties of exploratory drug molecules displayed a dramatic shift toward higher molecular weight and lipophilicity. In response, Lipinski and coworkers reported an analysis of compounds that successfully navigated Phase I and entered into Phase II clinical studies, and correlated the computed physicochemical properties of these molecules to their aqueous solubility, permeability, and oral bioavailability. In doing so, the authors created the “Rule of Five,” a mnemonic tool for medicinal chemists to use to quickly assess compounds during the drug discovery and optimization process with respect to the compounds’ likelihood to display good solubility and permeability profiles.” (Current Protocols in Pharmacology, 1 June 2010, pages 9.12.1-9.12.8) (Abstract) Pollastri further discloses “the Ro5 [Rule of Five] is frequently utilized within drug discovery projects. During the compound design phase, where medicinal chemists are developing ideas for the next round of analog synthesis, the level of Ro5 compliance is typically computed and factored into the compound designs. For parallel array design, large collections of molecules are frequently designed in silico and “shaped” by including or excluding those molecules that violate one or more of the rules… the molecular structure and other properties are frequently registered into a database… free, Web-based Ro5 calculators are available that are fast and easy to use for analysis of singleton compounds.” (Pages 9.12.4 -9.12.5, Implementation of the Ro5) In addition, deduplication is frequently utilized in the art when screening molecules. Dalecki et al teaches that “[a]s more data sources are added from diverse sources the structure processing becomes increasingly important, as standard representation and deduplication is essential for high quality machine learning model building.” (Metallomics, 22 February 2019, pages 696-706) (Page 698, Left column, lines 24-27) Therefore, one of ordinary skill in the art would have been motivated to utilize the known techniques of performing Lipinski’s Rule of Five and deduplicating a set of compounds obeying Lipinski’s Rule of Five in an artificial intelligence-based drug molecule processing method. The invention is therefore prima facie obvious.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over BJT in view of BJT and Yuan et al. (US12040094B2, 4 September 2019) (Herein referred to as Yuan.)
BJT teaches the methods of claims 1 and 2 as described above. BJT further teaches “the first candidate molecular library is constructed by a molecular library construction model, which is at least one of a molecular generation model, a substructure matching model, and a chemical property-based filtering model.” ([0015])
BJT does not teach “removing an enantiomer of a chiral compound from the plurality of filtered compounds, to obtain the plurality of candidate drug molecules for the target protein.”
Yuan teaches “[t]he method can further include defining each of the molecules (e.g., molecules in the first or second data sets) by a plurality of selected features… the selected features can further include chirality of the molecule. By including chirality, it is possible to convert the molecules into graphs without losing spatial information. It should be understood that some different molecules (e.g., enantiomers, diastereomers, etc.) can have the same SMILES notation but different spatial structures.” (Column 8, lines 43-45 and 50-55) Yuan further teaches “[d]uring preprocessing, SMILES with bio-activity data are read from each dataset. The input SMILES do not need to be canonical since the model will rearrange the input atoms in a preset order. Chirality information can be included in each input in order to distinguish between isomers.” (Column 26, lines 5-9)
It would have been prima facie obvious to one of ordinary skill in the art at the effective
filing date of the invention to have combined the drug screening method of BJT with the additional features of BJT and the method of Yuan. BJT discloses “[f]iltering based on chemical properties, such as filtering according to drug class rules, or selecting based on other chemical properties, can screen out molecules with superior properties.” ([0044]) Brooks et al discloses “[p]roteins are often enantioselective towards their binding partners. When designing small molecules to interact with these targets, one should consider stereoselectivity. As considerations for exploring structure space evolve, chirality is increasingly important. Binding affinity for a chiral drug can differ for diastereomers and between enantiomers. For the virtual screening and computational design stage of drug development, this problem can be compounded by incomplete stereochemical information in structure libraries leading to a “coin toss” as to whether or not the “ideal” chiral structure is present.” (Current Topics in Medicinal Chemistry, 1 April 2011, pages 760-770) (Abstract) Brooks et al further discloses “[t]he enantiomers may differ in their effects and so the exact composition of a racemate was of concern. Enantiomers that have the desired effect are termed eutomers. Enantiomers that do not have the desired effect or even a detrimental effect are termed distomers.” (Page 762, left column, section 2.2, lines 9-13) Brooks et al additionally discloses “[c]hirality in drug discovery and development has increased in importance since guidelines for chiral compounds in drug development and approval were issued by regulatory agencies in the 1992-2000 timeframe. The guidelines were issued due to the potential differing activities of enantiomers in chiral compounds. Although enantiomers will have the same properties in achiral environments, the effects can be quite different in biological environments where enzymes and receptors can be chirally selective. In some cases, the presence of the distomer in a racemic mixture can affect the results due to detrimental effects of the distomer or its con version to the eutomer configuration. The composition of the racemic mixture and its potential to change with time or de pending on the system and tests (ex. animal studies versus clinical trials) meant that the more active enantiomer in a single enantiomer drug would be a better option in many cases if scientifically and technically possible.” (Page 769, right column, Conclusion paragraph 1) Therefore, one of ordinary skill in the art would have been motivated to use known techniques of chemically filtering a plurality of screened compounds and removing an enantiomer of a chiral compound to obtain candidate drug molecules. The invention is therefore prima facie obvious.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over BJT in view of Yuan and Yang et al. (Journal of Chemical Information and Modeling, 30 July 2019, pages 3370-3388) (Herein referred to as Yang.) BJT teaches the methods of claims 1 and 5 as described above.
BJT does not teach “determining a molecular graph of the candidate drug molecule based on the molecular structure of the candidate drug molecule; and performing image encoding on the molecular graph of the candidate drug molecule, to obtain the embedding feature of the candidate drug molecule.”
Yuan teaches “ that DEEPCHEM (or other software tool) can be used to describe molecules by features of the atoms. When chirality is included as a selected features, the molecules can be described with thirty two of about seventy five features offered by DEEPCHEM. Following featurization, the method can further include converting the molecules defined by the selected features into a plurality of respective graphs associated with each of the molecules.” (Column 8, lines 56-64) Yuan further teaches “DeepChem's implementation of GCNN is used. This implementation offers the creation of architectures with graph convolutional layers, graph pooling layers, dropout layers, graph gather layers, and fully connected layers. The molecular graph is sorted via atom index in order to attain the same graph for canonical SMILES.” (Column 19, lines 3-9)
Yang teaches “we expand on the characteristics of Directed MPNN (D-MPNN) used in this paper…we refer to it as Directed MPNN to show it is a variant of the generic MPNN architecture…During training, the network takes molecular graphs as input and outputs a prediction for each molecule.” (Page 3371, right column, methods section paragraphs 1 and 4) Yang further teaches “D-MPNN uses messages associated with directed edges (bonds)” and “[u]sing Figure1 as an illustration, in D-MPNN, the message 1→2 will only be propagated to nodes 3 and 4 in the next iteration.” (Page 3372, left column, Directed MPNN section lines 5-6 and 10-12, and Figure 1) Note that the instant application states “image encoding is performed on the molecular graph of the candidate drug molecule through an image encoder (for example, DMPNN).” ([0082])
It would have been prima facie obvious to one of ordinary skill in the art at the effective
filing date of the invention to have combined the drug screening methods of BJT with the additional methods of Yuan and Yang. Yang discloses that a “main line of research is the optimization of the model architecture, whether the model is applied to descriptors or fingerprints or is directly applied to SMILES strings or the underlying graph of the molecule. Our model belongs to the last category of models, known as graph convolutional neural networks. In essence, such models learn their own expert feature representations directly from the data, and they have been shown to be very flexible and capable of capturing complex relationships given sufficient data.” (Page 3371, right column, paragraph 1) Yang further discloses “[t]he motivation of this design [re: D-MPNN] is to prevent totters, that is, to avoid messages being passed along any path of the form v1v2···vn where vi=vi+2 for some i. Such excursions are likely to introduce noise into the graph representation… as an illustration… creating an unnecessary loop in the message passing trajectory.” (Page 3372, left column, lines 6-10, 13) Therefore, one of ordinary skill in the art would have been motivated to utilize the known technique of creating molecular graphs to the improved methods of D-MPNN to provide better screening methods for potential drug candidates. The invention is therefore prima facie obvious.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over BJT in view of Gibson et al. (US12009064B2, 14 November 2018) (Herein referred to as Gibson.) BJT teaches the methods of claims 1 and 5 as described above.
BJT does not specifically teach “summing the embedding feature of the candidate drug molecule and the embedding feature of the target protein, to obtain the activity fusion feature; or concatenating the embedding feature of the candidate drug molecule and the embedding feature of the target protein, to obtain the activity fusion feature.”
Gibson teaches “[i]n some embodiments, combining the vectors is performed using a predetermined mathematical operation, such as concatenation, addition, mean, etc. In some embodiments, combining the vectors in performed using a learned mathematical operation.” (Column 58, lines 8-12) Gibson further teaches “[t]he dimension-reduced, whitened, and normalized instance vectors for each of the 1804 remaining compound in the candidate library were concatenated, resulting in 1804 concatenated compound vectors having approximately 800-dimensions, each compound vector representing the approximately 200 reduced-dimensionality feature measurements extracted from each of the four cell contexts.” (Column 67, lines 12-18)
It would have been prima facie obvious to one of ordinary skill in the art at the effective
filing date of the invention to have combined the drug screening methods of BJT with the additional method of Gibson. Gibson discloses “the method further includes reducing a dimension of each vector in the plurality of vectors using a dimension reduction technique…the dimension reduction technique is principal component analysis in which a plurality of principal components is identified based on a variance in the measurement of each different feature in the plurality of features…across each compound in the plurality of compounds, and each respective vector in the plurality of concatenated vectors for the cell context is re-expressed as a projection of the respective concatenated vector onto the plurality of principal components.” (Column 5, lines 48-59) Therefore, one of ordinary skill in the art would have been motivated to utilize the known and routine technique of concatenating vectors to obtain fusion information for more efficient and complete activity prediction for determining candidate drug molecules. The invention is therefore prima facie obvious.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over BJT in view of CN111462833A (Shenzhen Zhiyao Information Technology, 20 January 2019) and Maragakis et al. (Journal of Chemical Information and Modeling, 22 July 2020, pages 4487-4496) (Herein referred to as SZIT and Maragakis, respectively.) BJT teaches the methods of claims 1 and 5 as described above.
BJT does not specifically teach “mapping the embedding feature of the candidate drug molecule and the embedding feature of the target protein, to obtain an intermediate feature vector comprising the candidate drug molecule and the target protein; and performing affine transformation on the intermediate feature vector, to obtain the activity fusion feature.”
SZIT teaches “mapping the predetermined structural information of the compound's neighboring atoms and the protein's neighboring atoms into a structural information matrix group includes: mapping the atom type, partial charge number, distance from a reference atom, and amino acid residue type to which the protein's neighboring atoms belong into an atom type matrix, a partial charge number matrix, a distance matrix, and an amino acid residue type matrix, respectively; using a neural network to perform an embedding operation on the structural information matrix group, and obtaining a representation matrix of the ligand compound-target protein complex from the embedded structural information matrix group.” ([0011] of machine translation)
Maragakis teaches “DESMILES encodes the molecular fingerprint. The encoding of the fingerprint is then passed to a decoder RNN, which learns to generate the corresponding SMILES string. Since the fingerprint can be represented as a fixed-length vector, we used a simple two-layer feed-forward neural network with tanh activation functions for the encoder. The size of the output of the final layer was restricted to be the same size as that of the decoder hidden state. Each layer of the encoder consists of batch normalization, an affine transformation, and a hyperbolic-tangent activation function.” (Page 4492, left column, lines 20-30)
It would have been prima facie obvious to one of ordinary skill in the art at the effective
filing date of the invention to have combined the drug screening methods of BJT with the additional methods of SZIT and Maragakis. SZIT discloses “[m]olecular docking
allows for continuous adjustment of the binding conformation between ligand compounds
and target proteins, thereby predicting the optimal binding mode and corresponding binding
strength. This is of great significance for optimizing drug structure and elucidating
biochemical processes.” ([0005]) Maragakis discloses “that a deep-learning model could be developed to translate fingerprints to SMILES, and that additional fine tuning of such a model could make it effective for proposing modified molecules that are both structurally legible and have improved properties of interest.” (Page 4488, left column, lines 8-12) Regarding use of affine transformations in protein modeling, Gao et all discloses “recent advances in machine learning, especially in deep learning (DL)-related techniques, have opened up new avenues in many areas of protein modeling. DL is a set of machine learning techniques based on stacked neural network layers that parameterize functions in terms of compositions of affine transformations and non-linear activation functions. Their ability to extract domain-specific features that are adaptively learned from data for a particular task often enables them to surpass the performance of more traditional methods.” (Patterns. 11 December 2020, pages 1-23) (Page 2, left column, lines 7-17) Therefore, one of ordinary skill in the art would have been motivated to utilize the known techniques of mapping embedding features and performing affine transformation on feature vectors to improve activity prediction for determining candidate drug molecules. The invention is therefore prima facie obvious.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over BJT in view of Gaspar et al. (American Chemical Society: Frontiers in Molecular Design and Chemical Information Science, 1 January 2016, pages 211-241) (Herein referred to as Gaspar.) BJT teaches the methods of claims 1 and 5 as described above.
BJT does not specifically teach “mapping the activity fusion feature to a latent vector space, to obtain a latent vector of the activity fusion feature; and performing nonlinear mapping on the latent vector of the activity fusion feature, to obtain the activity information of the candidate drug molecule.”
Gaspar teaches “Generative Topographic Mapping or GTM, introduced by Bishop et al. in the 1990s, performs a dimensionality reduction from the initial D-dimensional data space to 2-dimensional latent space (“GTM” map) by embedding a 2-dimensional non-linear manifold into the data space.” (Page 211, Section 1, paragraph 1) Gaspar further teaches “Each node xk in the latent space is mapped to the corresponding manifold point yk in the initial data space by the non-linear mapping function y(x;W).” (Page 212, Figure 1)
It would have been prima facie obvious to one of ordinary skill in the art at the effective
filing date of the invention to have combined the drug screening methods of BJT with the mapping method of Gaspar. Gaspar discloses, regarding GTM, that “[i]ts probabilistic nature can be exploited in order to build regression or classification models, to define their applicability domain, to predict activity profiles of compounds, to compare large datasets, to screen for compounds of interest, and even to identify new molecules possessing desirable properties. Thus,
GTM can be seen as a sort of a multi-purpose Swiss knife, each of its blades being able to shape an answer to a specific chemoinformatics question, based on a unique map.” (Page 211, Introductory paragraph, lines 4-12) Gaspar further discloses that “such a manifold able to simultaneously support many models of completely unrelated and biologically relevant properties may claim to be a “universal” map of (drug-like) chemical space.” (Page 228, Section 7, paragraph 3) Therefore, one of ordinary skill in the art would have been motivated to utilize the known mapping techniques described by Gaspar to improve the activity prediction of the method of BJT to screen for potential drug candidate molecules. The invention is therefore prima facie obvious.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over BJT in view of SZIT and Srivastava. BJT teaches the methods of claims 1 and 12 as described above.
BJT does not explicitly teach “performing format transformation on the plurality of candidate drug molecules respectively, to obtain a transformation format of each of the plurality of candidate drug molecules; constructing a three-dimensional conformation of each of the plurality of candidate drug molecules based on the transformation format of each of the plurality of candidate drug molecules; determining a hydrogen atom addible position of each of the plurality of candidate drug molecules based on the three-dimensional conformation of each of the plurality of candidate drug molecules; and adding a hydrogen atom to a hydrogen atom addible position of a candidate drug molecule, to obtain the molecular conformation of the candidate drug molecule.”
SZIT teaches, regarding molecular docking, that “[t]he process involves using a molecular docking program to perform molecular docking on files containing the three-dimensional structures of ligand compounds and target proteins. After docking, the files containing structural information such as atom types, partial charges, distances from reference atoms, and the types of amino acid residues to which the atoms in the protein molecule belong are saved. In one embodiment of the present invention, the smina molecular docking program is used, and the molecular structure file format is pdbqt.” ([0055]) SZIT further teaches “Using the docked structure file, for each atom (reference atom) in the ligand compound molecule, its neighboring atoms (including itself) are determined according to the Euclidean distance in three-dimensional space within the compound molecule and in the target protein molecule. The predetermined structural information of these neighboring atoms is recorded, including the atom type, the number of bias charges, the distance from the reference atom, and the type of amino acid residue to which they belong (for neighboring atoms from the protein molecule)… Furthermore, using the simplified molecular input line entry specification (SMILES) sequence representation of the ligand compound molecule as input, the physicochemical properties and molecular fingerprint of the compound are calculated using cheminformatics software, and the results are saved as a numerical vector p0 (referred to as the physicochemical property vector) for use in a neural network. In one embodiment of the present invention, the reference atom is taken as 6 neighboring atoms in the ligand compound and 2 in the protein; the calculated physicochemical properties include the number of hydrogen bond donors, the number of hydrogen bond acceptors, the number of rotatable chemical bonds, the number of aromatic rings, the relative molecular mass, the topological polar surface area, and the n-octanol-water partition coefficient; the molecular fingerprint used is ECFP-4 with a length of 1024.” ([0056])
Srivastava teaches “[t]he valency and hydrogen bonding of the ligands as well as target proteins were subsequently satisfied through the Workspace module of Scigress Explorer software. Hydrogen atoms were added to protein targets for correct ionization and tautomeric states of amino acid residues such as His, Asp, Ser and Glu etc.” ([0065])
It would have been prima facie obvious to one of ordinary skill in the art at the effective
filing date of the invention to have combined the drug screening methods of BJT with the additional methods of SZIT and Srivastava. Regarding the benefits of molecular docking, SZIT discloses “[m]olecular docking allows for continuous adjustment of the binding conformation between ligand compounds and target proteins, thereby predicting the optimal binding mode and corresponding binding strength. This is of great significance for optimizing drug structure and elucidating biochemical processes.” ([0005]) SZIT further discloses “[o]ne of the key aspects of molecular docking is the scoring function, which scores the binding conformation of the ligand compound and the target protein as an approximation of the binding free energy, and is used to guide conformation sampling—selecting the optimal binding conformation by minimizing the scoring function (i.e. maximizing the absolute value of the binding energy). Commonly used scoring functions can be mainly divided into three categories: the first is force field-based scoring functions, which cover interactions such as van der Waals forces, electrostatic forces, and hydrogen bonding forces, and calculate the binding energy of molecules from scratch based on first-principles simulations; the second is prior knowledge-based scoring functions, which use known structural data and their binding energies in existing databases to generate some simplified coefficient terms to approximate complex physical interactions, such as establishing binding energy coefficients for pairs of all atom types and summing them as an approximation of the binding energy, which greatly reduces the amount of computation but increases the risk of overfitting; the third is experience-based scoring functions, which integrate force field-based and prior knowledge-based scoring functions, including physical parameters of some force fields, and also setting parameters such as hydrophobic interactions and desolvation interactions, which can be fitted by regression using existing known data.” ([0006]) Regarding the benefits of protonation in virtual screening and molecular docking, Berry et al teaches “[a]lthough most receptor preparation tools accurately complete processes that were not undertaken during X-ray crystal structure refinement, it is important to understand these processes and make adjustments where necessary. The most common receptor preparation procedures include adding hydrogens and atom-type charges, but it is also important to ensure that missing side-chains are added, missing bonds and molecule chain breaks are detected and fixed, bond orders are assigned, and where alternate locations are present, the atoms with highest frequencies must be selected. Other, more complicated, procedures in receptor preparation include accurate prediction of protonation states and identifying which water molecules (if any) should remain in the receptor structure. All of these procedures maximize the bio logical realism in the modeled system, which leads to the identification of a higher proportion of true bioactives.” (Emerging Trends in Computational Biology, Bioinformatics, and Systems Biology, Chapter 27: Practical Considerations in Virtual Screening and Molecular Docking, July 2015, pages 487-502) (Page 488, Section 2, paragraph 1) Therefore, one of ordinary skill in the art would have been motivated to utilize the known preprocessing techniques of SZIT and the use of protonation described Srivastava to improve the application of molecular docking in the method of BJT to screen for potential drug candidate molecules. The invention is therefore prima facie obvious.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. CN112053742A teaches a method of screening for molecular target proteins that can be applied to drug discovery and development. (23 July 2020) Hassan-Harrirou et al teaches an ensemble of three-dimensional (3D) Convolutional Neural Networks (CNNs), which combines voxelized molecular mechanics energies and molecular descriptors for predicting the absolute binding affinity of protein−ligand complexes to improve and accelerate drug development. (JCIM, 11 May 2020, pages 2791−2802) Agarwal and Nayarisseri teach an in silico approach using molecular docking and virtual screening approaches to develop treatments for PCOS. (MOL2NET'20, 20 February 2020, pages 1-4) Batool et al teaches a review of available methods and algorithms for structure-based drug design including virtual screening and de novo drug design, with a special emphasis on AI- and deep-learning-based methods used for drug discovery. (International Journal of Molecular Sciences, 6 June 2019, pages 1-18) Bajpai et al teaches a method using artificial intelligence and machine learning to screen candidate pharmaceutical molecules or compounds for pharmaceutical uses based on structural similarity between the candidate pharmaceutical molecules and various pharmaceutical targets. (US11127488B1, 21 January 2021)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHARON LEVINE GRAFF whose telephone number is (571)317-0219. The examiner can normally be reached Mon - Fri 7:30 AM - 4 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Karlheinz Skowronek can be reached at (571) 272-9047. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.L.G./Examiner, Art Unit 1687
/Karlheinz R. Skowronek/Supervisory Patent Examiner, Art Unit 1687