Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-20 are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by Norvaisas et al (US 2022/0351053), hereinafter Norvaisas.
Regarding claim 1, NORVAISAS discloses a method for providing explainable artificial intelligence (Fig. 37, para [0188] "The API module 3734 communicates with the artificial intelligence assistance system 3720 via API calls, sending assay results to the Al assistance system 3720, receiving instructions from the Al assistance system 3720, exchanging data and instructions with the other orchestration software components 3731-3733, and providing information to the AI assistance system 3720 about available resources such as chemical libraries and chemical synthesis capabilities."), comprising: receiving, at a computing device, preliminary data corresponding to one or more biological problems (Fig. 38, para [0189] "In this embodiment, the process involves iteratively performing physical assays on drug candidate molecules using an iterative physical testing system 3710, processing the results of the assays through an Al assistance system 3720 to generate new drug candidate molecules that may better fit a hypothesis or bioactivity goal, performing assays on the new drug candidate molecules, and repeating the process until some pre-defined parameter is met (e.g., a molecule is discovered with a better bioactivity than the drug candidate molecule, a molecule is discovered with precursors that meet certain criteria such as cost or availability, the drug candidate molecule and hypothesis are confirmed to a certain level of confidence, etc.)."); generating, based on a transformation of the preliminary data, a plurality of candidate features, each feature comprising a molecular mechanism (Fig. 7, para [0113] "FIG. 7 is a diagram illustrating an exemplary graph-based representation of molecules as simple relationships between atoms using a matrix of adjacencies 700, wherein atoms are represented as nodes and bonds between the atoms are represented as edges Further, the graph-based representation of a molecule can be stated in terms of two matrices, one for the node features (e.g., type of atom and its available bonds) and one for the edges (i.e., the bonds between the atoms). The combination of the nodes (atoms) and edges (bonds) represents the molecule. Each molecule represented in the matrix comprises a dimensionality and features that describe the type of bond between the atoms."), wherein a given molecular mechanism comprises: information that is determinative of the operation of a step in at least one biochemical pathway corresponding to the one or more biological problems (Fig. 10, para [0215] "At the training stage, the adjacency matrices 1011 and node features matrices 1012 for many molecules are input into the MPNN 1020 along with vector representations of known or suspected bioactivity interactions of each molecule with certain proteins."), wherein the determinative information may comprise one or more of: one or more aspects of molecules involved in the step of the at least one biochemical pathway (para [0215] "Based on the training data, the MPNN 1020 learns the characteristics of molecules and proteins that allow interactions and what the bioactivity associated with those interactions is. At the analysis stage, a drug candidate molecule is input into the MPNN 1020, and the output of the MPNN 1020 is a vector representation of that molecule's likely interactions with proteins and the likely bioactivity of those interactions."), or one or more environmental conditions corresponding to the step of the at least one biochemical pathway (Fig. 32, para [0171] "Each of the product nodes 3210 and the reactant nodes 3221-3222, 3231 in the hypergraph comprise a specific reactant (i.e., molecule). Once a product node 3210 or reactant node 3221-3222, 3231 has been visited, information about that specific reactant can be queries from the environment, Eq. From each node, an action can be taken to select a specific hypernode 3220, 3230."); reducing, by performing one or more feature reduction operations, the plurality of candidate features into an optimized feature set (Fig. 14, para [0125] The neural networks build a model from the training data. In the case of using an autoencoder (or a variational autoencoder), the encoder portion of the neural network reduces the dimensionality of the input molecules, learning a model from which the decoder portion recreates the input molecule. The significance of outputting the same molecule as the input is that the decoder may then be used as a generative function for new molecules.") "), wherein the one or more feature reduction operations comprise one or more of: selecting, based on one or more predictive scores indicating a likelihood of a given candidate feature to be predictive of a solution to a biological problem, a subset of candidate features (Fig. 30, para [0157] "The hypergraph search engine 3012 creates hypernodes from each set of retrosynthesized precursors, receives forward reaction prediction scoring for each node, assigns awards based on the scoring, selects a preferred node at each recursion stage, and back-propagates the rewards to the root node (the drug candidate molecule), thus creating a preferred chemical synthesis pathway from some set of preferred precursors (preferably commercially available) to the drug candidate molecule."), eliminating, based on a comparative indicator of the likelihood of a first candidate feature, of the plurality of candidate features, to be a false positive, the first candidate feature, wherein the comparative indicator indicates that the likelihood of the first candidate feature to be a false positive exceeds a threshold (Fig. 31, para [0166] "As retrosynthesis is a non-determinative process, many possible precursors and reaction pathways may be predicted. Of these, some will be theoretically plausible but invalid, some will be valid but unfeasible or impractical, and some will be both valid and practical. At each recursive stage, validation of the predicted precursors may be performed to determine a likely validity of the predicted precursors and the chemical reaction by which the precursors are predicted to result in the molecule at the next higher recursion stage. For example, after the first retrosynthesis step 3101, two intermediate molecules are predicted 3121, 3122 which, given some chemical reaction, are expected to produce the drug candidate molecule 3111."),
combining, based on one or more similarity scores, two or more candidate features with predictive effects on one or more solutions to the one or more biological problems and corresponding to matching similarity scores of the one or more similarity scores (Fig. 31, para [0166] "The forward reaction as predicted by the machine learning algorithm may be scored by one of several means (e.g., the forward reactant's log-probability, a synthesis score, forward-synthesis likelihood, etc.), and the results of the scoring may be fed to a hypergraph search engine which assigns rewards 3124, 3137 at each recursive stage as a form of reinforcement learning. As will be described later herein, a hypergraph in conjunction with an MCTS algorithm is useful in the context as it allows for association of actions (e.g., chemical reactions) with hypernodes in addition to objects (e.g., precursors), and because the MCTS algorithm constantly re- assesses the current preferred pathway each time rewards are updated at each recursive stage."); training, based on the optimized feature set, a predictive model, wherein training the predictive model configures the predictive model to output solutions to biological problems and representations of decisions made by the predictive model to output the solutions to biological problems (Fig. 10, para [0215] "At the training stage, the adjacency matrices 1011 and node features matrices 1012 for many molecules are input into the MPNN 1020 along with vector representations of known or suspected bioactivity interactions of each molecule with certain proteins. Based on the training data, the MPNN 1020 learns the characteristics of molecules and proteins that allow interactions and what the bioactivity associated with those interactions is. At the analysis stage, a drug candidate molecule is input into the MPNN 1020, and the output of the MPNN 1020 is a vector representation of that molecule's likely interactions with proteins and the likely bioactivity of those interactions."); outputting, using the predictive model, one or more solutions for the one or more biological problems (Fig. 34, para [0179] "A check is made to determine whether there is a set of precursors (reactants) that can produce the drug candidate molecule, all of which precursors are commercially available 3410. If there is no such set, the process is repeated recursively from step 3401 until such a set is found. If there is such a set, recursion is ended 3411 and the preferred chemical pathway and its commercially available precursors are returned as an output."); outputting, using the predictive model, a representation of decisions made by the predictive model to output the one or more solutions (Fig. 34, para [0179] "A check is made to determine whether there is a set of precursors (reactants) that can produce the drug candidate molecule, all of which precursors are commercially available 3410. If there is no such set, the process is repeated recursively from step 3401 until such a set is found. If there is such a set, recursion is ended 3411 and the preferred chemical pathway and its commercially available precursors are returned as an output."); and applying, based on the representation of decisions made by the predictive model to output the one or more solutions, one or more decisions made by the predictive model to at least one additional predictive model (Fig. 37, para [0180] FIG. 37 is an exemplary system architecture for a system for feedback-driven automated drug discovery. In this embodiment, the system 3700 combines an iterative physical testing system 3710 of drug candidate molecules for agreement with hypotheses with an artificial intelligence/machine learning system 3720 configured to generate predictions of additional drug candidate molecules for testing which may better match a given hypothesis.").
Regarding claim 2, NORVAISAS discloses the method of claim 1, wherein the applying the one or more decisions comprises: revising, through an iterative loop and based on the one or more solutions for the one or more biological problems (Fig. 38, para [0189] FIG. 38 is a high-level process flow diagram illustrating exemplary operation of a system for feedback-driven automated drug discovery. In this embodiment, the process involves iteratively performing physical assays on drug candidate molecules using an iterative physical testing system 3710, processing the results of the assays through an Al assistance system 3720 to generate new drug candidate molecules that may better fit a hypothesis or bioactivity goal. the predictive model, wherein the iterative loop comprises: receiving, based on outputting the representation of decisions made by the predictive model, feedback information corresponding to one or more human biology experts (Fig. 1, para [0099] "The data platform 110 in this embodiment comprises a knowledge graph 111, an exploratory drug analysis (EDA) interface 112, a data analysis engine 113, a data extraction engine 114, and web crawler/database crawler 115. The crawler 115 searches for and retrieves medical information such as published medical literature, clinical trials, dissertations, conference papers, and databases of known pharmaceuticals and their effects."); updating, based on the feedback information, the predictive model (Fig. 12, para [0228]); and repeating, during the outputting the one or more solutions for the one or more biological problems, the receiving feedback information and the updating until a solution has been outputted for each biological problem of the one or more biological problems (Fig. 39, para [0196]).
Regarding claim 3, NORVAISAS discloses the method of claim 1, wherein the applying the one or more solutions one or more decisions made by the predictive model to the at least one additional predictive model comprises: defining, based on the one or more decisions made by the predictive model, a decision making process for the at least one additional predictive model (Fig. 30, para [0157] "The recursion engine 3011 manages the recursive process of repeatedly retrosynthesizing larger molecules into smaller molecules that are precursors of that larger molecule. The process is recursive because each stage of the process is a one-to-many process, and there may be many possible chemical pathways to produce a drug candidate molecule. It is not necessary to know all of the precursors needed to synthesize the drug candidate molecule; each recursive stage of retrosynthesis will find the precursors of the previous stage."); and outputting, using the at least one additional predictive model, a solution for a problem adjacent to the one or more biological problems (Fig. 30, para [0157] "The hypergraph search engine 3012 creates hypernodes from each set of retrosynthesized precursors, receives forward reaction prediction scoring for each node, assigns awards based on the scoring, selects a preferred node at each recursion stage, and back-propagates the rewards to the root node (the drug candidate molecule), thus creating a preferred chemical synthesis pathway from some set of preferred precursors (preferably commercially available) to the drug candidate molecule."), wherein the problem adjacent to the one or more biological problems comprises one or more parameters matching at least one parameter of the one or more biological problems (Fig. 30, para [0157]).
Regarding claim 4, NORVAISAS discloses the method of claim 1, wherein the given molecular mechanism further comprises one or more of: molecular substructures involved in the step of the at least one biochemical pathway (Fig. 18, para [0145] "According to one aspect, various cheminformatics libraries may be used as a learned force-field for docking simulations, which perform gradient descent of the ligand atomic coordinates with respect to the binding affinity 1806 and pose score 1805 (the model outputs). This requires the task of optimizing the model loss with respect to the input features, subject to the constraints imposed upon the molecule by physics (i.e., the conventional intramolecular forces caused for example by bond stretches still apply and constrain the molecule to remain the same molecule)."), geometric constraints between one or more molecular substructures corresponding to a single molecule of a plurality of molecules involved in the step of the at least one biochemical pathway (Fig. 18, para [0145] "According to one aspect, various cheminformatics libraries may be used as a learned force-field for docking simulations, which perform gradient descent of the ligand atomic coordinates with respect to the binding affinity 1806 and pose score 1805 (the model outputs). This requires the task of optimizing the model loss with respect to the input features, subject to the constraints imposed upon the molecule by physics (i.e., the conventional intramolecular forces caused for example by bond stretches still apply and constrain the molecule to remain the same molecule)."), catalysts affecting a rate of the step of the at least one biochemical pathway, or environmental conditions affecting the step of the at least one biochemical pathway (Fig. 18, para [0145]).
Regarding claim 5, NORVAISAS discloses the method of claim 1, wherein the selecting of the subset of candidate features is based on one or more differential effects of each molecular mechanism, corresponding to the plurality of candidate features, on at least one outcome of solving the one or more biological problems (Fig. 10, para [0215]).
Regarding claim 6, NORVAISAS discloses the method of claim 1, wherein the eliminating the first candidate feature is based on the comparative indicator further indicating that instances of a first molecular mechanism corresponding to the first candidate feature being incorporated into instances of a second molecular mechanism, corresponding to at least one second candidate feature and having a size exceeding the first molecular mechanism, exceed a frequency threshold (Fig. 28, para [0155] "The training data presents a choice of a threshold bracket 2830. The threshold bracket is a trade-off between the average information contained in each datapoint, and the sheer quantity of data, assuming that datapoints with more extreme inactive/active IC50 values are indeed more typical of the kind of interactions that determine whether or not a protein-ligand pair is active or inactive. In the case of the 3D-model, using the dataset with no threshold performs consistently better across most metrics. The channels used for the data set are hydrophobic, hydrogen-bond donor or acceptor, aromatic, positive or negative ionizable, metallic and total excluded volume. Regardless of the choice of threshold, the data is then used to train a 3D -CNN to know the classification of a molecule regarding activation and pose propriety 2840.").
Regarding claim 7, NORVAISAS discloses the method of claim 1, wherein the combining candidate features comprises one or more of: combining a plurality of geometric constraints into a range of geometric constraints, or combining a plurality of molecular substructures sharing a threshold percentage of biochemical traits into a molecular substructure with a plurality of specification choices (Fig. 18, para [0145]), where the plurality of specification choices comprise one or more of: permitting substitution of a first atom with a second atom in the same column of the periodic table (Fig. 37, para [0185] "The module utilizes the knowledge graph 111 and data analysis engine 113 capabilities of the data platform 110, and in one embodiment is configured to identify ligands with certain properties based on three dimensional (3D) models 131 of known ligands and differentials of atom positions 132 in the latent space of the models after encoding by a 3D convolutional neural network (3D CNN), which is part of the data analysis engine 113. In one embodiment, the 3D model comprises a voxel image (volumetric, three dimensional pixel image) of the ligand. In cases where enrichment data is available, ligands may be identified by enriching the SMILES string for a ligand with information about possible atom configurations of the ligand and converting the enriched information into a plurality of 3D models of the atom."), or permitting substitution of a first catalyst with a second catalyst sharing a threshold percentage of similar traits (Fig. 37, para [0185] "The module utilizes the knowledge graph 11.1 and data analysis engine 113 capabilities of the data platform 110, and in one embodiment is configured to identify ligands with certain properties based on three dimensional (3D) models 131 of known ligands and differentials of atom positions 132 in the latent space of the models after encoding by a 3D convolutional neural network (3D CNN), which is part of the data analysis engine 113. In one embodiment, the 3D model comprises a voxel image (volumetric, three dimensional pixel image) of the ligand. In cases where enrichment data is available, ligands may be identified by enriching the SMILES string for a ligand with information about possible atom configurations of the ligand and converting the enriched information into a plurality of 3D models of the atom.").
Regarding claim 8, NORVAISAS discloses the method of claim 1, wherein the predictive model comprises: a ranking function (Fig. 39, para [0202] "During molecule ranking 3921, a large number of molecules are ranked according to a set of criteria."), decision tree (para [0175] "FIG. 33 is an exemplary diagram. showing the application of a Monte Carlo Tree Search as applied to automated retrosynthesis using a hypergraph." a random forest of decision trees (para [0175] "FIG. 33 is an exemplary diagram showing the application of a Monte Carlo Tree Search as applied to automated retrosynthesis using a hypergraph."), or a shallow neural network.
Regarding claim 9, NORVAISAS discloses the method of claim 1, wherein the transformation of the preliminary data comprises at least one of: an indication of an operation or non-operation corresponding to the step of the at least one biochemical pathway (Fig. 30, para [0157] "The recursion stage comprises a recursion engine 3011 and a hypergraph search engine 3012. The recursion engine 3011 manages the recursive process of repeatedly retrosynthesizing larger molecules into smaller molecules that are precursors of that larger molecule. The process is recursive because each stage of the process is a one-to-many process, and there may be many possible chemical pathways to produce a drug candidate molecule."), or an indication of one or more results of performing the step of the at least one biochemical pathway (Fig. 30, para [0157] "The hypergraph search engine 3012 creates hypernodes from each set of retrosynthesized precursors, receives forward reaction prediction scoring for each node, assigns awards based on the scoring, selects a preferred node at each recursion, stage, and back-propagates the rewards to the root-node (the drug candidate molecule), thus creating a preferred chemical synthesis pathway from some set of preferred precursors (preferably commercially available) to the drug candidate molecule.").
Regarding claim 10, NORVAISAS discloses a method for providing explainable artificial intelligence (Fig. 37, para [0188] "The API module 3734 communicates with the artificial intelligence assistance system 3720 via API calls, sending assay results to the Al assistance system 3720, receiving instructions from the Al assistance system 3720, exchanging data and instructions with the other orchestration software components 3731-3733, and providing information to the AI assistance system 3720 about available resources such as chemical libraries and chemical synthesis capabilities."), comprising: receiving, at a computing device, a domain-specific dataset corresponding to a domain specific problem (Fig. 38, para [0189] "In this embodiment, the process involves iteratively performing physical assays on drug candidate molecules using an iterative physical testing system 3710, processing the results of the assays through an Al assistance system 3720 to generate new drug candidate molecules that may better fit a hypothesis or bioactivity goal, performing assays on the new drug candidate molecules, and repeating the process until some pre-defined parameter is met (e.g., a molecule is discovered with a better bioactivity than the drug candidate molecule, a molecule is discovered with precursors that meet certain criteria such as cost or availability, the drug candidate molecule and hypothesis are confirmed to a certain level of confidence, etc.)."); identifying, based on balancing an estimated minimum of information required to solve the domain-specific problem and an estimated amount of computation time required to solve the domain-specific problem, an optimal level of description for a domain corresponding to the domain-specific problem (Fig. 39, para [0197] "The molecule ranking stage 3921 uses libraries 3970 extracted from the scientific and academic literature, or from private or internal databases, to prioritize molecules for in vitro validation which are likely to show activity against the target of choice. The precise ranking algorithm will depend on parameters deemed to be important, a non-inclusive list of parameters may be maximization of selectivity, solubility, permeability, metabolic stability, and low cytochrome P450 (CYP) inhibition. In some embodiments, the molecule ranking stage 3921 uses a hierarchy of ranking methodologies designed to balance accuracy against computational cost."); generating, based on the domain-specific dataset, a plurality of candidate features (Fig. 7, para [0113] "FIG. 7 is a diagram illustrating an exemplary graph-based representation of molecules as simple relationships between atoms using a matrix of adjacencies 700, wherein atoms are represented as nodes and bonds between the atoms are represented as edges Further, the graph-based representation of a molecule can be stated in terms of two matrices, one for the node features (e.g., type of atom and its available bonds) and one for the edges (i.e., the bonds between the atoms). The combination of the nodes (atoms) and edges (bonds) represents the molecule. Each molecule represented in the matrix comprises a dimensionality and features that describe the type of bond between the atoms."), wherein generating the plurality of candidate features comprises: transforming one or more portions of the domain-specific dataset to correspond to the optimal level of description (Fig. 10, para [0214] "FIG. 10 is a diagram illustrating an exemplary architecture for prediction of molecule bioactivity using concatenation of outputs from a graph-based neural network which analyzes molecules and their known or suspected bloactivities with proteins and a sequence-based neural network which analyzes protein segments and their known or suspected bioactivities with molecules. In this architecture, in a first neural network processing stream, SMILES data 1010 for a plurality of molecules is transformed at a molecule graph construction stage 1013 into a graph-based representation wherein each molecule is represented as a graph comprising nodes and edges, wherein each node represents an atom and each edge represents a connection between atoms of the molecule."); and generating, based on the transforming and using a combinatorial algorithm, a plurality of combinations of features corresponding to the optimal level of description and included in the domain-specific dataset (Fig. 16, para [0136] "According to an embodiment of active example generation, a 3-dimensional convolutional neural network (3D CNN) is used in which atom-type densities are reconstructed using a sequence of 3D convolutional layers and dense layers. Since the output atom densities are fully differentiable with respect to the latent space, a trained variational autoericoder (VAE) 1606 may connect to a bioactivity-prediction module 1604 comprising a trained 3D-CNN model with the same kind of atom densities (as output by the autoencoder) as the features, and then optimize the latent space with respect to the bioactivity predictions against one or more receptors."); reducing, by performing one or more feature reduction operations, the plurality of candidate features into an optimized feature set (Fig. 14, para [0125] "The neural networks build a model from the training data. In the case of using an autoencoder (or a variational autoencoder), the encoder portion of the neural network reduces the dimensionality of the input molecules, learning a model from which the decoder portion recreates the input molecule. The significance of outputting the same molecule as the input is that the decoder may then be used as a generative function for new molecules."), wherein the one or more feature reduction operations comprise one or more of: selecting, based on based on one or more predictive scores indicating a likelihood of a given candidate feature predicting a solution to the domain-specific problem, a subset of candidate features (Fig. 30, para [0157] "The hypergraph search engine 3012 creates hypernodes from each set of retrosynthesized precursors, receives forward reaction prediction scoring for each node, assigns awards based on the scoring, selects a preferred node at each recursion stage, and back-propagates the rewards to the root node (the drug candidate molecule), thus creating a preferred chemical synthesis pathway from some set of preferred precursors (preferably commercially available) to the drug candidate molecule."), eliminating one or more redundant features (Fig. 31, para [0166] "As retrosynthesis is a non-determinative process, many possible precursors and reaction pathways may be predicted. Of these, some will be theoretically plausible but invalid, some will be valid but unfeasible or impractical, and some will be both valid and practical. At each recursive stage, validation of the predicted precursors may be performed to determine a likely validity of the predicted precursors and the chemical reaction by which the precursors are predicted to result in the molecule at the next higher recursion stage. For example, after the first retrosynthesis step 3101, two intermediate molecules are predicted 3121, 3122 which, given some chemical reaction, are expected to produce the drug candidate molecule 3111."), compressing the plurality of candidate features to combine features exceeding a threshold similarity score (Fig. 18, para [0145] "Attempting to minimize the loss 1804 directly with respect to the input features without such constraints may end up with atom densities that do not correspond to realistic molecules. To avoid this, one embodiment uses an autoencoder that encodes/decodes from/to the input representation of the bioactivity model, as the compression of chemical structures to a smaller latent space, which produces only valid molecules for any reasonable point in the latent space. Therefore, the optimization is performed with respect to the values of the latent vector, then the optima reached corresponds to real molecules."); training, based on the optimized feature set, a predictive model, wherein training the predictive model configures the predictive model to output solutions to domain-specific problems and representations of decisions made by the predictive model to output the solutions to domain specific problems (Fig. 10, para [0215] "At the training stage, the adjacency matrices 1011 and node features matrices 1012 for many molecules are input into the MPNN 1020 along with vector representations of known or suspected bioactivity interactions of each molecule with certain proteins. Based on the training data, the MPNN 1020 learns the characteristics of molecules and proteins that allow interactions and what the bioactivity associated with those interactions is. At the analysis stage, a drug candidate molecule is input into the MPNN 1020, and the output of the MPNN 1020 is a vector representation of that molecule's likely interactions with proteins and the likely bioactivity of those interactions."); outputting, using the predictive model, a solution for the domain-specific problem and a representation of decisions made by the predictive model to output the solution (Fig. 34, para [0179] "A check is made to determine whether there is a set of precursors (reactants) that can produce the drug candidate molecule, all of which precursors are commercially available 3410. If there is no such set, the process is repeated recursively from step 3401 until such'a set is found. If there is such a set, recursion is ended 3411 and the preferred chemical pathway and its commercially available precursors are returned as an output."); and applying, based on the representation of decisions made by the predictive model to output the solution, one or more decisions made by the predictive model to at least one additional domain-specific problem (Fig. 37, para [0180] FIG. 37 is an exemplary system architecture for a system for feedback-driven automated drug discovery. In this embodiment, the system 3700 combines an iterative physical testing system 3710 of drug candidate molecules for agreement with hypotheses with an artificial intelligence/machine learning system 3720 configured to generate predictions of additional drug candidate molecules for testing which may better match a given hypothesis.").
Regarding claim 11, NORVAISAS discloses the method of claim 10, wherein the applying the one or more decisions comprises: revising, through an iterative loop and based on the solution to the domain-specific problem, the predictive model (Fig. 38, para [0189] FIG. 38 is a high-level process flow diagram illustrating exemplary operation of a system for feedback-driven automated drug discovery. In this embodiment, the process involves iteratively performing physical assays on drug candidate molecules using an Iterative physical testing system 3710, processing the results of the assays through an AI assistance system 3720 to generate new drug candidate molecules that may better fit a hypothesis or bioactivity goal. wherein the iterative loop comprises: receiving, based on outputting the representation of decisions made by the predictive model, feedback information corresponding to one or more human experts in the domain corresponding to the domain-specific problem (Fig. 1, para [0099] "The data platform 110 in this embodiment comprises a knowledge graph 111, an exploratory drug analysis (EDA) interface 112, a data analysis engine 113, a data extraction engine 114, and web crawler/database crawler 115. The crawler 115 searches for and retrieves medical information such as published medical literature, clinical trials, dissertations, conference papers, and databases of known pharmaceuticals and their effects."); updating, based on the feedback information, the predictive model (Fig. 12, para [0228]); and repeating, for one or more additional domain-specific problems, outputting a solution, the receiving feedback information, and the updating for each additional domain specific problem (Fig. 39, para [0196]).
Regarding claim 12, NORVAISAS discloses the method of claim 10, wherein the optimal level of description corresponds to descriptive information that has the following properties: the descriptive information indicates distinctions between candidate features corresponding to potential predictive values (Fig. 18, para [0146] "In the case of a 3D CNN bioactivity model, the 3D CNN autoencoder would thus form the input of the combined trained models. This embodiment allows both differentiable representations which also have an easily decodable many-to-one mapping to real moiecules since the latent space encodes the 3D structure of a particular rotation and translation of a particular conformation of a certain molecule, therefore many latent points can decode to the same molecule but with different arrangements in space. The derivative of the loss with respect to the atom density in a voxel allows for backpropagation of the gradients all the way through to the latent space, where optimization may be performed on the model output(s) 1805, 1806 with respect to, not the weights, but the latent vector values."); a size of the descriptive information is below a threshold storage capacity (Fig. 28, para [0155] "The training data presents a choice of a threshold bracket 2830. The threshold bracket is a trade-off between the average information contained in each datapoint, and the sheer quantity of data, assuming that datapoints with more extreme inactive/active IC50 values are indeed more typical of the kind of interactions that determine whether or not a protein-ligand pair is active or inactive. In the case of the 3D-model, using the dataset with no threshold performs consistently better across most metrics."); and the descriptive information comprises all information identified as necessary to predict a solution to the domain-specific problem (Figs. 11A, -B).
Regarding claim 13, NORVAISAS discloses the method of claim 10, wherein the predictive model comprises: a ranking function (Fig. 39, para [0202] "During molecule ranking 3921, a large number of molecules are ranked according to a set of criteria:"), a decision tree (para [0175] "FIG. 33 is an exemplary diagram showing the application of a Monte Carlo Tree Search as applied to automated retrosynthesis using a hypergraph."), a random forest of decision trees (para [0175] "FIG. 33 is an exemplary diagram showing the application of a Monte Carlo Tree Search as applied to automated retrosynthesis using a hypergraph."), or a shallow neural network.
Regarding claim 14, NORVAISAS discloses the method of claim 10, wherein eliminating the one or more redundant features comprises eliminating, based on a comparative indicator of the likelihood of a first candidate feature, of the plurality of candidate features, to be a false positive exceeding a threshold, the first candidate feature (Fig. 28, para [0155] "The training data presents a choice of a threshold bracket 2830. The threshold bracket is a trade-off between the average information contained in each datapoint, and the sheer quantity of data, assuming that datapoints with more extreme inactive/active IC50 values are indeed more typical of the kind of interactions that determine whether or not a protein-ligand pair is active or inactive. in the case of the 3D-model, using the dataset with no threshold performs consistently-better across most metrics. The channels used for the data set are hydrophobic, hydrogen-bond donor or acceptor, aromatic, positive or negative ionizable, metallic and total excluded volume, Regardless of the choice of threshold, the data is then used to train a 3D -CNN to know the classification of a molecule regarding activation and pose propriety 2840.").
Regarding claim 15, NORVAISAS discloses the method of claim 10, wherein the compressing comprises combining, based on comparing one or more similarity scores corresponding to respective features of the plurality of candidate features, two or more candidate features exceeding the threshold similarity score (Fig. 31, para [0166] "At each recursive stage, validation of the predicted precursors may be performed to determine a likely validity of the predicted precursors and the chemical reaction by which the precursors are predicted to result in the molecule at the next higher recursion stage. For example, after the first retrosynthesis step 3101, two intermediate molecules are predicted 3121, 3122 which, given some chemical reaction, are expected to produce the drug candidate molecule 3111. Forward reaction prediction scoring 3123, 3135-3136 may be used to perform this validation. In one embodiment, the forward reaction of the two intermediate molecules 3121, 3122 may be predicted by a machine learning algorithm such as an attention-based transformer.").
Regarding claim 16, NORVAISAS discloses a computing system comprising: one or more processors (para [0257]); and memory storing computer executable instructions that, when executed by the one or more processors (para [0257]). cause the computing system to: receive a domain-specific dataset corresponding to a domain-specific problem (Fig. 38, para [0189] "In this embodiment, the process involves iteratively performing physical assays on drug candidate molecules using an iterative physical testing system 3710, processing the results of the assays through an Al assistance system 3720 to generate new drug candidate molecules that may better fit a hypothesis or bioactivity goal, performing assays on the new drug candidate molecules, and repeating the process until some pre-defined parameter is met (e.g., a molecule is discovered with a better bioactivity than the drug candidate molecule, a molecule is discovered with precursors that meet certain criteria such as cost or availability, the drug candidate molecule and hypothesis are confirmed to a certain level of confidence, etc.)."); identify, based on balancing an estimated minimum of information required to solve the domain-specific problem and an estimated amount of computation time required to solve the domain-specific problem, an optimal level of description for a domain corresponding to the domain-specific problem (Fig. 39, para [0197] "The molecule ranking stage 3921 uses libraries 3970 extracted from the scientific and academic literature, or from private or internal databases, to prioritize molecules for in vitro validation which are likely to show activity against the target of choice. The precise ranking algorithm will depend on parameters deemed to be important, a non-inclusive list of parameters may be maximization of selectivity, solubility, permeability, metabolic stability, and low cytochrome P450 (CYP) inhibition. In some embodiments, the molecule ranking stage 3921 uses a hierarchy of ranking methodologies designed to balance accuracy against computational cost."); generate, based on the domain-specific dataset, a plurality of candidate features (Fig. 7, para [0113] "FIG. 7 is a diagram illustrating an exemplary graph-based representation of molecules as simple relationships between atoms using a matrix of adjacencies 700, wherein atoms are represented as nodes and bonds between the atoms are represented as edges Further, the graph-based representation of a molecule can be stated in terms of two matrices, one for the node features (e.g., type of atom and its available bonds) and one for the edges (i.e., the bonds between the atoms). The combination of the nodes (atoms) and edges (bonds) represents the molecule. Each molecule represented in the matrix comprises a dimensionality and features that describe the type of bond between the atoms."), wherein generating the plurality of candidate features comprises: transforming one or more portions of the domain-specific dataset to correspond to the optimal level of description (Fig. 10, para [0214] "FIG. 10 is a diagram illustrating an exemplary architecture for prediction of molecule bioactivity using concatenation of outputs from a graph-based neural network which analyzes molecules and their known or suspected bioactivities with proteins and a sequence-based neural network which analyzes protein segments and their known or suspected-bioactivities with molecules. In this architecture, in a first neural network processing stream, SMILES data 1010 for a plurality of molecules is transformed at a molecule graph construction stage 1013 into a graph-based representation wherein each molecule is represented as a graph comprising nodes and edges, wherein each node represents an atom and each edge represents a connection between atoms of the molecule."); and generating, based on the transforming and using a combinatorial algorithm, a plurality of combinations of features corresponding to the optimal level of description and included in the domain-specific dataset (Fig. 16, para [0136] "According to an embodiment of active example generation, a 3-dimensional convolutional neural network (3D CNN) is used in which atom-type densities are reconstructed using a sequence of 3D convolutional layers and dense layers. Since the output atom densities are fully differentiable with respect to the latent space, a trained variational autoencoder (VAE) 1606 may connect to a bioactivity-prediction module 1604 comprising a trained 3D-CNN model with the same kind of atom densities (as output by the autoencoder) as the features, and then optimize the latent space with respect to the bioactivity predictions against one or more receptors."); reduce, by performing one or more feature reduction operations, the plurality of candidate features into an optimized feature set (Fig. 14, para [0125] "The neural networks build a model from the training data. In the case of using an autoencoder (or a variational autoencoder), the encoder portion of the neural network reduces the dimensionality of the input molecules, learning a model from which the decoder portion recreates the input molecule. The significance of outputting the same molecule as the input is that the decoder may then be used as a generative function for new molecules."), wherein the one or more feature reduction operations comprise one or more of: selecting, based on based on one or more predictive scores indicating a likelihood of a given candidate feature predicting a solution to the domain-specific problem, a subset of candidate features (Fig. 30, para [0157] "The hypergraph search engine 3012 creates hypernodes from each set of retrosynthesized precursors, receives forward reaction prediction scoring for each node, assigns awards based on the scoring, selects a preferred node at each recursion stage, and back-propagates the rewards to the root node (the drug candidate molecule), thus creating a preferred chemical synthesis pathway from some set of preferred precursors (preferably commercially available) to the drug candidate molecule."), eliminating one or more redundant features (Fig. 31, para [0166] "As retrosynthesis is a non-determinative process, many possible precursors and reaction pathways may be predicted. Of these, some will be theoretically plausible but invalid, some will be valid but unfeasible or impractical, and some will be both valid and practical. At each recursive stage, validation of the predicted precursors may be performed to determine a likely validity of the predicted precursors and the chemical reaction by which the precursors are predicted to result in the molecule at the next higher recursion stage. For example, after the first retrosynthesis step 3101, two intermediate molecules are predicted 3121, 3122 which, given some chemical reaction, are expected to produce the drug candidate molecule 3111."), or compressing the plurality of candidate features to combine features exceeding a threshold similarity score (Fig. 18, para [0145] "Attempting to minimize the loss 1804 directly with respect to the input features without such constraints may end up with atom densities that do not correspond to realistic molecules. To avoid this, one embodiment uses an autoencoder that encodes/decodes from/to the input representation of the bioactivity model, as the compression of chemical structures to a smaller latent space, which produces only valid molecules for any reasonable point in the latent space. Therefore, the optimization is performed with respect to the values of the latent vector, then the optima reached corresponds to real molecules."); train, based on the optimized feature set, a predictive model, wherein training the predictive model configures the predictive model to output solutions to domain-specific problems and representations of decisions made by the predictive model to output the solutions to domain-specific problems (Fig. 10, para [0215] "At the training stage, the adjacency matrices 1011 and node features matrices 1012 for many molecules are input into the MPNN 1020 along with vector representations of known or suspected bioactivity interactions of each molecule with certain proteins. Based on the training data, the MPNN 1020 learns the characteristics of molecules and proteins that allow interactions and what the bioactivity associated with those interactions is. At the analysis stage, a drug candidate molecule is input into the MPNN 1020, and the output of the MPNN 1020 is a vector representation of that molecule's likely interactions with proteins and the likely bioactivity of those interactions."); output, using the predictive model, a solution for the domain-specific problem and a representation of decisions made by the predictive model to output the solution (Fig. 34, para [0179] "A check is made to determine whether there is a set of precursors (reactants) that can produce the candidate molecule, all of which precursors are commercially available 3410. If there is no such set, the process is repeated recursively from step 3401 until such a set is found. If there is such a set, recursion is ended 3411 and the preferred chemical pathway and its commercially available precursors are returned as an output."); and apply, based on the representation of decisions made by the predictive model to output the solution, one or more decisions made by the predictive model to at least one additional domain-specific problem (Fig. 37, para [0180] FIG. 37 is an exemplary system architecture for a system for feedback-driven automated drug discovery. In this embodiment, the system 3700 combines an iterative physical testing system 3710 of drug candidate molecules for agreement with hypotheses with an artificial intelligence/machine learning system 3720 configured to generate predictions of additional drug candidate molecules for testing which may better match a given hypothesis.").
Regarding claim 17, NORVAISAS discloses the computing system of claim 16, wherein the instructions, when executed, configure the computing system to apply the one or more decisions by: revising, through an iterative loop and based on the solution to the domain-specific problem, the predictive model (Fig. 38, para [0189] FIG. 38 is a high-level process flow diagram illustrating exemplary operation of a system for feedback-driven automated drug discovery. In this embodiment, the process involves iteratively performing physical assays on drug candidate molecules using an iterative physical testing system 3710, processing the results of the assays through an Al assistance system 3720 to generate new drug candidate molecules that may better fit a hypothesis or bioactivity goal..."), wherein the iterative loop comprises: receiving, based on outputting the representation of decisions made by the predictive model, feedback information corresponding to one or more human experts in the domain corresponding to the domain-specific problem (Fig. 1, para [0099] "The data platform 110 in this embodiment comprises a knowledge graph 111, an exploratory drug analysis (EDA) interface 112, a data analysis engine 113, a data extraction engine 114, and web crawler/database crawler 115. The crawler 115 searches for and retrieves medical information such as published medical literature, clinical trials, dissertations, conference papers, and databases of known pharmaceuticals and their effects."); updating, based on the feedback information, the predictive model (Fig. 12, para [0228]); and repeating, for one or more additional domain-specific problems, outputting a solution, the receiving feedback information, and the updating for each additional domain specific problem (Fig. 39, para [0196]).
Regarding claim 18, NORVAISAS discloses the computing system of claim 16, wherein the optimal level of description corresponds to descriptive information that has the following properties: the descriptive information indicates distinctions between candidáte features corresponding to potential predictive values (Fig. 18, para [0146] "In the case of a 3D CNN bioactivity model, the 3D CNN autoencoder would thus form the input of the combined trained models. This embodiment allows both differentiable representations which also have an easily decodable many-to-one mapping to real molecules since the latent space encodes the 3D structure of a particular rotation and translation of a particular conformation of a certain molecule, therefore many latent points can decode to the same molecule but with different arrangements in space. The derivative of the loss with respect to the atom density in a voxel allows for backpropagation of the gradients all the way through to the latent space, where optimization may be performed on the model output(s) 1805, 1806 with respect to, not the weights, but the latent vector values,"); a size of the descriptive information is below a threshold storage capacity (Fig. 28, para [0155] "The training data presents a choice of a threshold bracket 2830. The threshold bracket is a trade-off between the average information contained in each datapoint, and the sheer quantity of data, assuming that datapoints with more extreme inactive/active IC50 values are indeed more typical of the kind of interactions that determine whether or not a protein-ligand pair is active or inactive. In the case of the 3D-model, using the dataset with no threshold performs consistently better across most metrics."); and the descriptive information comprises all information identified as necessary to predict a, solution to the domain-specific problem (Figs. 11A-B).
Regarding claim 19, NORVAISAS discloses the computing system of claim 16, wherein the predictive model comprises: a ranking function (Fig. 39, para [0202] "During molecule ranking 3921, a large number of molecules are ranked according to a set of criteria."), a decision tree (para [0175] "FIG. 33 is an exemplary diagram showing the application of a Monte Carlo Tree Search as applied to automated retrosynthesis using a hypergraph."), a random forest of decision trees (para [0175] "FIG. 33 is an exemplary diagram showing the application of a Monte Carlo Tree Search as applied to automated retrosynthesis using a hypergraph."), or a shallow neural network.
Regarding claim 20, NORVAISAS discloses the computing system of claim 16, wherein the instructions, when executed, configure the computing system to compress the plurality of candidate features by combining, based on comparing one or more similarity scores corresponding to respective features of the plurality of candidate features, two or more candidate features exceeding the threshold similarity score (Fig. 31, para [0166] "At each recursive stage, validation of the predicted precursors may be performed to determine a likely validity of the predicted precursors and the chemical reaction by which the precursors are predicted to result in the molecule at the next higher recursion stage. For example, after the first retrosynthesis step 3101, two intermediate molecules are predicted 3121, 3122 which, given some chemical reaction, are expected to produce the drug candidate molecule 3111. Forward reaction prediction scoring 3123, 3135-3136 may be used to perform this validation. In one embodiment, the forward reaction of the two intermediate molecules 3121, 3122 may be predicted by a machine learning algorithm such as an attention-based transformer.").
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARK D FEATHERSTONE whose telephone number is (571)270-3750. The examiner can normally be reached Monday-Friday 9:00AM - 5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John Cottingham can be reached at 571-272-1400. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARK D FEATHERSTONE/Supervisory Patent Examiner, Art Unit 2111