DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Claims 1-4, 6-7, 9, 11-13, 15, 18-22, 25-26, and 30-31 are pending.
Priority
This application claims benefit of application no. 63/038,691, filed 06/12/2020 and application no. 63/306,958, filed 02/04/2022 and application 63/310,453, filed 02/15/2022. The instant application has the effective filing date of 15 February 2022.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/24/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement has been considered by the examiner.
Drawings
The drawings, submitted on 02/03/2023, are accepted by the examiner.
Claim Rejections – 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-4, 6, 9, 11-13, 15, 18-22, 25-26, and 30 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention for the following reasons.
The term “sufficiently fewer” in claims 1 and 30 is a relative term which renders the claim indefinite. The term is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. All dependent claims that do not remedy the issue are similarly rejected.
The term “at least about” in claim 9 is a relative term which renders the claim indefinite. The term “at least about” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-4, 6-7, 9, 11-13, 15, 18-22, 25-26, and 30-31 are rejected under U.S.C 101 because the claimed invention is directed to abstract ideas without significantly more as detailed in the analysis below.
Eligibility Step 1: Subject matter eligibility evaluation in accordance with MPEP § 2106:
Claims 1-4, 6-7, 9, 11-13, 15, 18-22, 25-26, and 30 are directed to a statutory category (method).
Claim 31 is directed to a statutory category (system).
Therefore, in accordance with MPEP § 2106.03, all claims have patent eligible subject matter.
[Eligibility Step 1: YES]
Eligibility Step 2A: This step determines whether a claim is directed to a judicial exception in accordance with MPEP § 2106.
Eligibility Step 2A -- Prong One: Limitations are analyzed to determine if the claims recite any concepts that could equate to a judicial exception (i.e. abstract idea, law of nature, or natural phenomenon). Possible judicial exceptions are explored below.
Recitations of Judicial Exceptions:
Claims 1 and 30: optimizing at least one loss function based at least in part on the plurality of latent descriptors and the plurality of reconstructions by adjusting the at least one parameter, such that the neural network learns a latent space comprising a denoised embedding for the plurality of polyamino acid descriptors. (mathematical concept, mental process)
Claim 2: wherein the output layer outputs a plurality of parameters for a probability distribution. (mathematical concept)
Claim 3: wherein the probability distribution is a zero inflated distribution. (mathematical concept)
Claim 4: wherein the zero inflated distribution is a zero inflated negative binomial distribution. (mathematical concept)
Claim 12: wherein the posterior distribution comprises a higher kurtosis than a normal distribution (mathematical concept)
Claim 13: wherein the at least one loss function comprises a Kullbeck-Leibler divergence loss function based at least in part on a difference between a sum of posterior distributions parameterized by the plurality of parameters and a prior distribution. (mathematical concept)
Claim 15: wherein the prior distribution comprises a higher kurtosis than a normal distribution. (mathematical concept)
Claim 22: classifying at least a first set of latent descriptors from a second set of latent descriptors, wherein the first set of latent descriptors is associated with a first biological state and the second set of latent descriptors is associated with a second biological state. (mental process)
Claim 31: (b) generating, in a latent space, a latent descriptor based at least in part on the polyamino acid descriptor, and wherein the latent descriptor comprises sufficiently fewer dimensions than the polyamino acid descriptor such that at least a portion of information in the polyamino acid descriptor is lost in the latent descriptor; (mathematical concept, mental process)
(c) determining, based at least in part on the latent descriptor, the biological state associated with the polyamino acid descriptor. (mental process)
Step 2A – Prong One Analysis:
Limitations such as calculating loss functions, distributing data, sums, Kullback-Leibler equations, or encoding data into latent spaces read on transformations and organizations of data via mathematical calculations, formulas and/or relationships. Such recitations fall under the mathematical concepts grouping of abstract ideas.
Limitations such as adjusting parameters or classifying states of data read on observations and mental determinations of data that can be performed within the human mind and with pen/paper. Such recitations fall under the mental process grouping of abstract ideas.
Eligibility Step 2A – Prong Two: A claim that integrates a judicial exception into a practical application will apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception. If the claim contains no additional claim elements beyond the abstract idea, the claim fails to integrate the abstract idea into a practical application (MPEP 2106.04(d)). Additional elements are recited, categorized, and analyzed below.
Data Gathering/Outputting Elements:
Claim 1: (a) providing a neural network comprising: i) an input layer configured to receive at least a polyamino acid descriptor; ii) a latent layer configured to output at least a latent descriptor, wherein the latent layer is connected to the input layer, and wherein the latent descriptor comprises sufficiently fewer dimensions than the polyamino acid descriptor such that at least a portion of information in the polyamino acid descriptor is filtered in the latent descriptor; iii) an output layer configured to output at least a reconstruction of the polyamino acid descriptor, wherein the output layer is connected to the latent layer; and iv) at least one parameter;
(b) providing training data comprising a plurality of polyamino acid descriptors, wherein the plurality of polyamino acid descriptors comprises at least one value for a polyamino acid in association with a given assay method; and
(c) training the neural network, by (i) inputting at least the plurality of polyamino acid descriptors at the input layer of the neural network, (ii) outputting a plurality of latent descriptors at the latent layer and a plurality of reconstructions at the output layer
Claim 6: wherein the polyamino acid descriptor comprises at least 100 dimensions.
Claim 7: wherein the latent descriptor comprises at most about 50% of the number of dimensions in the polyamino acid descriptor.
Claim 9: wherein at least about 10% of values in the plurality of polyamino acid descriptors in the training data are zero.
Claim 11: wherein the latent layer outputs a plurality of parameters for a posterior distribution.
Claim 25: wherein the at least one polyamino acid descriptor comprises an identification of at least one protein or protein group.
Claim 26: wherein the at least one polyamino acid descriptor comprises an identification of at least one peptide.
Claim 31: (a) receiving the polyamino acid descriptor comprising at least one dimension representing a polyamino acid association with a given assay method
Assay Elements:
Claim 18: wherein the given assay method comprises contacting a plurality of biomolecules with a given surface.
Claim 19: wherein the given surface is a surface of a particle.
Claim 20: wherein the given assay method comprises (i) performing mass spectrometry on cleaved derivatives of the plurality of biomolecules to obtain a plurality of peptide spectral signals and (ii) processing the plurality of peptide spectral signals to obtain a plurality of peptide identifications, wherein the plurality of polyamino acid descriptors comprises the plurality of peptide identifications.
Claim 21: wherein the given assay method comprises (i) performing mass spectrometry on cleaved derivatives of the plurality of biomolecules to obtain a plurality peptide spectral signals (ii) processing the plurality of peptide spectral signals to obtain a plurality of peptide identifications and (iii) processing the plurality of peptide identifications to obtain a plurality of intensities for plurality of protein or protein group identification, wherein the plurality of polyamino acid descriptors comprises the plurality of protein or protein group identifications.
Step 2A- Prong 2 Analysis:
The data gathering and outputting elements are drawn to receiving data, inputting it into layers of a neural network, and outputting latent descriptors and other secondary data. The elements do not appear to incorporate the elements into practical application per MPEP 2105.06 (g), nor recite a particular machine, transformation, or improvement to technology (MPEP 2106.04(d)).
The assay elements merely describe the type of data input into the neural network layers by limiting the experiment and measurements taken from the sample prior to input, which do not integrate the judicial elements into practical application, as it equates to insignificant extra-solution activity per MPEP 2105.06 (g).
As such, the additional elements when evaluated separately and in the context of a whole claimed invention, do not integrate the judicial elements into practical application.
Eligibility Step 2B: Claim elements are probed for inventive concept equating to significantly more than the judicial exception (MPEP 2106.04(II)). Additional elements are recited and categorized below.
Step 2B Analysis:
The data gathering elements drawn to inputting polyamino acid descriptors (page 8, column 2), that includes assay data (page 10, table 1) into latent layers (page 3, fig. 2), to obtain latent representations, with less dimensions (page 1, column 1); outputting reconstructed data (page 5, column 1); and accounting for zeroes within the training data (page 4, column 1) are found well-understood, routine, and conventional per Kopf et al. (Patterns 2; 2021), which reviews latent representation learning in biology and translational medicine.
The assay elements are found routine, well-understood, and conventional via Cunningham et al. (Frontiers in Biology; Vol. 7: 4; 2012), which reviews mass spectrometry-based proteomics and peptidomics; and teaches using the spectral processing techniques for peptide (page 8, column 2) and protein identifications (page 9, column 1).
As such, the additional elements are further found to lack inventive concept.
[Eligibility Step 2B: NO]
As such, claims 1-4, 6-7, 9, 11-13, 15, 18-22, 25-26, 30-31 are directed to judicial exceptions without significantly more and are rejected under 35 U.S.C 101.
Claim Rejections – 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-4, 11, 13, 18-19, 22, 25, and 30-31 are rejected under 35 U.S.C. 103 as being unpatentable over Colby et al. (Analytical Chemistry; Vol. 92; 2020) in view of Kopf et al. (Patterns 2; 2021).
Claims 1 and 30 are directed to methods and systems that store a neural network that includes i) an input layer that receives a polyamino acid descriptor; ii) a latent layer that can output a latent descriptor, has sufficiently fewer dimensions than the polyamino acid descriptor, such that portion of information in the polyamino acid descriptor is filtered in the latent descriptor; the latent layer is connected to the input layer; iii) an output layer that outputs a reconstruction of the polyamino acid descriptor, wherein the output layer is connected to the latent layer; and iv) at least one parameter.
Colby et al. describes DarkChem, a deep learning framework that generates in silico chemical property libraries and candidate molecules for small molecule identification.
Colby et al. teaches initially training DarkChem on ∼53 million SMILES inputs from PubChem (page 4, fig. 3); and using the SMILES input encoder and SMILES character decoding layer (page 2, column 2) to down select data sets into molecules containing only CHNOPS atoms and SMILES string lengths of 100 characters or less (page 2, column 2).
Colby et al. teaches the overall DarkChem architecture consists of: (1) the SMILES input encoder and convolutional layers, connected to (2) a latent space, which holds the vector representation of molecular structure, connected to (3) a decoder, consisting of convolutional layers and a SMILES character decoding layer, and (4) a property prediction layer (page 2, column 2), which yields outputs and is connected to the latent representation (page 3, fig. 1).
Colby et al. further teaches conducting a sweep over selected parameters, including latent dimension size, number of filters, kernel size, noise parameter epsilon, dropout, and embedding dimension, wherein each parameter was varied and performed one at a time (page 5, column 1).
Claims 1 and 30 are further directed to b) providing training data that includes polyamino acid descriptors, wherein descriptors comprise at least one value for a polyamino acid in association with a given assay method; (c) training the neural network, by (i) inputting the polyamino acid descriptors at the input layer of the neural network, and (ii) outputting latent descriptors at the latent layer, and reconstructions at the output layer.
Colby et al. teaches the latent space simultaneously encodes a numerical representation of structure and associated chemical properties (page 6, column 1), obtained from experimental measurements, such as mass spectra (m/z) and collision cross section (CCS) (page 1, abstract); using weights from this network to seed the next, by training on the ∼600 000 in silico data set, with m/z and in silico CCS labels (page 4, fig. 3); seeding the final training step with the further trained weights, using ∼500 inputs with m/z and experimental CCS (page 4, fig. 3); and yielding a mean per-character absolute difference between the input and predicted SMILES output sequences as a reconstruction accuracy of 78.5% (page 5, column 2).
Claims 1 and 30 are further directed to iii) optimizing at least one loss function based at least in part on the plurality of latent descriptors and the plurality of reconstructions by adjusting the at least one parameter, such that the neural network learns a latent space comprising a denoised embedding for the plurality of polyamino acid descriptors.
Colby et al. teaches using an objective function to evaluate property prediction loss as the mean absolute percent error between the predicted and target property vector (page 3, column 2); representing VAE loss by categorical cross entropy and KL-divergence losses (page 4, column 1); and adding a Kullback–Leibler divergence term to the objective function evaluation in order to penalize departures from a mean of 0 and a variance of 1, and ensure normally distributed noise was added to the latent representation during training, scaled by hyperparameter epsilon (page 3, column 1).
Colby et al. further teaches determining CCS values from SMILES found in the PubChem and in silico data sets through the trained DarkChem network to generate CCS for [M + H]+, [M – H]−, and [M + Na]+ adducts, without the normally distributed noise hyperparameter (epsilon) added to the latent representation during training (page 5, column 1).
Therefore Colby et al. trains a neural network on polyamino acid descriptors, outputs reconstructions, and optimizes a loss function, based on the latent descriptor by adjusting the epsilon hyperparameter and adding noise to an embedding.
Claim 2 is directed to the output layer outputting multiple parameters for a probability distribution.
Colby et al. teaches the model’s left side terms are the Kullback–Leibler divergence, DKL; expected and observed probabilities qϕ and pϕ, respectively, over a set of observed variables, x, and a set of latent variables, z, with joint distribution p(z, x) (page 3, column 1).
Claim 18 is directed to the assay method including contacting biomolecules with a given surface. Claim 19 is directed to the surface being a particle surface.
Colby et al. teaches CCS measures an ionized molecule’s effective interaction surface with a buffer gas from ion mobility spectroscopy separations (page 2, column 1).
Colby et al. does not teach the latent descriptor having sufficiently fewer dimensions than the polyamino acid descriptor, (claims 1 and 30); or the latent space including a denoised embedding for the amino acid descriptors (claims 1 and 30).
Kopf et al. describes latent representation learning (LRL) in biology and translational medicine.
With regards to claims 1 and 30 and the latent descriptor having sufficiently fewer dimensions than the polyamino acid descriptor, Kopf et al. teaches the latent representation is usually a compressed form of the empirical measurements; and consists of fewer latent variables than the dimensionality of the measurements (page 1, column 1).
With regards to claims 1 and 30 and the latent space including a denoised embedding for the amino acid descriptors, Kopf et al. teaches an overview of recent and forthcoming LRL approaches include classical factor analysis (FA) models, Gaussian process (GP) LVM and, deep model architectures, such as autoencoders (AEs), variational AEs (VAEs), or generative adversarial networks (page 2, column 1).
Kopf et al. teaches GP-LVM defines a Gaussian process prior with ε (epsilon) a noise term usually normally distributed around zero (page 4, footnote B); can be used to analyze single-cell qPCR expression data; and embed the dynamics of the data in a latent representation for classification (page 5, column 1). Kopf et al teaches VAE, where the encoder predicts the parameter’s mean and SD of a normally distributed latent representation, models different distributions of the latent representation (page 4; footnote D); can be used for batch correction, visualization, clustering, and differential expression analysis (page 6, column 2); and uses the denoised reconstructed output of the VAEs as input for RNNs for data generation purposes (page 6, column 2).
Claim 3 is directed to the probability distribution being a zero inflated distribution. Claim 4 is directed to the zero inflated distribution being a zero inflated negative binomial distribution.
Regarding claims 3-4, Kopf et al. teaches a scalable VAE architecture, Single-cell VI (scVI), assumes a zero-inflated negative binomial distribution conditioned on batch annotations for decoding the single-cell data and accounting for dropout effect (page 6, column 2).
Claim 11 is directed to the latent layer outputting multiple parameters for a posterior distribution.
Kopf et al. teaches GP-LVM models can be trained via optimizing the log likelihood function and finding its maximum, or in a more efficient way via VI, where the posterior distribution of the model is approximated (page 5, column 1); and a VAE case in which, the encoder defines the approximation of the true posterior distribution, and the neural network parameters are the variational parameters for interference (page 6, column 1).
Claim 13 is directed to at least one loss function comprising a Kullbeck-Leibler divergence function based at least in part on a difference between a sum of posterior distributions, parameterized by the plurality of parameters and a prior distribution.
Colby et al. teaches representing the VAE loss by categorical cross entropy and Kullback- Leibler divergence loss (page 4, column 1).
Kopf et al. teaches with VAE, the encoder 𝑞𝜑(𝒛|𝒙) defines the approximation of the true posterior distribution 𝑝(𝒛|𝒙), which is in general intractable; the neural network parameters 𝜑 are the variational parameters for the inference; the distributional assumption on the latent space is incorporated in the model via the prior distribution 𝑝(𝒛) (page 6, column 1); and optimizing the VAE model requires maximizing the evidence lower bound where log𝑝𝜽(𝒙|𝒛) is the log likelihood of the reconstructed data and 𝐷𝐾𝐿 is the Kullback-Leibler divergence, which ensures the distribution of the latent representation follow 𝑝(𝒛) when optimized (page 6, column 1).
Claim 22 is directed to classifying a first set of latent descriptors, associated with a first biological state, from a second set of latent descriptors, associated with a second biological state.
Claim 31 is directed to (a) receiving the polyamino acid descriptor comprising at least one dimension representing a polyamino acid association with a given assay method; (b) generating, in a latent space, a latent descriptor based at least in part on the polyamino acid descriptor, and wherein the latent descriptor comprises sufficiently fewer dimensions than the polyamino acid descriptor such that at least a portion of information in the polyamino acid descriptor is lost in the latent descriptor; and determining, based at least in part on the latent descriptor, the biological state associated with the polyamino acid descriptor.
Regarding claims 22 and 31, Kopf et al. teaches applying LVM to single-cell measurements of tissues originating from different cancer types to infer cancer specific subpopulations (page 11, column 1), in which the latent representation is usually a compressed form of the empirical measurements; and consists of fewer latent variables than the dimensionality of the measurements (page 1, column 1).
Kopf et al. further teaches tissue- or organism-level phenotypes, such as disease states, are frequently associated with, specific cell types or subpopulations and can inferred by LVM approaches (page 11, column 1); in which gene and protein expression play a central role in inferring the latent variable explaining the phenotype variation of interest, since they define the subpopulations, and constitute the basis of the mechanism conferring their association with the phenotype (page 11, column 1). Kopf et al. further teaches using latent variables to capture the variation along time or developmental stages of biological processes using the snapshot single-cell expression data of genes or proteins (page 11, column 1).
Claim 25 is directed to the at least one polyamino acid descriptor comprising an identification of at least one protein or protein group.
Kopf et al. teaches citing LVM approaches and their applications, in which the data type includes that of a protein sequence or protein structure data (page 9, table 1).
Kopf et al. further teaches such latent variable modeling (LVM) methods provide a solution to challenges that arise due to the variable length of the DNA, RNA, or protein information captured in sequential data, such as nucleotide or amino acid sequences (page 8, column 2)(page 10, column 1); and several deep learning models take SMILES or molecular graphs as input (page 9, table 1) for data representation, dimensionality representation, and representation learning (page 9, table 1).
Therefore Kopf et al. teaches applying the known technique of decreasing dimensionality among the latent layers yields predictable results within a neural network environment, applicable to Colby et al. Kopf et al. further teaches a number of identified, analogous, and predictable solutions to the latent representation learning problem within a polyamino acid framework that includes additive noise reconstruction, identical to the method of Colby et al.; and a VAE framework that results in a denoised embedding and a zero-inflated negative binomial distribution.
As such, it would be obvious to one of ordinary skill in the art to substitute the GP-LVM framework of Colby et al., with the denoising and zero-inflated binomial distribution VAE techniques, taught by Kopf et al. in order to accomplish deep latent representation learning functions, known in the art, with a reasonable expectation of success in polyamino acid predictions.
Claims 6-7 are rejected under 35 U.S.C. 103 as being unpatentable over Colby et al. (Analytical Chemistry; Vol. 92; 2020) in view of Kopf et al. (Patterns 2; 2021), as applied to claims 1-4, 11, 13, 18-19, 22, 25, and 30-31 above, and in further view of Schoenholz et al. (arXiv:1808.06576v2; 2018).
Colby et al. in view of Kopf et al. teach a neural network that inputs a dimensional polyamino acid descriptor, encodes it into a latent layer with fewer dimensions, and uses an optimized loss function, based on prior and posterior probabilities, to output polyamino acid reconstructions.
Claim 6 is directed to the polyamino acid descriptor comprising at least 100 dimensions.
Colby et al. further teaches networks of this nature take simplified molecular line entry system (SMILES) strings, chemical fingerprints, or molecular graphs as input to predict the same structural representation as output (page 2, column 1).
Colby et al. in view of Kopf et al. do not teach that the polyamino acid descriptor is at least 100 dimensions (claim 6).
Schoenholz et al. describes a deep learning peptide-spectra matching model.
Schoenholz et al. teaches although the use of deep learning models for mapping sequences to structured outputs is well studied, such models are applied in a blackbox manner (page 2, column 1); in contrast, we introduce an architecture that is partially a black-box, and is partially manually designed to reflect the structure of the peptide fragment problem (page 2, column 1); can be easily extended to other input structures, such as graphs or chemical structures using graph embeddings; and as a result can be applied to other mass spectrometry modalities (page 2, column 1).
Schoenholz et al. teaches mapping sequences of amino acids to a “fragment” representation, of dimension F, which contains information about the probabilities of different fragmentations as well as their m/z values (page 4, column 1); and identifying the best DeepMatch model to have a single layer bidirectional LSTM with hidden dimension of 600 (page 7, column 1).
Claim 7 is directed to the latent descriptor comprising at most about 50% of the number of dimensions in the polyamino acid descriptor.
Colby et al. teaches a 128-dimensional dense layer corresponding to the latent vector representation of molecular structure (page 3, column 1).
Schoenholz et al. teaches using a deep bidirectional LSTM with K layers and hidden dimension F/2 to construct the fragment representation, accomplishes the fragment representation goal of modelling the probabilities between different amino acids, the structure of the entire peptide, and accounts for nonlocal information (page 11, column 1).
Therefore, Schoenholz et al. provides sufficient motivation for one of ordinary skill in the art to represent amino acid sequences with a dimension of greater than one hundred (600), and encode the descriptor into a latent layer with 50% of the dimensions. Schoenholz et al. further teaches the techniques to be applicable to models which use deep learning to reconstruct polyamino acid structures, mass spectrometry data (m/z), and solving the peptide fragment representation problem with a reasonable expectation of success.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Colby et al. (Analytical Chemistry; Vol. 92; 2020) in view of Kopf et al. (Patterns 2; 2021), as applied to claims 1-4, 11, 13, 18-19, 22, 25, and 30-31 previously, and in further view of Lopez et al. (Nature Methods; Vol. 15; 2018).
Colby et al. in view of Kopf et al. teach a neural network that inputs a dimensional polyamino acid descriptor, encodes it into a latent layer with fewer dimensions, and uses an optimized loss function, based on prior and posterior probabilities and a zero-inflated negative binomial distribution, to output polyamino acid reconstructions.
Claim 9 is directed to at least approximately 10% of values in the plurality of polyamino acid descriptors in the training data being zero.
Kopf et al. further teaches dropout characteristics of single-cell RNA sequencing data leads to many zero entries in the data matrix (page 4, column 1).
Colby et al. in view of Kopf et al. do not teach approximately 10% of values in the plurality of polyamino acid descriptors in the training data being zero.
Lopez et al. describes deep generative modeling for single-cell transcriptomics.
Lopez et al. teaches evaluating the extent to which the methods fit the data by assessing their ability to accurately impute missing values on five datasets (page 2, column 2); using a zero-inflated negative binomial distribution, which accounts for the observed overdispersion and limited sensitivity (page 1, column 1); and setting set 9% of nonzero entries to zero (page 2, column 2).
Lopez et al. teaches in most cases, methods based on a ZINB distribution—namely, scVI, DCA, and ZINB-WaVE performed better than methods that use alternative distributions, such as log normal in ZIFA, thus supporting the suitability of ZINB for current scRNA-seq datasets (page 3, column 1).
Therefore Lopez et al. provides evidence that one of ordinary skill in the art could apply a zero-inflated negative binomial technique, as taught by Kopf et al. on analogous training data including at least about 10% (9) of data entries with a value of zero, with a reasonable expectation of success and improvement to the recognized problems of sparsity, overdispersion, and limited sensitivity within a deep learning environment. Though the percent is not exact, it falls within the generic range of at least about 10%, and the proportions are so close that prima facie one skilled in the art would expect them to exhibit the same properties (MPEP 2144.05 I).
Claims 12 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Colby et al. (Analytical Chemistry; Vol. 92; 2020) in view of Kopf et al. (Patterns 2; 2021), as applied to claims 1, 11, and 13 previously, and in further view of Kessler et al. (Nesug; 2012.).
Colby et al. in view of Kopf et al. teach a neural network that inputs a dimensional polyamino acid descriptor, encodes it into a latent layer with fewer dimensions, and uses an optimized loss function, based on prior and posterior probabilities to output polyamino acid reconstructions.
Claim 12 is directed to the posterior distribution having a higher kurtosis than a normal distribution. Claim 15 is directed to the prior distribution having a higher kurtosis than a normal distribution.
Colby et al. in view of Kopf et al. do not explicitly teach the posterior or prior distribution having a higher kurtosis than a normal distribution.
Kessler et al. describes a PROC FMM, a finite mixture model procedure.
Kessler et al. teaches Finite mixture models (FMM) provide a flexible framework for analyzing a variety of data (age 1, column 1); and if the corresponding data are multimodal, skewed, heavy-tailed, or exhibit kurtosis, they may not be representative of most known distributions, in which case, you often use a nonparametric method such as kernel density estimation to describe the distribution (page 1, column 1).
Kessler et al. teaches PROC FMM has the ability to produce a SAS data set with important statistics for interpreting mixture models, such as component log likelihoods and prior and posterior probabilities by using the OUTPUT statement (page 3, column 1, in which Output 15 includes the estimated skewness and excess kurtosis of approximately 15, when you would expect the measurements to be close to 0 for a normally distributed variable (page 15, column 1).
Colby et al. further shows a kernel density estimator for each principal component dimension, which emphasizes the density of the distribution (page, fig. 4).
Therefore Kessler et al. provides sufficient motivation for one of ordinary skill in the art to model prior and posterior distributions that demonstrate a higher kurtosis than a normal distribution of zero, with kernel density estimation. Colby et al. teaches using kernel density estimation to account for the density of the data. As such, the higher kurtoses of the data distributions are predictable based on the prior art.
Claims 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over Colby et al. (Analytical Chemistry; Vol. 92; 2020) in view of Kopf et al. (Patterns 2; 2021), as applied to claims 1-4, 11, 13, 18-19, 22, 25, and 30-31 previously, and in further view of Borges et al. (Chemical Reviews; Vol. 121; 2021).
Colby et al. in view of Kopf et al. teach a neural network that inputs a dimensional polyamino acid descriptor and associated assay value, encodes it into a latent layer with fewer dimensions, and uses an optimized loss function, based on prior and posterior probabilities to output polyamino acid reconstructions.
Claim 20 is directed to the given assay method including (i) performing mass spectrometry on cleaved derivatives of the plurality of biomolecules to obtain a plurality of peptide spectral signals and (ii) processing the plurality of peptide spectral signals to obtain a plurality of peptide identifications, wherein the plurality of polyamino acid descriptors comprises the plurality of peptide identifications.
Colby et al. further teaches DarkChem demonstrates novel molecule generation, focused on m/z obtained from mass spectrometry after ionization (page 2, column 1), by analyzing 10 synthetic mixtures and blanks using a drift tube ion mobility spectrometry-mass spectrometer and a 21-T Fourier transform-ion cyclotron resonance spectrometer-mass spectrometer (FTICR-MS) in both positive (+) and negative (−) ionization modes (page 5, column 2); and using the mass-to-charge ratio for compound identification and as the core feature around which most identifications are anchored in current, nontargeted, small molecule identification pipelines (page 2, column 1).
Colby et al. further teaches conducting CCS among multiple ion forms, such as adducts of parent molecules (page 3, column 2).
Colby et al. does not explicitly teach performing mass spectrometry on cleaved derivatives (claims 20-21); for obtaining peptide (claim 20) or protein/protein group identifications (claim 21).
Borges et al. reviews the generation of reference libraries for small-molecule identification; and using quantum chemistry to characterize complex samples.
Regarding claims 20-21, Borges et al. teaches the shotgun proteomics paradigm includes “sequencing” proteins using tandem mass spectrometry; where peptides are digested into their constituent; cleaved to generate peptides of manageable size (page 4, column 1); separated using ionization (ESI) and gas-phase fragmentation, such as collision-induced dissociation (CID); and dissociated to produce fragmentation spectra with constituent m/z (page 4, column 1).
Borges et al. further teaches various software tools use the process for in silico prediction of peptide fragmentation spectra that generate comprehensive reference libraries of predicted peptide spectra for every protein suspected of being present in the sample (page 4, column 1); peptide identification (page 4, column 1); and that deep learning is useful in such in silico property prediction pipelines (page 23, column 1).
Therefore Colby et al. teaches a deep learning prediction pipeline capable of in silico molecule identification via a process of analyzing ion mass spectrometry (m/z) and collision cross section (CCS) spectral data. Borges et al. teaches that identical processes are often used in software tools with cleaved molecule derivatives for the identification of peptides and proteins from a sample. As such, it would be obvious to one of ordinary skill in the art to apply the method of Colby et al. with cleaved derivatives, for the identification of peptides and proteins, as taught by Borges et al., with an expectation of predictable results within a known proteomics paradigm.
Claim 26 is rejected under 35 U.S.C. 103 as being unpatentable over Colby et al. (Analytical Chemistry; Vol. 92; 2020) in view of Kopf et al. (Patterns 2; 2021), as applied to claims 1-4, 11, 13, 18-19, 22, 25, and 30-31 previously, and in further view of Wang et al. (Int. J. Mol. Sci; Vol. 21: 5694; 2020).
Colby et al. in view of Kopf et al. teach a neural network that inputs a dimensional polyamino acid descriptor, encodes it into a latent layer with fewer dimensions, and uses an optimized loss function, based on prior and posterior probabilities to output polyamino acid reconstructions.
Claim 26 is directed to the at least one polyamino acid descriptor comprising an identification of at least one peptide.
Kopf et al. further teaches using LVM to solve challenges arising due to DNA, RNA, or protein information, captured in sequential data, such as nucleotide or amino acid sequences (page8, column 2).
Kopf et al. does not explicitly teach using a peptide identification associated with the polyamino acid descriptor (claim 26).
Wang et al. describes a method of predicting Drug-Target Interactions with Electrotopological State Fingerprints and Amphiphilic Pseudo Amino Acid Composition.
Wang et al. teaches using deep learning-based methods to address biological issues on large scale datasets, such as predicting and discriminating drug target interactions via adjusting a key number of hyperparameters and managing feature dimensions (page 2, column 1).
Wang et al. further teaches representing target proteins by amphiphilic pseudo amino acid composition (APAAC), which effectively reflect the sequence-order information, consider hydrophobicity and hydrophilicity of the constituent amino acids, and is considered effective in drug-target interactions, due to its successful application in protein representation for the prediction of enzymes subfamily, structure and interactions (page 10, column 1).
Wang et al. teaches using web server PROFEAT to calculate commonly used structural and physicochemical features of proteins and peptides from amino acid sequences; applying it to the APAACs (page 10, column 1); and finding that the low-dimensional features based on E-state fingerprints and APAAC obtained satisfactory results and further account for sequence-order information of amino acid sequences when compared to typical amino acid composition (AAC) data (page 11, column 1).
Therefore Wang et al. provides sufficient motivation for one of ordinary skill in the art to associate a polyamino acid descriptor with a peptide identifier, with a reasonable expectation of success and improvement to a deep learning biomolecule prediction framework, such as that of Colby et al. in view of Kopf et al.
Conclusion
No claims are currently allowed.
Correspondence
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Milana Thompson whose telephone number is (571)272-8740. The examiner can normally be reached Monday - Friday, 9:00-6:00 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Karlheinz Skowronek can be reached at (571) 272-1113. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.K.T./Examiner, Art Unit 1687
/Karlheinz R. Skowronek/Supervisory Patent Examiner, Art Unit 1687