Prosecution Insights
Last updated: October 01, 2026
Application No. 18/273,594

PREDICTING COMPLETE PROTEIN REPRESENTATIONS FROM MASKED PROTEIN REPRESENTATIONS

Non-Final OA §101§103
Filed
Jul 21, 2023
Priority
Mar 16, 2021 — provisional 63/161,789 +1 more
Examiner
LUO, JAMMY NMN
Art Unit
Tech Center
Assignee
DeepMind Technologies Limited
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
30 currently pending
Career history
24
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Status Claims 17-28 are cancelled. Claims 1-16 and 29-30 are currently pending and examined on the merits. Priority The instant application is a 371 of PCT/EP2022/051943 filed on 1/27/2022, which claims priority to U.S. Provisional Application 63/161,789 filed on 3/16/2021 At this point in examination, the effective filing date of claims 1-16 and 29-30 is 3/16/2021. Information Disclosure Statement The information disclosure statements (IDS) submitted on 12/28/2023, 8/21/2024, 2/12/2025, 3/19/2025, and 3/17/2026 are in compliance with the provisions of 37 CFR 1.97. A signed copy of the corresponding 1449 form has been included with this Office Action. Drawings The drawings filed on 7/21/2023 are accepted. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-16 and 29-30 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite: (a) mathematical concepts, (e.g., mathematical relationships, formulas or equations, mathematical calculations); and (b) mental processes, i.e., concepts performed in the human mind, (e.g., observation, evaluation, judgement, opinion). Subject matter eligibility evaluation in accordance with MPEP 2106: Eligibility Step 1: Claims 1-16 are directed to a method (process) for unmasking a masked representation of a protein. Claim 29 is directed to a system (machine). Claim 30 is directed to non-transitory computer storage media (machine). Therefore, these claims are encompassed by the categories of statutory subject matter and thus satisfy the subject matter eligibility requirements under Step 1. [Step 1: YES] Eligibility Step 2A: First, it is determined in Prong One whether a claim recites a judicial exception, and if so, then it is determined in Prong Two whether the recited judicial exception is integrated into a practical application of that exception. Eligibility Step 2A, Prong One: In determining whether a claim is directed to a judicial exception, examination is performed that analyzes whether the claim recites a judicial exception, i.e., whether a law of nature, natural phenomenon, or abstract idea is set forth described in the claim. Claims 1-3, 6, 9, 12-16, and 29-30 recite the following steps which fall within the mental processes and/or mathematical concepts groups of abstract ideas, as noted below. Independent claims 1 and 29-30 further recite: processing the masked representation of the protein using the protein reconstruction neural network to generate a respective predicted embedding corresponding to one or more masked embeddings that are included in the masked representation of the protein, wherein a predicted embedding corresponding to a masked embedding in the representation of the amino acid sequence of the protein defines a prediction for an identity of an amino acid at a corresponding position in the amino acid sequence, wherein a predicted embedding corresponding to a masked embedding in the representation of the structure of the protein defines a prediction for a corresponding structural feature of the protein (i.e., mental processes). Dependent claim 2 further recites: updating the masked representation of the protein by replacing a proper subset of the masked embeddings in the masked representation of the protein by corresponding predicted embeddings (i.e., mental processes); processing the updated masked representation of the protein using the protein reconstruction neural network to generate respective predicted embeddings corresponding to one or more remaining masked embeddings that are included in the masked representation of the protein (i.e., mental processes). Dependent claim 3 further recites: processing a predicted amino acid sequence of the protein, defined by replacing each masked embedding in the representation of the amino acid sequence by a corresponding predicted embedding, using a protein folding neural network to generate data defining a predicted protein structure of the predicted amino acid sequence (i.e., mental processes); processing both: (i) the masked representation of the protein, and (ii) the predicted protein structure of the predicted amino acid sequence, using the protein reconstruction neural network to generate a new predicted embedding corresponding to one or more masked embeddings that are included in the masked representation of the protein (i.e., mental processes). Dependent claim 6 further recites: wherein each predicted embedding corresponding to a masked embedding in the representation of the structure of the protein defines a prediction for a spatial distance between a corresponding pair of amino acids in the structure of the protein (i.e., mental processes; this is further information limiting the judicial exceptions). Dependent claim 9 further recites: wherein the protein reconstruction neural network comprises a sequence of update blocks (i.e., mental processes, mathematical concepts); wherein each update block has a respective set of update block parameters and performs operations comprising: receiving current pair embeddings and current single embeddings (i.e., mental processes, mathematical concepts); updating the current single embeddings, in accordance with values of the update block parameters of the update block, based on the current pair embeddings (i.e., mental processes, mathematical concepts); updating the current pair embeddings, in accordance with the values of the update block parameters of the update block, based on the updated single embeddings (i.e., mental processes, mathematical concepts); wherein a final update block in the sequence of update blocks generates final pair embeddings and final single embeddings (i.e., mental processes, mathematical concepts). Dependent claim 12 further recites: wherein updating the current single embeddings based on the current pair embeddings comprises: updating the current single embeddings using attention over the current single embeddings, wherein the attention is conditioned on the current pair embeddings (i.e., mental processes, mathematical concepts). Dependent claim 13 further recites: wherein updating the current single embeddings using attention over the current single embeddings comprises: generating, based on the current single embeddings, a plurality of attention weights (i.e., mental processes, mathematical concepts); generating, based on the current pair embeddings, a respective attention bias corresponding to each of the attention weights (i.e., mental processes, mathematical concepts); generating a plurality of biased attention weights based on the attention weights and the attention biases (i.e., mental processes, mathematical concepts); updating the current single embeddings using attention over the current single embeddings based on the biased attention weights (i.e., mental processes, mathematical concepts). Dependent claim 14 further recites: wherein updating the current pair embeddings based on the updated single embeddings comprises: applying a transformation operation to the updated single embeddings (i.e., mental processes, mathematical concepts); updating the current pair embeddings by adding a result of the transformation operation to the current pair embeddings (i.e., mental processes, mathematical concepts). Dependent claim 15 further recites: wherein the transformation operation comprises an outer product operation (i.e., mental processes, mathematical concepts; this is further information limiting the judicial exceptions). Dependent claim 16 further recites: wherein updating the current pair embeddings based on the updated single embeddings further comprises, after adding the result of the transformation operation to the current pair embeddings: updating the current pair embeddings using attention over the current pair embeddings, wherein the attention is conditioned on the current pair embeddings (i.e., mental processes, mathematical concepts). The abstract ideas recited in the claims are evaluated under the broadest reasonable interpretation (BRI) of the claim limitations when read in light of and consistent with the specification. As the claims are currently recited, the method could be performed by writing down masked representations of protein sequences and their respective updated sequences and making predictions using those sequences. Writing out the data and making decisions on them can be practically performed in the human mind. Additionally, the recited limitations that are identified as judicial exceptions from the mathematical concepts grouping of abstract ideas are abstract ideas irrespective of whether or not the limitations are practical to perform in the human mind. The claims recite sequences of update blocks, which can be represented as a series of mathematical equations used to update single and pair embeddings based on attention weights. Applying transformation operations to embeddings can also be read as inputting embeddings into a mathematical equation to output new embeddings. Therefore, solving these equations to generate updated embeddings is considered a mathematical concept. Therefore, claims 1-3, 6, 9, 12-16, and 29-30 recite an abstract idea. [Step 2A, Prong One: YES] Eligibility Step 2A, Prong Two: In determining whether a claim is directed to a judicial exception, further examination is performed that analyzes if the claim recites additional elements that, when examined as a whole, integrates the judicial exception(s) into a practical application (MPEP 2106.04(d)). A claim that integrates a judicial exception into a practical application will apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception. The claimed additional elements are analyzed to determine if the abstract idea is integrated into a practical application (MPEP 2106.04(d)(I); MPEP 2106.05(a-h)). If the claim contains no additional elements beyond the abstract idea, the claim fails to integrate the abstract idea into a practical application (MPEP 2106.04(d)(III)). The judicial exceptions identified in Eligibility Step 2A, Prong One are not integrated into a practical application because of the reasons noted below. Claim 1 recites processing the masked representation of the protein using the protein reconstruction neural network. Claim 2 recites processing the updated masked representation of the protein using the protein reconstruction neural network. Claim 3 recites processing a predicted amino acid sequence of the protein using a protein folding neural network and processing both: (i) the masked representation of the protein, and (ii) the predicted protein structure of the predicted amino acid sequence, using the protein reconstruction neural network. The limitations recite “using a neural network”, which provides nothing more than mere instructions to implement an abstract idea on a generic computer. The neural network is generically recited because there is no structural information recited about the neural network. Therefore, it is equivalent to a generic computer. See MPEP 2106.05(f). Therefore, the claimed additional elements do not integrate the abstract ideas into a practical application. Claims 1, 3-5, and 7-11 recite the additional non-abstract elements of data gathering: receiving the masked representation of the protein, wherein the masked representation of the protein comprises: (i) a representation of an amino acid sequence of the protein that comprises a plurality of embeddings that each correspond to a respective position in the amino sequence of the protein, and (ii) a representation of a structure of the protein that comprises a plurality of embeddings that each correspond to a respective structural feature of the protein, wherein at least one of the embeddings included in the masked representation of the protein is masked (claim 1); wherein the representation of the amino acid sequence of the protein comprises one or more masked embeddings (claim 3; this is further information limiting the data gathered); wherein each masked embedding included in the masked representation of the protein is a default embedding (claim 4; this is further information limiting the data gathered); wherein the default embedding comprises a vector of zeros (claim 5; this is further information limiting the data gathered); wherein at least one of the embeddings of the representation of the amino acid sequence of the protein is masked (claim 7; this is further information limiting the data gathered); wherein at least one of the embeddings of the representation of the structure of the protein is masked (claim 8; this is further information limiting the data gathered); wherein the representation of the amino acid sequence of the protein comprises a plurality of single embeddings that each correspond to a respective position in the amino acid sequence of the protein (claim 9; this is further information limiting the data gathered); wherein the representation of the structure of the protein comprises a plurality of pair embeddings that each corresponding to a respective pair of positions in the amino acid sequence of the protein (claim 9; this is further information limiting the data gathered); generating the predicted embedding for the masked single embedding based on the corresponding final single embedding generated by the final update block (claim 10); generating the predicted embedding for the masked pair embedding based on the corresponding final pair embedding generated by the final update block (claim 11). Data gathering steps are not an abstract idea, they are extra-solution activity, as they collect the data needed to carry out the JE. The data gathering does not impose any meaningful limitation on the JE, or how the JE is performed. The additional limitation (data gathering) must have more than a nominal or insignificant relationship to the identified judicial exception. (MPEP 2106.04/.05, citing Intellectual Ventures LLC v. Symantee Corp, McRO, TLI communications, OIP Techs. Inc. v. Amason.com Inc., Electric Power Group LLC v. Alstrom S.A.). Claims 1 and 29-30 recite the additional non-abstract element (EIA) of a general-purpose computer system or parts thereof: one or more data processing apparatus (claim 1); a system comprising computers and storage devices (claim 29); one or more non-transitory computer storage media (claim 30). The EIA do not provide any details of how specific structures of the computer elements are used to implement the JE. The claims require nothing more than a general-purpose computer to perform the functions that constitute the judicial exceptions. The computer elements of the claims do not provide improvements to the functioning of the computer itself (as in DDR Holdings, LLC v. Hotels.com LP); they do not provide improvements to any other technology or technical field (as in Diamond v. Diehr); nor do they utilize a particular machine (as in Eibel Process Co. v. Minn. & Ont. Paper Co.). Hence, these are mere instructions to apply the JE using a computer, and therefore the claim does not recite integrate that JE into a practical application. Thus, the additionally recited elements merely invoke a computer as a tool, and/or amount to insignificant extra-solution data gathering activity, and as such, when all limitations in claims 1-16 and 29-30 have been considered as a whole, the claims are deemed to not recite any additional elements that would integrate a judicial exception into a practical application. Claims 1-5, 7-11, and 29-30 contain additional elements that would not integrate a judicial exception into a practical application and are further probed for inventive concept in Step 2B. [Step 2A, Prong Two: NO] Eligibility Step 2B: Because the claims recite an abstract idea, and do not integrate that abstract idea into a practical application, the claims are probed for a specific inventive concept. The judicial exception alone cannot provide that inventive concept or practical application (MPEP 2106.05). Identifying whether the additional elements beyond the abstract idea amount to such an inventive concept requires considering the additional elements individually and in combination to determine if they amount to significantly more than the judicial exception (MPEP 2106.05A i-vi). The claims do not include any additional elements that are sufficient to amount to significantly more than the judicial exception(s) because of the reasons noted below. With respect to claims 1, 3-5, and 7-11: The courts have recognized that computer functions such as receiving or transmitting data over a network and storing and retrieving information in memory are well-understood, routine, and conventional. This is evidenced by Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014) (computer receives and sends information over a network); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). See MPEP 2106.05(d)(II). With respect to claims 1 and 29-30: The limitations identified above as non-abstract elements (EIA) related to general-purpose computer systems do not rise to the level of significantly more than the judicial exception. These elements do not improve the functioning of the computer itself, or comprise an improvement to any other technical field (Trading Technologies Int’l v. IBG, TLI Communications). They do not require or set forth a particular machine (Ultramercial v. Hulu, LLC., Alice Corp. Pty. Ltd v. CLS Bank Int’l), they do not affect a transformation of matter, nor do they provide an unconventional step. Simply appending well understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception are insufficient to provide significantly more (as discussed in Alice Corp., CyberSource v. Retail Decisions, Parker v. Flook, Versata Development Group v. SAP America). The additional elements of processing the masked representation of the protein using the protein reconstruction neural network (claim 1), processing the updated masked representation of the protein using the protein reconstruction neural network (claim 2), and processing a predicted amino acid sequence of the protein using a protein folding neural network and processing both: (i) the masked representation of the protein, and (ii) the predicted protein structure of the predicted amino acid sequence, using the protein reconstruction neural network (claim 3) is conventional. Evidence for conventionality is shown by Kuhlman et al. (Nature Reviews Molecular Cell Biology, 2019, 20, 681-697). Kuhlman et al. reviews using deep convolutional neural networks in protein structural analysis by extracting features from protein sequences to output protein structure predictions (pg. 687, Box 2). This shows processing protein representations and structural information using neural networks, which makes it a conventional practice in the art. [Step 2B: NO] Therefore, claims 1-16 and 29-30 are patent ineligible under 35 U.S.C. § 101. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-2, 6-8, and 29-30 are rejected under 35 U.S.C. 103 as being unpatentable over Strokach et al. (Cell Systems, 2020, 11(4), 402-411.e4), in view of Anand et al. (NIPS'18: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 2018, 1-26). With respect to claims 1 and 29-30: Claim 1 recites one or more data processing apparatus. Claim 29 recites a system comprising one or more computers and one or more storage devices coupled to the computers. Claim 30 recites one or more non-transitory computer storage media. Strokach et al. discloses source code for ProteinSolver as well as a web server implementation which allows the user to run a trained ProteinSolver model to generate new protein sequences matching geometric constraints extracted from a reference structure and is deployed to virtual machines equipped with an NVIDIA Tesla T4 GPU, with at most eight concurrent users sharing a single GPU (pg. 411.e2, Section “Network Implementation”; pg. 411.e3, Section “Web Server Implementation”). The source code and web server requires an operating system of a general-purpose computer such as Windows, Mac, or LINUX. These general-purpose computers require the recited data processing apparatus, system comprising computers and storage devices, and non-transitory computer storage media. Regarding the recited receiving the masked representation of the protein, wherein the masked representation of the protein comprises: (i) a representation of an amino acid sequence of the protein that comprises a plurality of embeddings that each correspond to a respective position in the amino sequence of the protein, and (ii) a representation of a structure of the protein that comprises a plurality of embeddings that each correspond to a respective structural feature of the protein, wherein at least one of the embeddings included in the masked representation of the protein is masked, Strokach et al. discloses training a deep graph neural network ProteinSolver by providing as input a partially masked amino acid sequence with adjacency matrices of distances between pairs of amino acids adapted from structural templates of proteins and minimizing the cross-entropy loss between network predictions and the identities of the masked amino acid residues (pg. 402, Summary, lines 2-5; pg. 403, col. 2, para. 2, lines 2-6; pg. 404, Figure 1; pg. 406, col. 1, para. 1; pg. 406, col. 2, para. 1, lines 6-11). This teaches a masked representation of proteins comprising masked amino acid sequences and an adjacency matrix representation that encodes Euclidean distances between amino acids and the relative positions of those amino acids along the amino acid chain, as depicted in Figure 1. Regarding the recited processing the masked representation of the protein using the protein reconstruction neural network to generate a respective predicted embedding corresponding to one or more masked embeddings that are included in the masked representation of the protein, wherein a predicted embedding corresponding to a masked embedding in the representation of the amino acid sequence of the protein defines a prediction for an identity of an amino acid at a corresponding position in the amino acid sequence, Strokach et al. discloses training a deep graph neural network ProteinSolver by providing as input a partially masked amino acid sequence and outputting a proposed sequence with predicted amino acids corresponding to the masked amino acids at their respective positions (pg. 404, Figure 1; pg. 406, col. 2, para. 1, lines 6-11; pg. 406, Figure 2(A), lines 1-2). This teaches processing masked amino acid sequences using a neural network to generate a predicted amino acid sequence, recovering the identities of masked amino acid sequences. Strokach et al. does not disclose wherein a predicted embedding corresponding to a masked embedding in the representation of the structure of the protein defines a prediction for a corresponding structural feature of the protein. However, Anand et al. discloses using trained Generative Adversarial Networks (GANs) to infer contextually correct missing portions of protein structures by formulating this problem as an inpainting problem, where for a subset of residues, all pairwise distances are eliminated and the task is to fill in these distances reasonably given the context of the rest of the uncorrupted structure (pg. 1, Abstract, lines 5-9; pg. 7, Section “4.2 Inpainting for protein design”, pg. 8, Section “4.2.2 Inpainting results”, para. 1; pg. 22, Figure S9). Figure S9 depicts masked input data corresponding to deletion of all pairwise distances for 20 consecutive residues and fake samples generated by the model to fill in the masked regions. Therefore, this teaches predicting a structural feature of distances for a corresponding masked portion of a protein. With respect to claim 2: Anand et al. does not disclose updating the masked representation of the protein by replacing a proper subset of the masked embeddings in the masked representation of the protein by corresponding predicted embeddings; processing the updated masked representation of the protein using the protein reconstruction neural network to generate respective predicted embeddings corresponding to one or more remaining masked embeddings that are included in the masked representation of the protein. However, Strokach et al. discloses incremental generation of the most probable protein sequence using ProteinSolver, which passes inputs through the network once for every missing label (pg. 406, col. 2, para. 1, lines 6-11; pg. 404, Figure 1; pg. 411.e2, Section “Generating Novel Protein Sequences”, para. 1, lines 3-5). At each iteration, the label the network makes the most confident prediction is accepted and that label is treated as given in all subsequent iterations. This teaches iteratively updating masked amino acid sequences by making predictions on the masked amino acids and processing updated sequences in subsequent iterations through the neural network to generate further predictions. With respect to claim 6: Strokach et al. does not disclose wherein each predicted embedding corresponding to a masked embedding in the representation of the structure of the protein defines a prediction for a spatial distance between a corresponding pair of amino acids in the structure of the protein. However, Anand et al. discloses using trained Generative Adversarial Networks (GANs) to infer contextually correct missing portions of protein structures by formulating this problem as an inpainting problem, where for a subset of residues, all pairwise distances are eliminated and the task is to fill in these distances reasonably given the context of the rest of the uncorrupted structure (pg. 1, Abstract, lines 5-9; pg. 7, Section “4.2 Inpainting for protein design”, pg. 8, Section “4.2.2 Inpainting results”, para. 1; pg. 22, Figure S9). Figure S9 depicts masked input data corresponding to deletion of all pairwise distances for 20 consecutive residues and fake samples generated by the model to fill in the masked regions. New inpainted maps were folded into structures to test whether the inpainted portions correspond to legitimate reconstructions of missing parts of the protein. Also, further discloses 3D structures of proteins encoded as 2D pairwise distances between α -carbons on the protein backbone (pg. 4, para. 1, lines 2-3). This teaches predicting spatial distances corresponding to masked distances between α -carbons of amino acids in the structure of proteins. With respect to claim 7: Anand et al. does not disclose wherein at least one of the embeddings of the representation of the amino acid sequence of the protein is masked. However, Strokach et al. discloses training a deep graph neural network ProteinSolver by providing as input a partially masked amino acid sequence with adjacency matrices of distances between pairs of amino acids adapted from structural templates of proteins and minimizing the cross-entropy loss between network predictions and the identities of the masked amino acid residues (pg. 402, Summary, lines 2-5; pg. 404, Figure 1; pg. 406, col. 2, para. 1, lines 6-11). This teaches embeddings of amino acid sequences of proteins are masked, as depicted in Figure 1. With respect to claim 8: Strokach et al. does not disclose wherein at least one of the embeddings of the representation of the structure of the protein is masked. However, Anand et al. discloses using trained Generative Adversarial Networks (GANs) to infer contextually correct missing portions of protein structures by formulating this problem as an inpainting problem, where for a subset of residues, all pairwise distances are eliminated and the task is to fill in these distances reasonably given the context of the rest of the uncorrupted structure (pg. 1, Abstract, lines 5-9; pg. 7, Section “4.2 Inpainting for protein design”, pg. 8, Section “4.2.2 Inpainting results”, para. 1; pg. 22, Figure S9). Figure S9 depicts masked input data corresponding to deletion of all pairwise distances for 20 consecutive residues and fake samples generated by the model to fill in the masked regions. This teaches masking distances of proteins. It would have been prima facie obvious to one of ordinary skill in the art to combine the protein sequence prediction method disclosed by Strokach et al. with masked structural information disclosed by Anand et al. One would be motivated to combine the sequence prediction method with masked structural information because Anand et al. discloses that their methods are capable of quickly producing structurally plausible solutions in predicting completions of corrupted protein structures (pg. 1, Abstract, lines 12-14). Therefore, masked structural information of corrupted protein structures can be quickly predicted in the prediction method. There is a likelihood of success, since both methods are of protein sequence analysis and protein structure prediction, which are well known techniques in the field of bioinformatics. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Strokach et al. (Cell Systems, 2020, 11(4), 402-411.e4) and Anand et al. (NIPS'18: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 2018, 1-26) as applied to claims 1-2, 6-8, and 29-30 above, in view of Thomas et al. (Berkeley Artificial Intelligence Research, 2019, 1-13) and Qin et al. (Extreme Mechanics Letters, 2020, 36, 1-11). Strokach et al. and Anand et al. are applied to claims 1-2, 6-8, and 29-30 above. With respect to claim 3: Anand et al. does not disclose processing both: (i) the masked representation of the protein, and (ii) the predicted protein structure of the predicted amino acid sequence, using the protein reconstruction neural network to generate a new predicted embedding corresponding to one or more masked embeddings that are included in the masked representation of the protein. However, Strokach et al. discloses training a deep graph neural network ProteinSolver by providing as input a partially masked amino acid sequence with adjacency matrices of distances between pairs of amino acids adapted from structural templates of proteins and minimizing the cross-entropy loss between network predictions and the identities of the masked amino acid residues (pg. 402, Summary, lines 2-5; pg. 403, col. 2, para. 2, lines 2-6; pg. 404, Figure 1; pg. 406, col. 1, para. 1; pg. 406, col. 2, para. 1, lines 6-11). This teaches processing a masked representation of proteins comprising masked amino acid sequences and protein structure information using a neural network to generate predicted amino acid sequences. Strokach et al. and Anand et al. do not disclose wherein the representation of the amino acid sequence of the protein comprises one or more masked embeddings, and further comprising: processing a predicted amino acid sequence of the protein, defined by replacing each masked embedding in the representation of the amino acid sequence by a corresponding predicted embedding, using a protein folding neural network to generate data defining a predicted protein structure of the predicted amino acid sequence. However, Thomas et al. discloses using Bidirectional Encoder Representations from Transformers (BERT) language model to predict masked portions of amino acid sequences and using a structure prediction model to output a predicted protein structure from the predicted amino acid sequences (pg. 6-8, Section “Learning the Language of Proteins via Self-Supervision”). This teaches predicting an amino acid sequence from a masked amino acid sequence using the BERT model, which is used as input to the structure prediction model to generate a predicted protein structure of the predicted amino acid sequence. Thomas et al. does not disclose using a protein folding neural network to generate data defining a predicted protein structure. However, Qin et al. discloses training a Multi-scaled Neighborhood-based Neural Network (MNNN) model with amino acid sequences and labeled dihedral angles to predict phi-psi angles of any sequence and build atomic structures (pg. 1, Abstract, lines 5-7; pg. 3, Figure 1; pg. 4, Figure 2). This teaches a protein folding neural network that generates data defining a protein structure. It would have been prima facie obvious to one of ordinary skill in the art to substitute the protein structure information disclosed by Strokach et al. and Anand et al. for the predicted protein structures disclosed by Thomas et al. and the protein folding neural network disclosed by Qin et al. One would be motivated to substitute protein structure information for predicted protein structures and protein folding neural network because Thomas et al. discloses that the BERT model can gain implicit understanding of how proteins are constructed from learning partial information from other families of proteins to provide informative features for protein structure prediction (pg. 8, para. 3). Therefore, predicted protein structures utilizing this BERT model will be most accurate and informative for protein structure prediction in the unmasking method. Qin et al. discloses that the MNNN multi-stage modeling process can effectively give the fast prediction of the 3D protein structure for any known or unknown protein sequences, which provides a helpful tool to sequence design (pg. 3, col. 1, para. 1, lines 4-7). This protein folding neural network will improve the speed at which the unmasking method makes predictions. There is a likelihood of success, since all methods are of protein sequence analysis and protein structure prediction, which are well known techniques in the field of bioinformatics. Claims 4-5 are rejected under 35 U.S.C. 103 as being unpatentable over Strokach et al. (Cell Systems, 2020, 11(4), 402-411.e4) and Anand et al. (NIPS'18: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 2018, 1-26) as applied to claims 1-2, 6-8, and 29-30 above, in view of Wagner et al. [US11544537B2]. Strokach et al. and Anand et al. are applied to claims 1-2, 6-8, and 29-30 above. With respect to claim 4: Strokach et al. and Anand et al. do not disclose wherein each masked embedding included in the masked representation of the protein is a default embedding. However, Wagner et al. discloses mapping masked tokens in sequences of words to vectors of zeroes (pg. 9, col. 2, lines 32-34 and 55-59; pg. 10, col. 4, lines 60-65). A neural network is trained to guess what the hidden masked token is. The instant specification defines an embedding as being “masked” if the embedding is a predefined embedding or an embedding represented as a vector of zeroes (para. [0051], lines 7-9). Therefore, this teaches masked tokens represented as vectors of zeroes as default embeddings. With respect to claim 5: Strokach et al. and Anand et al. do not disclose wherein the default embedding comprises a vector of zeros. However, Wagner et al. discloses mapping masked tokens in sequences of words to vectors of zeroes (pg. 9, col. 2, lines 32-34 and 55-59; pg. 10, col. 4, lines 60-65). A neural network is trained to guess what the hidden masked token is. The instant specification defines an embedding as being “masked” if the embedding is a predefined embedding or an embedding represented as a vector of zeroes (para. [0051], lines 7-9). Therefore, this teaches default embeddings comprising vectors of zeroes. It would have been prima facie obvious to one of ordinary skill in the art to modify the unmasking method disclosed by Strokach et al. and Anand et al. to incorporate default embeddings disclosed by Wagner et al. One would be motivated to incorporate default embeddings into the unmasking method because Wagner et al. discloses that sequence-based neural networks can be applied to biological sequence analysis (pg. 10, col. 4, lines 1-15). Therefore, masked embeddings represented as vectors of zeroes from sequence-based neural networks can be used in neural networks for masked protein sequence analysis. There is a likelihood of success, since neural networks used in protein and general word sequence analyses are relevant and well known techniques in the field of bioinformatics. Claims 9 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Strokach et al. (Cell Systems, 2020, 11(4), 402-411.e4) and Anand et al. (NIPS'18: Proceedings of the 32nd International Conference on Neural Information Processing Systems, 2018, 1-26) as applied to claims 1-2, 6-8, and 29-30 above, in view of Jumper et al. (Fourteenth Critical Assessment of Techniques for Protein Structure Prediction 13, 2020, 1-42). Strokach et al. and Anand et al. are applied to claims 1-2, 6-8, and 29-30 above. With respect to claim 9: Strokach et al. and Anand et al. do not disclose wherein the representation of the amino acid sequence of the protein comprises a plurality of single embeddings that each correspond to a respective position in the amino acid sequence of the protein. However, Jumper et al. discloses a multiple sequence alignment (MSA) figure comprising single embeddings corresponding to respective positions in the amino acid sequence of a protein from an animal (pg. 10). This teaches protein sequence representations comprising single embeddings. Strokach et al. and Anand et al. do not disclose wherein the representation of the structure of the protein comprises a plurality of pair embeddings that each corresponding to a respective pair of positions in the amino acid sequence of the protein. However, Jumper et al. discloses structural templates of proteins and paired embeddings corresponding to respective positions in the amino acid sequence of a protein from an animal (pg. 10). This teaches structural information of proteins comprising pair embeddings. Strokach et al. and Anand et al. do not disclose wherein the protein reconstruction neural network comprises a sequence of update blocks. However, Jumper et al. discloses layers of a neural network comprising a sequence of update blocks (pg. 10). This teaches a neural network with update blocks. Strokach et al. and Anand et al. do not disclose wherein each update block has a respective set of update block parameters and performs operations comprising: receiving current pair embeddings and current single embeddings. However, Jumper et al. discloses update block parameters such as attention, and the update blocks receive pair embeddings and single embeddings (pg. 10). This teaches update block parameters and update block operations such as receiving pair embeddings and single embeddings. Strokach et al. and Anand et al. do not disclose updating the current single embeddings, in accordance with values of the update block parameters of the update block, based on the current pair embeddings. However, Jumper et al. discloses updating single embeddings with respect to the pair embeddings using attention (pg. 10). This teaches updating single embeddings according to attention, based on pair embeddings. Strokach et al. and Anand et al. do not disclose updating the current pair embeddings, in accordance with the values of the update block parameters of the update block, based on the updated single embeddings. However, Jumper et al. discloses updating pair embeddings with respect to the single embeddings using attention (pg. 10). This teaches updating pair embeddings according to attention, based on single embeddings. Strokach et al. and Anand et al. do not disclose wherein a final update block in the sequence of update blocks generates final pair embeddings and final single embeddings. However, Jumper et al. discloses final update blocks in the sequence of update blocks generating final pair embeddings and final single embeddings used in predicting a 3D structure of the protein (pg. 10). This teaches final update blocks generating final pair embeddings and final single embeddings used in the 3D structure of the protein. With respect to claim 12: Strokach et al. and Anand et al. do not disclose wherein updating the current single embeddings based on the current pair embeddings comprises: updating the current single embeddings using attention over the current single embeddings, wherein the attention is conditioned on the current pair embeddings. However, Jumper et al. discloses updating single embeddings in the MSA using attention, where the attention is conditioned on the pair embeddings (pg. 10). This teaches updating current single embeddings using attention that is conditioned on the current pair embeddings. It would have been prima facie obvious to one of ordinary skill in the art to modify the unmasking method disclosed by Strokach et al. and Anand et al. to incorporate updating single and pair embeddings using attention disclosed by Jumper et al. One would be motivated to incorporate updating single and pair embeddings using attention in the unmasking method because Jumper et al. discloses AlphaFold 2 as a deep learning tool comprising a full end-to-end pipeline for protein structure prediction (pg. 40). Therefore, updating single and pair embeddings will improve efficiency in the unmasking method pipeline. There is a likelihood of success, since all methods are of protein sequence analysis and protein structure prediction, which are well known techniques in the field of bioinformatics. Claims Free from Prior Art Claim 10 recites generating the predicted embedding for the masked single embedding based on the corresponding final single embedding generated by the final update block. Claim 11 recites generating the predicted embedding for the masked pair embedding based on the corresponding final pair embedding generated by the final update block. Claim 13 recites generating, based on the current single embeddings, a plurality of attention weights; generating, based on the current pair embeddings, a respective attention bias corresponding to each of the attention weights; generating a plurality of biased attention weights based on the attention weights and the attention biases; and updating the current single embeddings using attention over the current single embeddings based on the biased attention weights. Claim 14 recites applying a transformation operation to the updated single embeddings; and updating the current pair embeddings by adding a result of the transformation operation to the current pair embeddings. Claim 15 recites wherein the transformation operation comprises an outer product operation. Claim 16 recites updating the current pair embeddings using attention over the current pair embeddings, wherein the attention is conditioned on the current pair embeddings. These limitations are free of the art. Conclusion No claims are allowed. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jammy Luo whose telephone number is (571)272-2358. The examiner can normally be reached Monday - Friday, 9:00 AM - 5:00 PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Larry D Riggs can be reached at (571)270-3062. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.N.L./Examiner, Art Unit 1686 /OLIVIA M. WISE/Supervisory Patent Examiner, Art Unit 1685
Read full office action

Prosecution Timeline

Jul 21, 2023
Application Filed
Sep 09, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month