Prosecution Insights
Last updated: October 02, 2026
Application No. 18/321,044

NATURAL LANGUAGE PROCESSING TO PREDICT PROPERTIES OF PROTEINS

Non-Final OA §101§102§103§112
Filed
May 22, 2023
Priority
May 24, 2022 — provisional 63/345,128
Examiner
ANDERSON-FEARS, KEENAN NEIL
Art Unit
Tech Center
Assignee
Glaxosmithkline Biologicals S.A.
OA Round
1 (Non-Final)
12%
Grant Probability
At Risk
1-2
OA Rounds
1y 0m
Est. Remaining
53%
With Interview

Examiner Intelligence

Grants only 12% of cases
12%
Career Allowance Rate
3 granted / 25 resolved
-48.0% vs TC avg
Strong +41% interview lift
Without
With
+41.3%
Interview Lift
resolved cases with interview
Typical timeline
4y 4m
Avg Prosecution
50 currently pending
Career history
72
Total Applications
across all art units

Statute-Specific Performance

§101
31.2%
-8.8% vs TC avg
§103
40.0%
+0.0% vs TC avg
§102
9.2%
-30.8% vs TC avg
§112
11.8%
-28.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 25 resolved cases

Office Action

§101 §102 §103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Acknowledgment is made of applicant’s claim for priority. Application claims the benefit of U.S. Provisional Application 63/345,128 filed 5/24/2022. As such, the effective filing date of claims 1-10,12-18,21-23 and 26-27 is 5/24/2022. Information Disclosure Statement The information disclosure statement (IDS) submitted on 9/18/2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Status Claims 1-10,12-18,21-23 and 26-27 are pending. Claims 1-10,12-18,21-23 and 26-27 are rejected. Drawings The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: Figure 2A, Item 252; Figure 2B, Item 2200. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance. Nucleotide and/or Amino Acid Sequence Disclosures REQUIREMENTS FOR PATENT APPLICATIONS CONTAINING NUCLEOTIDE AND/OR AMINO ACID SEQUENCE DISCLOSURES Items 1) and 2) provide general guidance related to requirements for sequence disclosures. 37 CFR 1.821(c) requires that patent applications which contain disclosures of nucleotide and/or amino acid sequences that fall within the definitions of 37 CFR 1.821(a) must contain a "Sequence Listing," as a separate part of the disclosure, which presents the nucleotide and/or amino acid sequences and associated information using the symbols and format in accordance with the requirements of 37 CFR 1.821 - 1.825. This "Sequence Listing" part of the disclosure may be submitted: In accordance with 37 CFR 1.821(c)(1) via the USPTO patent electronic filing system (see Section I.1 of the Legal Framework for Patent Electronic System (https://www.uspto.gov/PatentLegalFramework), hereinafter "Legal Framework") as an ASCII text file, together with an incorporation-by-reference of the material in the ASCII text file in a separate paragraph of the specification as required by 37 CFR 1.823(b)(1) identifying: the name of the ASCII text file; ii) the date of creation; and iii) the size of the ASCII text file in bytes; In accordance with 37 CFR 1.821(c)(1) on read-only optical disc(s) as permitted by 37 CFR 1.52(e)(1)(ii), labeled according to 37 CFR 1.52(e)(5), with an incorporation-by-reference of the material in the ASCII text file according to 37 CFR 1.52(e)(8) and 37 CFR 1.823(b)(1) in a separate paragraph of the specification identifying: the name of the ASCII text file; the date of creation; and the size of the ASCII text file in bytes; In accordance with 37 CFR 1.821(c)(2) via the USPTO patent electronic filing system as a PDF file (not recommended); or In accordance with 37 CFR 1.821(c)(3) on physical sheets of paper (not recommended). When a “Sequence Listing” has been submitted as a PDF file as in 1(c) above (37 CFR 1.821(c)(2)) or on physical sheets of paper as in 1(d) above (37 CFR 1.821(c)(3)), 37 CFR 1.821(e)(1) requires a computer readable form (CRF) of the “Sequence Listing” in accordance with the requirements of 37 CFR 1.824. If the "Sequence Listing" required by 37 CFR 1.821(c) is filed via the USPTO patent electronic filing system as a PDF, then 37 CFR 1.821(e)(1)(ii) or 1.821(e)(2)(ii) requires submission of a statement that the "Sequence Listing" content of the PDF copy and the CRF copy (the ASCII text file copy) are identical. If the "Sequence Listing" required by 37 CFR 1.821(c) is filed on paper or read-only optical disc, then 37 CFR 1.821(e)(1)(ii) or 1.821(e)(2)(ii) requires submission of a statement that the "Sequence Listing" content of the paper or read-only optical disc copy and the CRF are identical. Specific deficiencies and the required response to this Office Action are as follows: Specific deficiency - This application fails to comply with the requirements of 37 CFR 1.821 - 1.825 because it does not contain a "Sequence Listing" as a separate part of the disclosure or a CRF of the “Sequence Listing.”. Required response - Applicant must provide: A "Sequence Listing" part of the disclosure; together with An amendment specifically directing its entry into the application in accordance with 37 CFR 1.825(a)(2); A statement that the "Sequence Listing" includes no new matter as required by 37 CFR 1.821(a)(4); and A statement that indicates support for the amendment in the application, as filed, as required by 37 CFR 1.825(a)(3). If the "Sequence Listing" part of the disclosure is submitted according to item 1) a) or b) above, Applicant must also provide: A substitute specification in compliance with 37 CFR 1.52, 1.121(b)(3) and 1.125 inserting the required incorporation-by-reference paragraph, consisting of: A copy of the previously-submitted specification, with deletions shown with strikethrough or brackets and insertions shown with underlining (marked-up version); A copy of the amended specification without markings (clean version); and A statement that the substitute specification contains no new matter. If the "Sequence Listing" part of the disclosure is submitted according to item 1) c) or d) above, applicant must also provide: A CRF in accordance with 37 CFR 1.821(e)(1) or 1.821(e)(2) as required by 1.825(a)(5); and A statement according to item 2) a) or b) above. Specific deficiency – Nucleotide and/or amino acid sequences appearing in the drawings are not identified by sequence identifiers in accordance with 37 CFR 1.821(d). Sequence identifiers for nucleotide and/or amino acid sequences must appear either in the drawings or in the Brief Description of the Drawings. Required response – Applicant must provide: Replacement and annotated drawings in accordance with 37 CFR 1.121(d) inserting the required sequence identifiers; AND/OR A substitute specification in compliance with 37 CFR 1.52, 1.121(b)(3) and 1.125 inserting the required sequence identifiers into the Brief Description of the Drawings, consisting of: A copy of the previously-submitted specification, with deletions shown with strikethrough or brackets and insertions shown with underlining (marked-up version); A copy of the amended specification without markings (clean version); and A statement that the substitute specification contains no new matter. Specific deficiency – Nucleotide and/or amino acid sequences appearing in the specification are not identified by sequence identifiers in accordance with 37 CFR 1.821(d). Required response – Applicant must provide: A substitute specification in compliance with 37 CFR 1.52, 1.121(b)(3) and 1.125 inserting the required sequence identifiers, consisting of: A copy of the previously-submitted specification, with deletions shown with strikethrough or brackets and insertions shown with underlining (marked-up version); A copy of the amended specification without markings (clean version); and A statement that the substitute specification contains no new matter. Specification The disclosure is objected to because it contains an embedded hyperlink and/or other form of browser-executable code. Applicant is required to delete the embedded hyperlink and/or other form of browser-executable code; references to websites should be limited to the top-level domain name without any prefix such as http:// or other browser-executable code. See MPEP § 608.01. Claim Objections Applicant is advised that should claim 26 be found allowable, claim 27 will be objected to under 37 CFR 1.75 as being a substantial duplicate thereof. When two claims in an application are duplicates or else are so close in content that they both cover the same thing, despite a slight difference in wording, it is proper after allowing one claim to object to the other as being a substantial duplicate of the allowed claim. See MPEP § 608.01(m). Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2-3, 6, 14, and 22 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The term “about” in claims 6, 14, and 22 is a relative term which renders the claim indefinite. The term “about” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. The term is used in conjunction with the percentage of masked amino acids within the dataset for training purposes, and as such renders unclear the number of masked amino acids necessary in order to achieve the trained system which is claimed. Claim 2 recites the limitation "the output" in lines 6 and 8. There is insufficient antecedent basis for this limitation in the claim. Claim 3 recites the limitation "the output" in line 3. There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 2-10,12-18,21-23 and 26-27 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract ideas without significantly more. The claims recite a method of training an NLP model, a method for predicting binding affinity, a system for predicting binding affinity and a CRM for the same. This judicial exception is not integrated into a practical application because while claims 2-10,12-18,21-23 and 26-27 attempt to integrate the exception into a practical application, said practical application is a generically recited computer element that does not add meaningful limitations to the abstract idea as it is simply implementing the abstract idea on a computer. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the computer elements only store and retrieve information in memory as well as perform basic calculations that are known to be well-understood, routine and conventional computer functions as recognized by the decisions listed in MPEP § 2106.05(d). Framework with which to Analyze Subject Matter Eligibility: Step 1: Are the claims directed to a category of statutory subject matter (a process, machine manufacture, or composition of matter)? [see MPEP § 2106.03] Claims are directed to statutory subject matter, specifically methods (claims 1-7, 9-10, and 12-17), a system (claims 18, and 21-22), and a CRM (claims 23, and 26-27). Step 2A Prong One: Do the claims recite a judicially recognized exception, i.e., an abstract idea, a law of nature, or a natural phenomenon? [see MPEP § 2106.04(a)] With respect to the Step 2A Prong One evaluation, the instant claims are found herein to recite abstract ideas that fall into the grouping of mental processes and mathematical concepts. The following claims recite abstract ideas (mental processes and mathematical concepts): Claim 2: Computing the cross attention to improve prediction of cross entropy is a process of calculating information that can be done via pen and paper or within the human mind and is therefore an abstract idea, specifically a mental process. Computing the cross attention to improve prediction of cross entropy is a verbal articulation of a mathematical process which is an abstract idea, specifically a mathematical concept. Claim 3: Determining binding probabilities is a process of calculating information that can be done via pen and paper or within the human mind and is therefore an abstract idea, specifically a mental process. Each tuple including the specified information is merely further limiting the data itself which is an abstract idea, specifically a mental process. Claim 4: The TCR dataset comprising the specified information is merely further limiting the data itself which is an abstract idea, specifically a mental process. Performing tokenization of the data is process of calculating information that can be done via pen and paper or within the human mind and is therefore an abstract idea, specifically a mental process. Claims 5 and 21: The TCR dataset and epitope dataset comprising the specified information is merely further limiting the data itself which is an abstract idea, specifically a mental process. Performing tokenization of the data is process of calculating information that can be done via pen and paper or within the human mind and is therefore an abstract idea, specifically a mental process. Claims 6 and 22: Masking a portion of the epitope and TCR datasets is a process of selecting information that can be done via pen and paper or within the human mind and is therefore an abstract idea, specifically a mental process. Claim 7: The transformer model comprising the specified information is merely further limiting the data itself which is an abstract idea, specifically a mental process. Claim 9: Preprocessing the TCR dataset, adding caps for each sequence, categorizing the HLA sequence, and clustering/generating the sequences are processes of calculating, grouping, identifying, and selecting information that can be done via pen and paper or within the human mind and are therefore abstract ideas, specifically mental processes. Claims 10, 18, and 23: Generating a prediction of one or more binding affinities is a process of calculating information that can be done via pen and paper or within the human mind and is therefore an abstract idea, specifically a mental process. The first and second phases comprising the specified information is merely further limiting the data itself which is an abstract idea, specifically a mental process. Claim 12: The property being binding affinity is merely further limiting the data itself which is an abstract idea, specifically a mental process. Claims 13 and 26/27: Using the specified information for training is merely further limiting the data itself which is an abstract idea, specifically a mental process. Claim 14: Using the specified information for training is merely further limiting the data itself which is an abstract idea, specifically a mental process. Claim 15: Generating information that indicates a contribution of the amino acids to the binding affinity is a process of calculating information that can be done via pen and paper or within the human mind and is therefore an abstract idea, specifically a mental process. Claim 17: Analyzing the amino acid sequences, and predicting whether the sequences bind to a TCR epitope are processes of calculating information that can be done via pen and paper or within the human mind and are therefore abstract ideas, specifically mental processes. Step 2A Prong Two: If the claims recite a judicial exception under prong one, then is the judicial exception integrated into a practical application? [see MPEP § 2106.04(d)] Because the claims do recite judicial exceptions, direction under Step 2A Prong Two provides that the claims must be examined further to determine whether they integrate the abstract ideas into a practical application. The following claims recite the following additional elements in the form of non-abstract elements: Claim 2: Training a first and second transformer, and providing the output to a cross-attention module are insignificant extra solution activities, specifically mere data gathering and necessary data outputting (See Performing clinical tests on individuals to obtain input for an equation, In re Grams, 888 F.2d 835, 839- 40; 12 USPQ2d 1824, 1827-28 (Fed. Cir. 1989) and Determining the level of a biomarker in blood, Mayo, 566 U.S. at 79, 101 USPQ2d at 1968. See also PerkinElmer, Inc. v. Intema Ltd., 496 Fed. App'x 65, 73, 105 USPQ2d 1960, 1966 (Fed. Cir. 2012) (assessing or measuring data derived from an ultrasound scan, to be used in a diagnosis)) [See MPEP § 2106.05(g)]. A computer is a generic and nonspecific element of a computer that does not improve the functioning of any computer or technology described herein [See MPEP § 2106.04(d)(1) and MPEP § 2106.05(d)]. Claim 3: Providing the output to inputs of a second neural network is an insignificant extra solution activity, specifically mere data gathering and necessary data outputting (See Performing clinical tests on individuals to obtain input for an equation, In re Grams, 888 F.2d 835, 839- 40; 12 USPQ2d 1824, 1827-28 (Fed. Cir. 1989) and Determining the level of a biomarker in blood, Mayo, 566 U.S. at 79, 101 USPQ2d at 1968. See also PerkinElmer, Inc. v. Intema Ltd., 496 Fed. App'x 65, 73, 105 USPQ2d 1960, 1966 (Fed. Cir. 2012) (assessing or measuring data derived from an ultrasound scan, to be used in a diagnosis)) [See MPEP § 2106.05(g)]. Claims 10, 18, and 23: A computer, computer program product, computer readable storage medium, system, device, processors, instructions, and display screen are generic and nonspecific elements of a computer that do not improve the functioning of any computer or technology described herein [See MPEP § 2106.04(d)(1) and MPEP § 2106.05(d)]. Providing a trained NLP system, training the NLP system, receiving an input query, and displaying the predicted properties are insignificant extra solution activities, specifically mere data gathering and necessary data outputting (See Performing clinical tests on individuals to obtain input for an equation, In re Grams, 888 F.2d 835, 839- 40; 12 USPQ2d 1824, 1827-28 (Fed. Cir. 1989) and Determining the level of a biomarker in blood, Mayo, 566 U.S. at 79, 101 USPQ2d at 1968. See also PerkinElmer, Inc. v. Intema Ltd., 496 Fed. App'x 65, 73, 105 USPQ2d 1960, 1966 (Fed. Cir. 2012) (assessing or measuring data derived from an ultrasound scan, to be used in a diagnosis)) [See MPEP § 2106.05(g)]. Claim 13: Training the NLP system is an insignificant extra solution activity, specifically mere data gathering (See Performing clinical tests on individuals to obtain input for an equation, In re Grams, 888 F.2d 835, 839- 40; 12 USPQ2d 1824, 1827-28 (Fed. Cir. 1989) and Determining the level of a biomarker in blood, Mayo, 566 U.S. at 79, 101 USPQ2d at 1968. See also PerkinElmer, Inc. v. Intema Ltd., 496 Fed. App'x 65, 73, 105 USPQ2d 1960, 1966 (Fed. Cir. 2012) (assessing or measuring data derived from an ultrasound scan, to be used in a diagnosis)) [See MPEP § 2106.05(g)]. Claim 14: Training the NLP system is an insignificant extra solution activity, specifically mere data gathering (See Performing clinical tests on individuals to obtain input for an equation, In re Grams, 888 F.2d 835, 839- 40; 12 USPQ2d 1824, 1827-28 (Fed. Cir. 1989) and Determining the level of a biomarker in blood, Mayo, 566 U.S. at 79, 101 USPQ2d at 1968. See also PerkinElmer, Inc. v. Intema Ltd., 496 Fed. App'x 65, 73, 105 USPQ2d 1960, 1966 (Fed. Cir. 2012) (assessing or measuring data derived from an ultrasound scan, to be used in a diagnosis)) [See MPEP § 2106.05(g)]. Claim 16: Compiling the system into an executable file is a generic and nonspecific elements of a computer that do not improve the functioning of any computer or technology described herein [See MPEP § 2106.04(d)(1) and MPEP § 2106.05(d)]. Claim 21: Training the NLP system is an insignificant extra solution activity, specifically mere data gathering (See Performing clinical tests on individuals to obtain input for an equation, In re Grams, 888 F.2d 835, 839- 40; 12 USPQ2d 1824, 1827-28 (Fed. Cir. 1989) and Determining the level of a biomarker in blood, Mayo, 566 U.S. at 79, 101 USPQ2d at 1968. See also PerkinElmer, Inc. v. Intema Ltd., 496 Fed. App'x 65, 73, 105 USPQ2d 1960, 1966 (Fed. Cir. 2012) (assessing or measuring data derived from an ultrasound scan, to be used in a diagnosis)) [See MPEP § 2106.05(g)]. Claim 22: Training the NLP system is an insignificant extra solution activity, specifically mere data gathering (See Performing clinical tests on individuals to obtain input for an equation, In re Grams, 888 F.2d 835, 839- 40; 12 USPQ2d 1824, 1827-28 (Fed. Cir. 1989) and Determining the level of a biomarker in blood, Mayo, 566 U.S. at 79, 101 USPQ2d at 1968. See also PerkinElmer, Inc. v. Intema Ltd., 496 Fed. App'x 65, 73, 105 USPQ2d 1960, 1966 (Fed. Cir. 2012) (assessing or measuring data derived from an ultrasound scan, to be used in a diagnosis)) [See MPEP § 2106.05(g)]. Step 2B: If the claims do not integrate the judicial exception, do the claims provide an inventive concept? [see MPEP § 2106.05] Because the additional claim elements do not integrate the abstract ideas into a practical application, the claims are further examined under Step 2B, which evaluates whether the additional elements, individually and in combination, amount to significantly more than the judicial exception itself by providing an inventive concept. The claims do not recite additional elements that are sufficient to amount to significantly more than the judicial exceptions because the claims recite additional elements that are generic, conventional, nonspecific, or insignificant extra solution activity. These additional elements include: The additional elements of a computer, computer program product, computer readable storage medium, system, device, processors, instructions, compiling the system into an executable file, and display screen are generic and nonspecific elements of a computer that are well-understood, routine and conventional within the art and therefore do not improve the functioning of any computer or technology described therein (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information), Performing repetitive calculations, Flook, 437 U.S. at 594, 198 USPQ2d at 199 (recomputing or readjusting alarm limit values), and Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015)) [See § MPEP 2106.05(d)(II)]. Therefore, taken both individually and as a whole, the additional elements do not amount to significantly more than the judicial exception by providing an inventive concept. The additional elements of providing a trained NLP system (Conventional:), training the NLP system (Conventional:), receiving an input query, training a first and second transformer (Conventional:), and providing the output to a cross-attention module (Conventional:), and displaying the predicted properties are insignificant extra solution activities, specifically mere data gathering and outputting, that are recognized as well understood, routine and conventional by the courts (See Analyzing DNA to provide sequence information or detect allelic variants, Genetic Techs. Ltd., 818 F.3d at 1377; 118 USPQ2d at 1546, and Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information)) [See MPEP § 2106.05(g)]. Therefore, taken both individually and as whole, the additional elements do not amount to significantly more than the judicial exception by providing an inventive concept. Therefore, claims 2-10,12-18,21-23 and 26-27, when the limitations are considered individually and as a whole, are rejected under 35 U.S.C. § 101 as being directed to non-statutory subject matter. It should be noted that claim is NOT rejected under 35 U.S.C. 101 as it is directed to the training, only, and as such contains no judicial exceptions, only additional elements within the claim. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claims 1-2, 4-7, 10-14, 16-18, 21-23, and 26-27 are rejected under 35 U.S.C. 102(a)(I) as being anticipated by Wu et al. (bioRxiv 2021.11.18.469186 (2021) 1-39). Claim 1 is directed to a method for training a natural language processing system for predicting binding affinity. Claim 10 is directed a method for predicting binding affinity or a level thereof using a natural language processing system. Claim 18 is directed to a system/apparatus for predicting binding affinity or a level thereof using a natural language processing system. Claim 23 is directed to a computer program product for predicting binding affinity or a level thereof using a natural language processing system. Wu et al. teaches in the abstract “due to TCRs’ staggering diversity and the complex binding dynamics underlying TCR antigen recognition, it is challenging to predict which antigens a given TCR may bind to. Here, we present TCR-BERT, a deep learning model that applies self-supervised transfer learning to this problem. TCR-BERT leverages unlabeled TCR sequences to learn a general, versatile representation of TCR sequences, enabling numerous downstream applications. TCR-BERT can be used to build state-of-the-art TCR-antigen binding predictors with improved generalizability compared to prior methods. Simultaneously, TCR-BERT’s embeddings yield clusters of TCRs likely to share antigen specificities. It also enables computational approaches to challenging, unsolved problems such as designing novel TCR sequences with engineered binding affinities”, reading on a computer-implemented method for training a predictive protein language NLP system to predict binding affinity or a level thereof using natural language processing (NLP) comprising: in a first phase, training the predictive protein language NLP system comprising a first neural network on TCR sequence datasets and epitope sequence datasets in a self- supervised manner. Wu et al. teaches on page 2, paragraph 1 “Conventional methods like GLIPH25,26, TCRMatch27, and TCRdist28,29 rely on sequence motif comparisons and manually tuned heuristics to predict which TCRs likely share antigen binding partners. More recently, researchers have applied various machine learning methods to predicting TCR-antigen binding3. While specific methodologies differ, these share the same supervised learning strategy training a classifier on TCR sequences with known antigen binding specificities. However, many TCRs do not have such antigen labels, which greatly limits the amount of data that supervised learning can leverage. Furthermore, it is likely that existing labels are incomplete owing to cross-reactivity. Thus, there is a unique opportunity to develop new approaches that leverage unlabeled data in concert with labelled examples to improve these models’ robustness and generalizability”, and page 2, paragraphs 4-5 “We pre-train TCR-BERT to capture the language of TCR CDR3 sequences by optimizing two objectives sequentially. First, we use unlabeled TCR sequences to learn the grammar of the naturally occurring TCR sequence space. We randomly hide 15% of the residues in each sequence and train TCR-BERT to impute these based on surrounding residues. This masked amino acid (MAA) pre-training does not require knowledge of the antigen specificity of each TCR sequence. MAA pre-training is performed using 88,403 predominantly human TRA and TRB sequences drawn from the VDJdb38 and PIRD39 datasets (Figure 1). After MAA pre-training, we leverage the fact that some TCRs are “labelled” with known antigen binding to train TCR-BERT to predict antigen specificity given TCR sequence in the form of a multi-class prediction problem”, reading on in a second phase, training the predictive protein language NLP system comprising a second neural network with an annotated dataset in a supervised manner to predict binding affinity or a level thereof, wherein the predictive protein language NLP system comprises features from the first phase of training. Wu et al. teaches on page 2, paragraph 4 “TCR-BERT is built upon a modified BERT architecture (see Methods for details) that takes TCR CDR3 amino acid sequences as input”, in the abstract “TCR-BERT can be used to build state-of-the-art TCR-antigen binding predictors”, on page 5, paragraph 1 “The predicted GP33 affinity grows with successive iterations, reaching a minimum 95% predicted probability of binding after 7 iterations”, and Figure 5C “Novel Engineered GP33-binding motifs”, read on wherein the predictive protein language NLP system comprises features from the first phase of training; receiving an input query, from a user interface device coupled to the trained predictive protein language NLP system, comprising a candidate amino acid sequence; generating, by the trained predictive protein language NLP system using one or more processors, a prediction including one or more binding affinities or levels thereof for the candidate TCR sequence and epitope sequence; and displaying, on a display screen of a device, the predicted one or more biophysiochemical properties for the candidate amino acid sequence. Claim 2 is directed to the method of claim 1 but further specifies the providing of the output of the transformers to a cross-attention module that computes a cross attention. Wu et al. teaches on page 4, paragraph 3 “We average each TRA/TRB submodule’s per residue attentions across TRA/TRB pairs of equal length (12 and 14 residues for TRA and TRB, spanning binding and 121 non-binding examples, Figure 4A). TCR-BERT’s attentions tend to be concentrated towards the central region of both the TRA and TRB”, and in Figure 4 “the horizontal axis illustrates each of the 12 attention heads within TCR-BERT. Attentions tend to be concentrated to the center of the TCRs. (B) We relate these averaged attentions”, reading on in a first phase, training a first transformer of a first neural network with a first dataset comprising TCR sequences in a self-supervised manner; in a first phase, training a second transformer of a first neural network with a second dataset comprising epitope sequences in a self-supervised manner; providing the output of the first transformer and second transformer to a cross attention module, wherein the cross attention module computes cross attention using one or more processors between the output of the first transformer model and the output of the second transformer model to improve the prediction of binding affinity. Claim 4 is directed to the method of claim 1 but further specifies that the TCR dataset comprising sequences that have undergone tokenization. Claim 21 is directed to the system/apparatus of claim 18 but further specifies that the TCR dataset comprising sequences that have undergone tokenization. Claim 26 is directed to the computer program product of claim 23 but further specifies that the TCR dataset comprising sequences that have undergone tokenization. Claim 27 is directed to the computer program product of claim 23 but further specifies that the TCR dataset comprising sequences that have undergone tokenization. Wu et al. teaches on page 14, paragraph 3 “TCR-BERT’s input is TCR sequence of length M, formatted as a series of M tokens spanning the set of 20 amino acids”, reading on wherein the TCR sequence dataset comprises a plurality of TCR sequences that have each undergone tokenization at an individual amino acid-level, an n-mer level, or a sub-word level. Claim 5 is directed to the method of claim 1 but further specifies that the TCR dataset comprising sequences that have undergone tokenization and/or the epitope sequence dataset has undergone tokenization. Claim 13 is directed to the method of claim 10 but further specifies that the training is done using TCR sequences that have undergone tokenization and/or epitope sequences that have undergone tokenization. Wu et al. teaches on page 14, paragraph 3 “TCR-BERT’s input is TCR sequence of length M, formatted as a series of M tokens spanning the set of 20 amino acids”, reading on wherein the TCR sequence dataset comprises a plurality of TCR sequences that have each undergone tokenization at an individual amino acid-level, an n-mer level, or a sub-word level, and/or wherein the epitope sequence dataset comprises a plurality of epitope sequences that have each undergone tokenization at an individual amino acid- level, an n-mer level, or a sub-word level. Claim 6 is directed to the method of claim 1 but further specifies the masking of about 10-20% of the sequences. Claim 14 is directed to the method of claim 10 but further specifies the masking of about 10-20% of the sequences. Claim 22 is directed to the system/apparatus of claim 18 but further specifies the masking of about 10-20% of the sequences. Wu et al. teaches on page 2, paragraph 4 “We randomly hide 15% of the residues in each sequence and train TCR-BERT to impute these based on surrounding residues. This masked amino acid (MAA) pre-training does not require knowledge of the antigen specificity of each TCR sequence”, reading on wherein about 10 - 20% of the amino acids in the epitope sequence dataset are masked, and/or about 10-20% of the amino acids in the TCR sequence dataset are masked. Claim 7 is directed to the method of claim 1 but further specifies that the first and second transformers comprise a BERT model. Wu et al. teaches on page 14, paragraphs 2-3 “TCR-BERT is implemented in Python primarily using the PyTorch and Transformers libraries. TCR-BERT uses a lightly modified version of the BERT language modelling architecture…TCR-BERT’s input is TCR sequence of length M, formatted as a series of M tokens spanning the set of 20 amino acids. These input tokens are then padded with special tokens: a classification token C as a prefix, and a separator token S as a suffix. The padded input tokens are then passed through a trained embedding layer that maps each token (amino acid) to a continuous representation of 768 dimensions. As this token embedding does not capture positional information, our TCR-BERT model follows the BERT model in adding a positional encoding to the amino acid embedding. This summed sequence embedding is then fed through a series of 12 transformer blocks to arrive at the overall sequence embedding. This sequence embedding represents each input amino acid chain as a (𝑀 + 2) × 768 matrix. The sequence embedding can then be fed into various “heads” that perform pretraining and downstream tasks such as masked amino acid prediction or sequence classification”, reading on wherein the first transformer model with self-attention further comprises a robustly optimized bidirectional encoder representations from transformers approach model, and the second transformer model with self-attention further comprises a robustly optimized bidirectional encoder representations from transformers approach model. Claim 12 is directed to the method of claim 10 but further specifies the biochemical property being predicted is binding affinity or a level thereof. Wu et al. teaches in the abstract “TCR-BERT leverages unlabeled TCR sequences to learn a general, versatile representation of TCR sequences, enabling numerous downstream applications. TCR-BERT can be used to build state-of-the-art TCR-antigen binding predictors with improved generalizability compared to prior methods. Simultaneously, TCR-BERT’s embeddings yield clusters of TCRs likely to share antigen specificities. It also enables computational approaches to challenging, unsolved problems such as designing novel TCR sequences with engineered binding affinities”, reading on wherein the biophysiochemical property is binding affinity or level thereof of a TCR to an epitope. Claim 16 is directed to the method of claim 10 but further specifies that the trained system is compiled into an executable file. Wu et al. teaches on page 6, paragraph 1 “The TCR-BERT model and model weights are available at github.com/wukevin/tcr-bert”, and on page 14, paragraph 2 “TCR-BERT is implemented in Python primarily using the PyTorch and Transformers libraries”, reading on wherein the trained system is compiled into an executable file. Claim 17 is directed to the method of claim 10 but further specifies the receiving of candidate sequences, analyzing the sequences, and predicting whether the sequences bind to a TCR epitope. Wu et al. teaches on page 2, paragraph 3 “TCR-BERT explicitly leverages unlabeled TCR sequences to achieve state-of-the-art performance on a variety of downstream tasks in TCR analysis, including antigen specificity prediction and exploratory clustering analyses, and even enabling in silico design of TCR sequences with specific binding characteristics”, reading on further comprising: receiving a plurality of candidate amino acid sequences; analyzing the candidate amino acid sequences; and predicting whether the candidate amino acid sequences bind to a TCR epitope. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. (bioRxiv 2021.11.18.469186 (2021) 1-39). Claim 3 is directed to the method of claim 1 but further specifies that the data input be in the form of a tuple, including an epitope and a TCR. Wu et al. teaches the method of claim 1. Wu et al. does not teach the use of a tuple data format. It would have been obvious at the time of first filing to have modified the teachings of Wu et al. for the method of claim 1, with the use of a tuple for information storage as this would be obvious to optimize, specifically because there are only a handful of ways to store data, lists, sets, hash tables, dictionaries, etc., which would be obvious to test to optimize model performance. One would have had a reasonable expectation of success given that there are only so many forms of data input, and tuples are implemented in python, which is the language TCR-Bert is implemented in. therefore, it would have been obvious at the time of first filing to have modified the teachings of each and to be successful. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. (bioRxiv 2021.11.18.469186 (2021) 1-39) as applied to claim 1-2, 4-7, 10-14, 16-18, 21-23, and 26-27 above, and further in view of Moris et al. (Breifings in Bioinformatics (2021) 1-12) and Springer et al. (Frontiers in immunology (2021) 1-11). Claim 9 is directed to the method of claim 1 but further specifies the preprocessing, addition of caps, categorizing the dataset, and clustering. Wu et al. teaches the method of claim 1. Wu et al. teaches on page 16, paragraph 2 “we can use TCR-BERT to obtain a sequence embedding by averaging across each input amino acid’s embedding. For this embedding, we select a representation layer from TCR-BERT that is most conducive to downstream tasks like clustering or building classifiers”, reading on clustering the sequences and generating datasets for training. Wu et al. does not teach the preprocessing, addition of caps, and categorization of the dataset. Moris et al. teaches on page 3, column 1, paragraph 4 “This dataset was reduced to 19,842 unique CDR3-epitope pairs by selecting only human TCR sequences (68,506), removing all spurious CDR3 sequences as defined by VDJdb (66,597), retaining only MHC class I entries (64,386), omitting all entries originating from the 10x Genomics demonstration study (24,513), limiting the length of the CDR3 and epitope sequences to lie between 10-20 and 8-11 amino acids respectively (24,294), and finally removing any duplicate sequence pair. This mixed chain dataset was further split into an alpha chain (5,654 CDR3 alpha (TRA) sequences) and a beta chain (14,188 CDR3 beta (TRB) sequences) dataset. In some experiments the beta chain dataset was down sampled to address the epitope imbalance (see results for exact numbers)”, reading on preprocessing the TCR sequence dataset by selecting for sequences with a specified HLA class and categorizing the HLA sequences and filtering the dataset based on sequence size. Moris et al. does not teach adding caps at the N-terminus and C-terminus of each TCR sequence. Springer et al. teaches in the abstract “We have previously developed ERGO-I (pEptide tcR matchinG predictiOn), a sequence-based T-cell receptor (TCR)-peptide binding predictor that employs natural language processing (NLP) -based methods. We improved it to create ERGO-II by adding the CDR3 alpha segment, the MHC typing, V and J genes, and T cell type (CD4+ or CD8+) as to the predictor”, reading on adding caps at the N-terminus and C-terminus of each TCR sequence. It would have been obvious at the time of first filing to have modified the teachings of Wu et al. for the method of claim 1, with the teachings of Moris et al. for the filtering and selecting for HLA and sequence size, and the teachings of Springer et al. for the addition of the N/C-terminus regions as Moris et al. teaches in the abstract “Our results indicate that while extrapolation to unseen epitopes remains a difficult challenge, ImRex makes this feasible for a subset of epitopes that are not too dissimilar from the training data. We show that appropriate feature engineering methods and rigorous benchmark standards are required to create and validate TCR-epitope predictive models”, and Springer et al. teaches in the abstract “ERGO-II provides for the first time high accuracy prediction of TCR-peptide for previously unseen peptides. For most tested peptides and all measures of binding prediction accuracy, the main contribution was from the beta chain CDR3 sequence, followed by the beta chain V and J and the alpha chain, in that order”. One would have had a reasonable expectation of success given that this is merely further limiting the data based upon specified criteria and including additional sequence information, not altering the NLP method. Therefore, it would have been obvious to a person skilled in the art to have modified the teachings of each and to be successful. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. (bioRxiv 2021.11.18.469186 (2021) 1-39) as applied to claim 1-2, 4-7, 10-14, 16-18, 21-23, and 26-27 above, and further in view of Chrysostomou et al. (Proceedings of the 2021 conference on empirical methods in natural language processing (2021) 8189-8200). Claim 15 is directed to the method of claim 10 but further specifies the use of a salience module that indicates the contribution of amino acids to the prediction of the binding affinity. Wu et al. teaches the method of claim 10. Wu et al. does not teach the use of a salience module that indicates the contribution of amino acids to the prediction of the binding affinity. Chrysostomou et al. teaches in the abstract “In this paper, we hypothesize that salient information extracted a priori from the training data can complement the task-specific information learned by the model during fine-tuning on a downstream task. In this way, we aim to help BERT not to forget assigning importance to informative input tokens when making predictions by proposing SALOSS; an auxiliary loss function for guiding the multi-head attention mechanism during training to be close to salient information extracted a priori using TextRank”, reading on wherein the predictive protein language NLP system comprises a salience module, further comprising: generating, for display on a display screen, information from the salience module that indicates a contribution of respective amino acids to the prediction of the binding affinity of an epitope to a TCR. It would have been obvious at the time of first filing to have modified the teachings of Wu et al. for the method of claim 1, with the teachings of Chrysostomou et al. for the use salience modules for determining the contribution of individual components of the text to the attention and classification, as the latter specifically designs the module for use within the BERT model and states within the abstract “models trained with SALOSS consistently provide more faithful explanations across four different feature attribution methods compared to vanilla BERT. Using the rationales extracted from vanilla BERT and SALOSS models to train inherently faithful classifiers, we further show that the latter result in higher predictive performance in downstream tasks”. One would have had a reasonable expectation of success given that the latter is made for use within the model of the former. Therefore, it would have been obvious at the time of first filing to have modified the teachings of each and to be successful. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to KEENAN NEIL ANDERSON-FEARS whose telephone number is (571)272-0108. The examiner can normally be reached M-Th, alternate F, 8-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Karlheinz Skowronek can be reached at 571-272-9047. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /K.N.A./Examiner, Art Unit 1687 /LARRY D RIGGS II/Supervisory Patent Examiner, Art Unit 1686
Read full office action

Prosecution Timeline

May 22, 2023
Application Filed
Aug 20, 2026
Non-Final Rejection mailed — §101, §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12592298
Hardware Execution and Acceleration of Artificial Intelligence-Based Base Caller
5y 1m to grant Granted Mar 31, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
12%
Grant Probability
53%
With Interview (+41.3%)
4y 4m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 25 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month