Prosecution Insights
Last updated: August 17, 2026
Application No. 18/951,572

Injecting Text in Self-Supervised Speech Pre-training

Non-Final OA §101§DP
Filed
Nov 18, 2024
Priority
Jun 30, 2021 — provisional 63/202,950 +1 more
Examiner
PATEL, SHREYANS A
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
89%
Grant Probability
Favorable
1-2
OA Rounds
3m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 89% — above average
89%
Career Allowance Rate
364 granted / 411 resolved
+28.6% vs TC avg
Moderate +8% lift
Without
With
+8.5%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 0m
Avg Prosecution
34 currently pending
Career history
457
Total Applications
across all art units

Statute-Specific Performance

§101
26.4%
-13.6% vs TC avg
§103
41.1%
+1.1% vs TC avg
§102
21.4%
-18.6% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 411 resolved cases

Office Action

§101 §DP
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 Abstract Idea. Claims 1 and 11 are directed to an abstract idea because, step 1/step 2A, prong one, the claims are about collecting information, analyzing information using mathematical techniques, and training a model based on the analysis. The claims receive two type of data (unspoken text and un-transcribed speech), generates encoded and masked encoded representations, determines a self-supervised loss, and uses that loss to train an audio encoder. These operations describe mathematical calculations and mental processes for organizing and analyzing information rather than an improvement to computer technology itself. The focus of the claims is on processing data and performing machine learning operations, which are abstract ideas. Step 2A, prong two, the claims do not integrate the abstract idea into a practical application. Although the claims are performed on “data processing hardware” and uses an “audio encoder”, these components are described only at a high level and are used as generic computing tools to carry out the mathematical training process. The claims do not recite any specific improvement to computer hardware, audio processing hardware, memory, or any other technological field. Instead, the computer merely performs conventional data processing steps – receiving data, generating representations, computing loss function, and updating a model – which amount to applying the abstract idea on a generic computer. Under Step 2B, the additional claim elements, considered individually and as an ordered combination, do not provide an inventive concept sufficient to transform the abstract idea into patent-eligible subject matter. The claimed operations of encoding data, masking representations, calculating a self-supervised loss, and pre-training a machine learning model are routine machine learning techniques performed using generic computing hardware. The claims do not recite any unconventional algorithm for implementing these operations or any technological improvement beyond using know machine learning methods. Accordingly, the claims merely applies an abstract mathematical training process on generic computer hardware and therefor do not amount to significantly more than the abstract idea itself. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the claims are (i) mere instructions to implement the idea on a computer, and/or (ii) recitation of generic computer structure that serves to perform generic computer functions that are well-understood, routine, and conventional activities previously known to the pertinent industry. Viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. There is further no improvement to the computing device. Dependent claims 2-10, 12-20 further recite an abstract idea performable by a human and do not amount to significantly more than the abstract idea as they do not provide steps other than what is conventionally known. Claims 2 and 12, conventional machine learning techniques that further describe how the abstract model training is performed. Claims 3 and 13, additional data generation and mathematical model training that does not improve computer technology. Claims 4 and 14, mathematical calculations used to train a machine learning model. Claims 5 and 15, merely specifies the type of data used in the abstract training process. Claims 6 and 16, additional mathematical calculations for training the model. Claims 7 and 17, merely identifies known machine learning architectures for carrying out the abstract training process. Claims 8 and 18, another form of collecting and preparing data for the abstract model training process. Claims 9 and 19, merely specifies the source of the training data without improving computer functionality. Claims 10 and 20, conventional machine learning training step that further applies the abstract idea using generic computer technology. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-9 and 11-19 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 13-14, 16-19 and 22-23 of U.S. Patent No. 12,159,619. Although the claims at issue are not identical, they are not patentably distinct from each other because of the following: Pending Application No. 18/951,572 US Patent No. 12,159,619 Claim 1 and 11: receiving training data comprising: unspoken textual utterances, each unspoken textual utterance not paired with any corresponding spoken utterance of non-synthetic speech; and un-transcribed non-synthetic speech utterances, each un-transcribed non- synthetic speech utterance not paired with a corresponding transcription; for each un-transcribed non-synthetic speech utterance: generating a corresponding encoded representation of the un-transcribed non-synthetic speech utterance; generating a corresponding masked encoded representation of the corresponding encoded representation; and determining a self-supervised loss between the corresponding encoded representation and the corresponding masked encoded representation; and pre-training an audio encoder on the self-supervised loss determined for each un- transcribed non-synthetic speech utterance and the unspoken textual utterances to teach the audio encoder to jointly learn shared speech and text representations. Claim 13: receiving training data comprising: unspoken textual utterances, each unspoken textual utterance not paired with any corresponding spoken utterance of non-synthetic speech; and un-transcribed non-synthetic speech utterances, each un-transcribed non-synthetic speech utterance not paired with a corresponding transcription; generating, using a text-to-speech model, a corresponding synthetic speech representation for each unspoken textual utterance of the received training data; and for each synthetic speech representation: generating a corresponding encoded representation of the synthetic speech representation; generating a corresponding masked encoded representation of the corresponding encoded representation; and determining a contrastive loss between the corresponding encoded representation and the corresponding masked encoded representation; and pre-training an audio encoder on the contrastive loss determined for each synthetic speech representation, the synthetic speech representations generated for the unspoken textual utterances, and the un-transcribed non-synthetic speech utterances to teach the audio encoder to jointly learn shared speech and text representations. Claim 2 and 12 corresponds Claim 14 Claim 3 and 13 corresponds Claim 13 Claim 4 and 14 corresponds Claim 16 Claim 5 and 15 corresponds Claim 17 Claim 6 and 16 corresponds Claim 18 Claim 7 and 17 corresponds Claim 19 Claim 8 and 18 corresponds Claim 22 Claim 9 and 19 corresponds Claim 23 Allowable Subject Matter Claims 1-20 would be allowable if the Applicant can overcome the Double Patenting and 101 Abstract Idea set forth. The following is a statement of reasons for the indication of allowable subject matter: Hsu et al. teaches most of the speech-related limitations of claims 1 and 11. It teaches using un-transcribed speech for self-supervised pre-training. In the introduction, Hsu states that “self-supervised learning methods do not rely on any linguistic resources during training,” allowing the model to learn from speech without transcripts (see pg. 3451). Hsu also teaches generating encoded speech representations and masked encoded representations. In Section II(E), Hsu states the model includes “a convolutional waveform encoder, a BERT encoder” and that “the audio encoded features are then randomly masked” before being processed by the BERT encoder (see pg. 3454). Hsu further teaches determining a self-supervised loss by stating that “the final loss is only computed over the masked timesteps,” forcing the model to predict the masked speech from surrounding context (see eq. 1-2, pg. 3454). Baevski et al. (2020) teaches the core speech encoder pre-training limitations of claims 1 and 11. Baevski states that it presents “a framework for self-supervised learning of representations from raw audio data” and that its approach “encodes speech audio via a multi-layer convolutional neural network” (see pg. 1). In section 2, Baevski teaches that “the feature encoder … outputs latent speech representations” and that “we mask a proportion of the feature encoder outputs before feeding them to the context network” (see sections 2 and 3.1, pgs. 2-3). Baevski also teaches computing a self-supervised loss, teaching that “the model is trained via a contrastive task” and that the overall objective is ”L = Lm + aLd” (see section 3.2, eq. 2, pg. 3). Devlin et al. (2019) teaches the text related limitations of claims 1 and 11. In the Abstract, Devlin states that “BERT is designed to pre-train deep bidirectional representation from unlabeled text” (see pg. 4171). Section 3.1 further explains that BERT uses a “masked LM” objective, stating, “we simply mask some percentage of the input tokens at random, and then predict those masked tokens” (see Task #1: Masked LM, pg. 4174). Devlin also identifies the text used for pre-training, stating, “for the pre-training corpus we use the BooksCorpus (800M words) and English Wikipedia (2,500M words)” (see section 3.1, pg. 4175). The difference between the prior art and the claimed invention is that Hsu, Baevski nor Devlin explicitly teach pre-training an audio encoder on the self-supervised loss determined for each un- transcribed non-synthetic speech utterance and the unspoken textual utterances to teach the audio encoder to jointly learn shared speech and text representations. Therefore, it would not have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the teachings of Hsu, Baevski and Devlin to include pre-training an audio encoder on the self-supervised loss determined for each un- transcribed non-synthetic speech utterance and the unspoken textual utterances to teach the audio encoder to jointly learn shared speech and text representations. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Varerkar et al. (US 2018/0300556) – A mechanism is described for facilitating person tracking and data security in machine learning at autonomous machines. A method of embodiments, as described herein, includes detecting, by a camera associated with one or more trackers, a person within a physical vicinity, where detecting includes capturing one or more images the person. The method may further include tracking, by the one or more trackers, the person based on the one or more images of the person, where tracking includes collect tracking data relating to the person. The method may further include selecting a tracker of the one or more trackers as a preferred tracker based on the tracking data. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHREYANS A PATEL whose telephone number is (571)270-0689. The examiner can normally be reached Monday-Friday 8am-5pm PST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. SHREYANS A. PATEL Primary Examiner Art Unit 2653 /SHREYANS A PATEL/Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Nov 18, 2024
Application Filed
Jul 31, 2026
Non-Final Rejection mailed — §101, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12659658
ACOUSTIC ECHO CANCELLATION SYSTEM AND ASSOCIATED METHOD
2y 10m to grant Granted Jun 16, 2026
Patent 12646496
METHODS AND SYSTEMS OF TEXT-CONDITIONED AUDIO-VISUAL SPEECH GENERATION WITH MULTI-MODAL LATENT DIFFUSION MODELS
2y 2m to grant Granted Jun 02, 2026
Patent 12608559
METHOD AND SYSTEM FOR ENHANCING A MUTIMODAL INPUT CONTENT
3y 0m to grant Granted Apr 21, 2026
Patent 12609128
METHOD FOR IMPROVING FAR-FIELD SPEECH INTERACTION PERFORMANCE, AND FAR-FIELD SPEECH INTERACTION SYSTEM
2y 0m to grant Granted Apr 21, 2026
Patent 12586597
ENHANCED AUDIO FILE GENERATOR
3y 6m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
89%
Grant Probability
97%
With Interview (+8.5%)
2y 0m (~3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 411 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month