Prosecution Insights
Last updated: October 02, 2026
Application No. 18/532,871

System and Method for Secure Speech Feature Extraction

Final Rejection §DP
Filed
Dec 07, 2023
Examiner
OPSASNICK, MICHAEL N
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
4 (Final)
82%
Grant Probability
Favorable
5-6
OA Rounds
4m
Est. Remaining
92%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
754 granted / 922 resolved
+19.8% vs TC avg
Moderate +10% lift
Without
With
+10.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 2m
Avg Prosecution
36 currently pending
Career history
965
Total Applications
across all art units

Statute-Specific Performance

§101
19.2%
-20.8% vs TC avg
§103
33.9%
-6.1% vs TC avg
§102
30.1%
-9.9% vs TC avg
§112
5.2%
-34.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 922 resolved cases

Office Action

§DP
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s claim amendments, filed 06/15/2026, with respect to the independent claims, have been fully considered and are persuasive. The 103 prior art rejection, has been withdrawn. The Obviousness – Type Double Patenting rejection is maintained, and the only rejection remaining. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1,4,5,14,17-19,22-25, 27-30, 32, 33 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-6,11-14,17,18,20 of copending Application No. 18/532900.(reference application), in view of Li et al (20190287515). Claims 1,4,5,14,17-19,21-33 are provisionally rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1,8-20 of copending Application No. 18/532893.(reference application) in view of Li et al (20190287515). Although the claims at issue are not identical, they are not patentably distinct from each other because, as to the ‘900 application, the extra decoding steps are not necessary to realize the functionality of the claims in the instant invention; as to the ‘893 application, the further detailing of the second set of features/embeddings towards acoustic information, is more detailed than the claim scope of the claims in the instant invention. Furthermore, the ‘900 and ‘893 application is silent to the student-teacher neural network relationship, wherein the teacher model is speaker invariant, and the student model is trained such that the network parameter results are similar to the teacher model; however, Li et al (20190287515) a speaker recognition defining at least two artificial networks (abstract), with the relationship defined as a student-teacher model, wherein the teacher model is trained as a full model, with the learned parameters transferred to a smaller “student” model – abstract, wherein the teacher model is speaker invariant – see para 0016, the teacher model is based on clean speech, and speaker invariant as well – para 0026, and is used to train the smaller student model (which handles noisy speech environments – para 0016; and the system uses an adversarial network to train the student model to be more similar/in-line with the teacher model – para 0017. Therefore, it would have been obvious to one of ordinary skill in the art of speaker recognition to further define the neural networks as described in the ‘900/’893, as a student-teacher model relationship, as taught by Li et al (20190287515), because it would advantageously carry the effectiveness of the teacher model, while using the smaller student model, and reducing the divergence between the teacher-student model (Li et al (20190287515) – para 0019, 0020). Examiner further notes that the reduced student model, can be taught/trained for taking into account speech content that is voice invariant. This is a provisional nonstatutory double patenting rejection because the patentably indistinct claims have not in fact been patented. The claims of each patent application, are reproduced below in their entirety, for the ease of evaluating all claim features of each application, in the event claim amendments are contemplated and as not to accidentally trigger further obviousness-type double patenting issues. Examiner notes that the table below, maps the method claims 1,4,5,32,33, and claim 14; the remaining claims are similar in scope and content, and are not repeated below. To see those further mapping relationships, see the presentation under the 35 USC 103 rejection. 18/532871 18/532893 18/532900 1A computer-implemented method, executed on a computing device, comprising: receiving a speech signal comprising acoustic features of a speaker content; altering the acoustic features in the speech signal to obtain an augmented speech signal, the augmented speech signal being a speaker invariant speech signal; and training, over multiple iterations, a student neural network to mimic a trained teacher neural network based on training data that includes the augmented speech signal and the speech signal, the trained teacher neural network being an automated speech recognition (ASR) model trained for speaker invariant speech recognition,….that achieve a threshold level of similarity to the speaker invariant embeddings generated by the trained teacher neural network. 32. The computer-implemented method of claim 1, wherein altering a component of the speaker information comprises adding a perturbation to the speech signal. 33. The computer-implemented method of claim 1, wherein the perturbation includes at least one of voice conversion, pitch shifting, and vocal tract length normalization. 4. The computer-implemented method of claim 1, wherein altering the acoustic features in the speech signal to obtain the augmented speech signal includes adding a loss function constraint to a processing of the received speech signal.. 5. The computer-implemented method of claim 4, wherein the loss function constraint includes at least one of speaker dispersion, speaker identification, and content clustering. 14. A computing system comprising: at least one processor; a memory storing programming instructions for execution by the at least one processor, the programming instructions, upon execution by the at least one processor, causing the system to perform the following operations, receiving a speech signal comprising acoustic features of a speaker; altering the acoustic features in the speech signal to obtain an augmented speech signal, the augmented speech signal being a speaker invariant speech signal; and training, over multiple iterations, a student neural network to mimic a trained teacher neural network based on training data that includes the augmented speech signal and the speech signal, the trained teacher neural network being an automated speech recognition (ASR) model trained for speaker invariant speech recognition. 1. A computer-implemented method, executed on a computing device, comprising: receiving a speech signal; extracting background information from the speech signal to generate a background acoustics embedding; extracting speaker information from the speech signal to generate a speaker acoustics embedding; applying a first loss factor to the background acoustics embedding to decrease speaker information therein to generate a processed background acoustics embedding using machine learning; applying a second loss factor to the speaker acoustics embedding to decrease background information therein to generate a processed speaker acoustics embedding using machine learning; and outputting at least one of the processed background acoustics embedding and the processed speaker acoustics embedding to a speech processing system. 2. The computer-implemented method of claim 1, further including identifying at least one background acoustics metric based on the background acoustics embedding. 3. The computer-implemented method of claim 2, wherein the first loss factor is based on the at least one background acoustics metric. 4. The computer-implemented method of claim 1, further including identifying at least one speaker acoustics metric based on the speaker acoustics embedding. 5. The computer-implemented method of claim 4, wherein the second loss factor is based on the at least one speaker acoustics metric. 6. The computer-implemented method of claim 3, wherein the at least one background acoustics metric comprises measures of at least one of reverberation, noise, and sound quality. 7. The computer-implemented method of claim 5, wherein the at least one speaker acoustics metric comprises measures of at least one of pitch, vocal tract length, gender, accent, language, and age. 8. The computer-implemented method of claim 1, further including combining features of the background information and the speaker information prior to generating the background acoustics embedding and the speaker acoustics embedding. 9. The computer-implemented method of claim 1, further including training a speech processing system with at least one of the processed background acoustics embedding and the processed speaker acoustics embedding. 10. The computer-implemented method of claim 1, further including applying a clustering constraint in the generation of at least one of the background acoustics embedding and the speaker acoustics embedding. 11. A computing system comprising: a memory; and a processor to: receive a speech signal; extract first feature information from the speech signal; generate a first acoustics embedding from the first feature information; generate a second acoustics embedding from the first feature information; identify at least one first acoustics metric based on the first acoustics embedding using machine learning; identify at least one second acoustics metric based on the second acoustics embedding using machine learning; and outputting at least one of the first acoustics metric and the second acoustics metric with a speech processing system. 12. The computing system of claim 11, further including training a speech processing system with at least one of the first acoustics metric and the second acoustics metric. 13. The computing system of claim 11, further including generating a loss factor based on the first acoustics embedding. 14. The computing system of claim 13, further including applying the loss factor to the first acoustics embedding to maximize first information therein and to minimize second information therein. 15. The computing system of claim 14, wherein the loss factor is further based on the second acoustics embedding. 16. The computing system of claim 11 wherein the at least one first acoustics metric comprises background information of the speech signal. 17. The computing system of claim 15, further including applying the loss factor to the second acoustics embedding to maximize second information therein and to minimize first information therein. 18. The computing system of claim 11 wherein the at least one second acoustics metric comprises speaker information of the speech signal. 19. The computing system of claim 18, wherein the background information of the speech signal comprises at least one of reverberation, noise, and sound quality. 20. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising: receiving a speech signal; extracting first feature information from the speech audio signal; extracting second feature information from the speech signal; generating a first acoustics embedding from the first feature information; generating a second acoustics embedding from the first feature information; identifying at least one first acoustics metric based on the first acoustics embedding using machine learning; identifying at least one second acoustics metric based on the second acoustics embedding using machine learning; and training a speech processing system with at least one of the first acoustics metric and the second acoustics metric. 1. A computer-implemented method, executed on a computing device, comprising: receiving, at an encoder, a speech signal comprising a content component and a speaker component, resulting in a received speech signal; processing, using machine learning, the speaker component of the speech signal to generate a representation of speaker information in the speaker component; processing, using machine learning and based at least on the representation of the speaker information, the content component of the audio signal, to generate a representation of content information in the content component having minimized speaker information; transmitting the representation of content information in the content component to a decoder; and decoding the representation of content information in the content component to generate at least a portion of the received speech signal. 2. The computer-implemented method of claim 1, wherein processing the speaker component of the speech signal to generate a representation of speaker information in the speaker component comprises generating a speaker embedding. 3. The computer-implemented method of claim 2, wherein processing the content component of the speech signal to generate a representation of content information in the content component comprises generating a content embedding. 4. The computer-implemented method of claim 3, further comprising generating an estimate of the speaker embedding from the content embedding. 5. The computer-implemented method of claim 4, further comprising comparing the estimate of the speaker embedding to the speaker embedding to generate a loss factor. 6. The computer-implemented method of claim 5, further comprising using the loss factor when generating the content embedding to minimize speaker information within the content embedding. 7. The computer-implemented method of claim 6, further comprising quantizing the content embedding prior to transmitting the content embedding to the decoder. 8. The computer-implemented method of claim 6, further comprising scrambling the content embedding prior to transmitting the content embedding to the decoder. 9. The computer-implemented method of claim 6, further comprising applying a neural watermark including the speaker information to the content embedding. 10. The computer-implemented method of claim 9, wherein the speaker information is encoded as side data of the content embedding. 11. A computing system comprising: a memory; and a processor to: receive, at an encoder, a speech signal comprising a content component and a speaker component of a first voice; process, using machine learning, the speaker component of the speech signal to generate a speaker embedding; process, using machine learning and based at least on the speaker embedding, the content component of the speech signal, to generate a content embedding having minimized speaker information; and transmit the content embedding to a decoder. 12. The computing system of claim 11, further comprising generating an estimate of the speaker embedding from the content embedding. 13. The computing system of claim 12, further comprising comparing the estimate of the speaker embedding to the speaker embedding to generate a loss factor. 14. The computing system of claim 13, further comprising using the loss factor when generating the content embedding to minimize speaker information within the content embedding. 15. The computing system of claim 14, further comprising quantizing the content embedding prior to transmitting the content embedding to the decoder. 16. The computing system of claim 13, further comprising applying a neural watermark including speaker information of the speaker component to the content embedding. 17. The computing system of claim 14, further comprising decoding the speaker information and the content embedding to generate the speech signal. 18. The computing system of claim 16, wherein the speaker information is encoded as side data of the content embedding. 19. The computing system of claim 13, further comprising performing voice conversion on the speech signal wherein the content embedding is transmitted with speaker information of a second voice different from the first voice. 20. A computer program product residing on a non-transitory computer readable medium having a plurality of instructions stored thereon which, when executed by a processor, cause the processor to perform operations comprising: receiving, at an encoder, a speech signal comprising a content component and a speaker component of a first voice, resulting in a received speech signal; processing, using machine learning, the speaker component of the speech signal to generate a speaker embedding; processing, using machine learning and based at least on the speaker embedding, the content component of the voice signal, to generate a content embedding having minimized speaker information, by using a loss factor generated from the speaker embedding when generating the content embedding; and decoding the representation of content information in the content component to generate the received speech signal. Allowable Subject Matter Claims 1,4,5,14,17-19,22-25,27-30,32,33 are allowed over the prior art of record. The following is a statement of reasons for the indication of allowable subject matter: the independent claims reflect that a student-teacher training approach is used to train a student neural network (NN) to perform speaker invariant speech recognition by virtue of training the student NN to mimic a trained teacher NN, which serves to configure the student NNs to recognize speech content invariant to acoustic features of a particular speaker. Li does not suggest training a student NN for speaker invariant speech recognition. In particular, the claimed student-teacher training approach is directed toward recognizing the content of a speech signal without relying on acoustic features of the speaker. In contrast, Li's student-teacher training approach relies on the speaker's voice characteristics to perform speech recognition - which is far simpler than performing speech recognition in a manner that is invariant to the speaker's voice characteristics. See Li, para 26 ("A domain is a definition of the speaker characteristics (e.g., accent, user background) and the characteristics of the acoustic environment (e.g., level of noise, distance to the microphone)." see also Li para 34 ("[Teacher/Student] learning is a form of transfer learning that implicitly handles the speaker and environment variability of the speech signal."); see also Li, para 76 ("Although the T/S framework may perform domain transform, there may still be some speech- recognition problems, such as recognizing speech in the target domain (e.g., the student model). For example, there may [be] data corresponding to different types of noise or different types of speakers."). Due to these fundamental differences, a person of ordinary skill in the art (POSITA) would not have found it obvious to modify Li's student-teacher training technique for speaker invariant speech recognition. Zhang does not cure the above-mentioned deficiencies of Li with respect to amended claim 1 because Zhang does not suggest using a teacher-student training approach to configure a student neural network for speaker-invariant speech recognition. Kim et al (20200357384) teaches a student/teacher model using a discriminator model to minimize any output differential between the two (para 0022, 0048, 0043). Meng et al (20200334538) teaches minimizing loss between the teacher/student models with different domains for each model (para 0032, 0033). Rebryk (20220157316) teaches the combinations of audio encoding and speaker encoding embeddings (first and second) for speech synthesis (see fig. 2), using neural network structures – fig. 3. Wang et al (20230100259) teaches parallel tracks of speaker embeddings and speech feature embedding, and combining for a final similarity calculation and output – see Fig. 4. However, none of the above example prior art references explicitly teach the claim features of the independent claims as discussed above. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The further detailed applicable references are discussed above in the Reasons for Allowable subject Matter. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /Michael N Opsasnick/Primary Examiner, Art Unit 2658 08/24/2026
Read full office action

Prosecution Timeline

Show 7 earlier events
Jan 09, 2026
Examiner Interview Summary
Feb 27, 2026
Request for Continued Examination
Mar 02, 2026
Response after Non-Final Action
Mar 24, 2026
Non-Final Rejection mailed — §DP
Jun 03, 2026
Examiner Interview Summary
Jun 03, 2026
Applicant Interview (Telephonic)
Jun 15, 2026
Response Filed
Aug 27, 2026
Final Rejection mailed — §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12737554
GENERATING TEXT PROMPTS FOR DIGITAL IMAGES UTILIZING VISION-LANGUAGE MODELS AND CONTEXTUAL PROMPT LEARNING
3y 2m to grant Granted Sep 15, 2026
Patent 12711002
GENERATING SYNTHETIC DATA FOR TRAINING LLMS WITH TOOL USE CAPABILITIES
2y 6m to grant Granted Aug 18, 2026
Patent 12707212
SYSTEM FOR FACILITATING IN-PERSON INTERACTION BETWEEN MULTIUSER VIRTUAL ENVIRONMENT USERS WHOSE AVATARS HAVE INTERACTED VIRTUALLY
4y 4m to grant Granted Aug 11, 2026
Patent 12699201
Audio rendering of an electromagnetic metal detection signal
3y 9m to grant Granted Aug 04, 2026
Patent 12676164
System and Method for Modulation Domain-Based Audio Signal Encoding
3y 0m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
82%
Grant Probability
92%
With Interview (+10.3%)
3y 2m (~4m remaining)
Median Time to Grant
High
PTA Risk
Based on 922 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month