Prosecution Insights
Last updated: October 02, 2026
Application No. 18/974,383

Delta Models for Providing Privatized Speech-to-Text During Virtual Meetings

Non-Final OA §103§DP
Filed
Dec 09, 2024
Priority
Apr 29, 2022 — continuation of 12/165,646
Examiner
GAUTHIER, GERALD
Art Unit
Tech Center
Assignee
Zoom Video Communications Inc.
OA Round
1 (Non-Final)
91%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 91% — above average
91%
Career Allowance Rate
1670 granted / 1834 resolved
+31.1% vs TC avg
Moderate +7% lift
Without
With
+6.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
24 currently pending
Career history
1845
Total Applications
across all art units

Statute-Specific Performance

§101
10.4%
-29.6% vs TC avg
§103
31.0%
-9.0% vs TC avg
§102
27.5%
-12.5% vs TC avg
§112
7.0%
-33.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1834 resolved cases

Office Action

§103 §DP
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on December 09, 2024, is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application is currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 1-3, 6, 8-10, 13 and 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Robert Jose et al. (US 2022/0301561 A1) in view of Gorzela et al. (US 20180137472 A1). As to claim 1, Robert Jose discloses a method [§ 0001] comprising: receiving, from a remote server, a local model for speech recognition, wherein the local model comprises a copy of a centralized model [“Enabling, on a local device, a voice processing system that limits the amount of data needed to be transmitted to a remote server in translating a voice input into executable data.” § 0003]. performing, by the first client device, speech recognition using the local model on at least one audio stream of the one or more audio streams, wherein performing speech recognition comprises [“The local speech-to-text model is a neural network model or machine learning model supplied to the local device by a remote server which is pre-trained to recognize a limited set of words corresponding to actions that the local device can perform.” § 0023]: identifying, based on a vocabulary database for a user [A transcription of the query], user-specific vocabulary within the at least one audio stream [“Server processing circuitry is in generating a transcription of the query, identify a plurality of phonemes corresponding to the individual sounds of the query.” § 0026-0027]; and generating, based on the user-specific vocabulary, a private transcription of the one or more audio streams, wherein the private transcription comprises the user-specific vocabulary [“Local device processing circuitry processes the voice query received by the local device via input circuitry using a local speech-to-text model to generate a transcription of the voice query.” § 0028]. Robert Jose fails to disclose joining a virtual meeting, by a first client device. However, Gorzela teaches joining, by a first client device, a virtual meeting having a plurality of participants, the virtual meeting involving exchanging one or more audio streams between the participants [§ 0023]. Robert Jose and Gorzela are analogous because they are all directed at participant speech recognition management system. One of ordinary skill in the art before the effective filing date of the claimed invention would have found obvious to modify Robert Jose reference with the teaching of Gorzela, so that the local device would include the joining of the meeting option in the system of Robert Jose, would have been combined into a joining of a virtual meeting, for the obvious purpose of providing the user the time for the meeting communications, by combining prior art elements according to known methods to yield predictable results. As to claim 2, Robert Jose discloses the method of claim 1, further comprising generating a neutralized transcription based on the private transcription and a modification of the user-specific vocabulary within the private transcription [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]. As to claim 3, Robert Jose discloses the method of claim 1, further comprising: receiving a transcript of the one or more audio streams [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]; and generating the private transcription is based on the transcript [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]. As to claim 6, Robert Jose discloses the method of claim 1, further comprising obtaining the vocabulary database from a user profile associated with the user [“The transcription is received from the remote server, it is stored in a data structure, such as a table or database at the local device.” § 0003]. As to claim 8, Robert Jose discloses a system [FIG. 1] comprising: a communications interface [106 on FIG. 1]; a non-transitory computer-readable medium [§ 0128]; and one or more processors [310 on FIG. 3] configured to execute processor-executable instructions stored in the non-transitory computer-readable medium [§ 0128] to: receive, from a remote server, a local model for speech recognition, wherein the local model comprises a copy of a centralized model [“Enabling, on a local device, a voice processing system that limits the amount of data needed to be transmitted to a remote server in translating a voice input into executable data.” § 0003]; perform, by the first client device, speech recognition using the local model on at least one audio stream of the one or more audio streams [“The local speech-to-text model is a neural network model or machine learning model supplied to the local device by a remote server which is pre-trained to recognize a limited set of words corresponding to actions that the local device can perform.” § 0023]; identify, based on a vocabulary database for a user, user-specific vocabulary within the at least one audio stream [“Server processing circuitry is in generating a transcription of the query, identify a plurality of phonemes corresponding to the individual sounds of the query.” § 0026-0027]; and generate, based on the user-specific vocabulary, a private transcription of the one or more audio streams, wherein the private transcription comprises the user-specific vocabulary [“Local device processing circuitry processes the voice query received by the local device via input circuitry using a local speech-to-text model to generate a transcription of the voice query.” § 0028]. Robert Jose fails to disclose join a virtual meeting, by a first client device. However, Gorzela teaches join, by a first client device, a virtual meeting having a plurality of participants, the virtual meeting involving exchanging one or more audio streams between the participants [§ 0023]. Robert Jose and Gorzela are analogous because they are all directed at participant speech recognition management system. One of ordinary skill in the art before the effective filing date of the claimed invention would have found obvious to modify Robert Jose reference with the teaching of Gorzela, so that the local device would include the joining of the meeting option in the system of Robert Jose, would have been combined into a joining of a virtual meeting, for the obvious purpose of providing the user the time for the meeting communications, by combining prior art elements according to known methods to yield predictable results. As to claim 9, Robert Jose discloses the system of claim 8, wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to generate a neutralized transcription based on the private transcription and a modification of the user-specific vocabulary within the private transcription [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]. As to claim 10, Robert Jose discloses the system of claim 8, wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to: receive a transcript of the one or more audio streams [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]; and generate the private transcription is based on the transcript [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]. As to claim 13, Robert Jose discloses the system of claim 8, wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to obtain the vocabulary database from a user profile associated with the user [“The transcription is received from the remote server, it is stored in a data structure, such as a table or database at the local device.” § 0003]. As to claim 15, Robert Jose discloses a non-transitory computer-readable medium comprising processor-executable instructions [§ 0128] configured to cause one or more processors to: receive, from a remote server, a local model for speech recognition, wherein the local model comprises a copy of a centralized model [“Enabling, on a local device, a voice processing system that limits the amount of data needed to be transmitted to a remote server in translating a voice input into executable data.” § 0003]; perform, by the first client device, speech recognition using the local model on at least one audio stream of the one or more audio streams [“The local speech-to-text model is a neural network model or machine learning model supplied to the local device by a remote server which is pre-trained to recognize a limited set of words corresponding to actions that the local device can perform.” § 0023]; identify, based on a vocabulary database for a user, user-specific vocabulary within the at least one audio stream [“Server processing circuitry is in generating a transcription of the query, identify a plurality of phonemes corresponding to the individual sounds of the query.” § 0026-0027]; and generate, based on the user-specific vocabulary, a private transcription of the one or more audio streams, wherein the private transcription comprises the user-specific vocabulary [“Local device processing circuitry processes the voice query received by the local device via input circuitry using a local speech-to-text model to generate a transcription of the voice query.” § 0028]. Robert Jose fails to disclose join a virtual meeting, by a first client device. However, Gorzela teaches join, by a first client device, a virtual meeting having a plurality of participants, the virtual meeting involving exchanging one or more audio streams between the participants [§ 0023]. Robert Jose and Gorzela are analogous because they are all directed at participant speech recognition management system. One of ordinary skill in the art before the effective filing date of the claimed invention would have found obvious to modify Robert Jose reference with the teaching of Gorzela, so that the local device would include the joining of the meeting option in the system of Robert Jose, would have been combined into a joining of a virtual meeting, for the obvious purpose of providing the user the time for the meeting communications, by combining prior art elements according to known methods to yield predictable results. As to claim 16, Robert Jose discloses the non-transitory computer-readable medium of claim 15, further comprising processor-executable instructions configured to cause the one or more processors to generate a neutralized transcription based on the private transcription and a modification of the user-specific vocabulary within the private transcription [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]. As to claim 17, Robert Jose discloses he non-transitory computer-readable medium of claim 15, further comprising processor-executable instructions configured to cause the one or more processors to: receive a transcript of the one or more audio streams [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]; and generate the private transcription is based on the transcript [“Local device processing circuitry uses the metadata to further train the local or private speech-to-text model to recognize future instances of the query.” § 0026-0027]. Allowable Subject Matter Claims 4-5, 7, 11-12, 14 and 18-20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12,165,646 B2. Although the claims at issue are not identical, they are not patentably distinct from each other because at least one claim of the instant application is being taught by the claims of the U.S. Patent. Patented claim 8 recites a method which perform the feature of the generating, based on the user-specific vocabulary, a private transcription of the one or more audio streams, wherein the private transcription comprises the user-specific vocabulary. The pending claim 1 recites a method which perform the similar feature of generating, based on the user-specific vocabulary, a private transcription of the one or more audio streams, wherein the private transcription comprises the user-specific vocabulary. Therefore, the patented claim 8 anticipates the pending 1. Pending claims Patented claims 1. A method comprising: joining, by a first client device, a virtual meeting having a plurality of participants, the virtual meeting involving exchanging one or more audio streams between the participants; receiving, from a remote server, a local model for speech recognition, wherein the local model comprises a copy of a centralized model; performing, by the first client device, speech recognition using the local model on at least one audio stream of the one or more audio streams, wherein performing speech recognition comprises: identifying, based on a vocabulary database for a user, user-specific vocabulary within the at least one audio stream; and generating, based on the user-specific vocabulary, a private transcription of the one or more audio streams, wherein the private transcription comprises the user-specific vocabulary. 8. A method comprising: joining, by a first client device, a virtual meeting having a plurality of participants, each participant of the plurality of participants exchanging one or more audio streams via the virtual meeting; receiving, from a video conference provider, a local model for speech recognition, wherein the local model comprises a copy of a centralized model; storing, by the first client device, the copy of the centralized model as a local recognizer; performing, by the first client device, speech recognition using the local model on the one or more audio streams, wherein performing speech recognition comprises: identifying, by the local recognizer, audio feature data within the one or more audio streams; identifying, based on a vocabulary database, user-specific vocabulary within the audio feature data; and generating, based on the user-specific vocabulary, a private transcription of the one or more audio streams, wherein the private transcription comprises the user-specific vocabulary. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO-892 form. Dines et al. (US 2014/0244252 A1) discloses method for providing participants to a multiparty meeting with a transcript of the meeting, comprising the steps of: establishing an meeting among two or more participants; exchanging during said meeting voice data as well as documents; uploading at least a part of said voice data and at least a part of said documents to a remote speech recognition server. Any inquiry concerning this communication or earlier communications from the examiner should be directed to GERALD GAUTHIER whose telephone number is (571)272-7539. The examiner can normally be reached 8:00 AM to 4:30 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CAROLYN R EDWARDS can be reached at (571) 270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GERALD GAUTHIER/Primary Examiner, Art Unit 2692 September 4, 2026
Read full office action

Prosecution Timeline

Dec 09, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §103, §DP (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748929
SUMMARIZING COLLABORATIVE DATA BASED ON DETERMINED USER INTEREST AND STATUS
3y 10m to grant Granted Sep 29, 2026
Patent 12749481
GENERATING AUDIO USING AUTO-REGRESSIVE GENERATIVE NEURAL NETWORKS
2y 4m to grant Granted Sep 29, 2026
Patent 12749491
AUTHENTICATION BY SPEECH AT A MACHINE
2y 2m to grant Granted Sep 29, 2026
Patent 12746471
SYSTEMS AND METHODS FOR IDENTIFYING A LOCATION OF A SOUND SOURCE
2y 2m to grant Granted Sep 29, 2026
Patent 12739562
SYSTEMS AND METHODS FOR MINIMIZING SPEAKER DISTORTION
3y 6m to grant Granted Sep 15, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
91%
Grant Probability
98%
With Interview (+6.6%)
2y 7m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1834 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month