Prosecution Insights
Last updated: August 17, 2026
Application No. 18/633,092

MODIFYING AUDIO DATA IN A VIRTUAL MEETING TO INCREASE UNDERSTANDABILITY

Non-Final OA §103
Filed
Apr 11, 2024
Examiner
ALBERTALLI, BRIAN LOUIS
Art Unit
2656
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
3 (Non-Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
4m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
706 granted / 862 resolved
+19.9% vs TC avg
Strong +17% interview lift
Without
With
+16.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
20 currently pending
Career history
883
Total Applications
across all art units

Statute-Specific Performance

§101
15.6%
-24.4% vs TC avg
§103
36.5%
-3.5% vs TC avg
§102
25.1%
-14.9% vs TC avg
§112
16.7%
-23.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 862 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 18 June 2026 has been entered. Response to Arguments Applicant’s arguments with respect to the rejection(s) of claim(s) 1, 9, and 17 under 35 U.S.C. 103(a) have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Karishma et al. Applicant has amended claims 1, 9, and 17 to require the AI model to be trained with training data comprising speech disrupted by a speech disorder and target audio data that corresponds to the sample audio data and comprises speech without the speech disorder. As noted by Applicant, Malkin does not disclose training data comprising target audio data that corresponds to the sample audio data and comprises speech without the speech disorder. However, Karishma et al. disclose a method for correcting speech disrupted by a speech disorder, using an AI model trained with training data comprising speech disrupted by a speech disorder and target audio data that corresponds to the sample audio data and comprises speech without the speech disorder. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the method of correcting speech disrupted by a speech disorder disclosed by Karishma et al. in the method of Nguyen for the reasons provided below. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nguyen et al. (U.S. Patent Application Pub. No. 2024/0098218, hereinafter “Nguyen”), in view of Karishma et al. (Dysarthric Speech Enhancement using Hybrid model involving Dynamic Time Warping and Feed Forward Neural Network, hereinafter “Karishma”). In regard to claim 1, Nguyen discloses a method (Fig. 7, 700), comprising: causing a virtual meeting user interface (UI) to be presented during a virtual meeting between a plurality of participants, the virtual meeting UI providing first audio data associated with an audio stream produced by a client device of a first participant of the plurality of participants (see Fig. 5A, a GUI 500 for virtual conferences is displayed, the GUI 500 including multiple participants 502 and 504, paragraph [0088]; the virtual conference associated with an audio stream including audio data from a client device, paragraphs [0104-0106]); determining that the first audio data associated with the audio stream produced by the client device of the first participant is to be modified during the virtual meeting (a request to convert audio from a participant in the virtual conference is received, paragraph [0107]); generating, using an artificial intelligence (AI) model and using the audio stream produced by the client device of the first participant as input to the AI model, a modified audio stream to improve understandability of the first audio data by one or more participants of the plurality of participants (an accent conversion (AC) model generates from a received audio stream including speech in a source accent to a second audio stream including speech in a target accent, paragraph [0112]; the AC model comprising one or more machine learning models, paragraphs [0013-0014]; allowing participants to more easily understand each other, paragraphs [0016-0017]); and causing second audio data associated with the modified audio stream to be provided during the virtual meeting in place of the first audio data (the second audio stream is transmitted to client devices in the virtual conference, paragraph [0115]). Nguyen does not expressly disclose the modification to improve understandability is applied to first audio data comprising speech disrupted by a speech disorder and that the modified audio stream comprises speech with a removed speech disorder. Karishma discloses a method for correcting speech disrupted by a speech disorder to generate a modified audio stream comprising speech with a removed speech disorder (correction of dysarthric speech, see Abstract) using an AI model, wherein the Al model is trained on a plurality of items of training data to correct speech disrupted by a speech disorder (a feed forward neural network is trained using training data, section III, first paragraph and section III-B), and wherein an item of training data comprises sample audio data comprising speech disrupted by a speech disorder and target audio data that corresponds to the sample audio data and comprises speech without the speech disorder (input training data comprises dysarthric speech and the corresponding target data comprises normal speech, section III, first paragraph and section III-B). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Nguyen to generate a modified audio stream comprising speech with a removed speech disorder using a trained AI model trained using speech disrupted by a speech disorder and target audio comprising speech without the disorder, because the method objectively enhances the dysarthric speech more effectively than alternative methods, as taught by Karishma (sections IV and V). In regard to claim 2, Nguyen does not disclose the item of training data comprises speech disrupted by the speech disorder and target audio data that comprises speech without the speech disorder. Karishma discloses the item of training data comprises: the sample audio data comprising the speech disrupted by a speech disorder (input dysarthric speech, section III-B); and the target audio data that comprises speech without the speech disorder (target normal speech, section III-B), wherein the sample audio data and the target audio data are associated with the same speaker or with different speakers having similar vocal qualities (the speech signals comprise female dysarthric speech signals matched with female non-dysarthric speech signals and male dysarthric speech signals matched with male non-dysarthric speech signals, section III; where dynamic time warping is performed on the spectrogram of the speech signals to determine a best match between any two provided sequences, section III-A). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use training data wherein the sample audio data and the target audio data are associated with the same speaker or with different speakers having similar vocal qualities, because the method objectively enhances the dysarthric speech more effectively than alternative methods, as taught by Karishma (sections IV and V). In regard to claim 3, Nguyen does not disclose the speech comprises a speech disorder. Malik discloses the speech disorder comprises at least one of verbal apraxia; cluttering; aphasia; stuttering; or a speech sound disorder (stutters, etc., paragraphs [0014], [0030], and [0032]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to remove speech disorders comprising at least one of the above speech disorders, because it would allow the participant’s speech impairment to be corrected and remove impairments that are not intended as part of speech when the participant speaks, as suggested by Malik (paragraphs [0013-0014]). In regard to claim 4, Nguyen discloses generating the modified audio stream comprises using the AI model to perform at least one of: increase a pitch of the audio stream; or change a timbre of the audio stream (the AC model is trained to allow a user to select a desired voice as well as an accent, paragraph [0114]; where performing voice conversion (VC) comprises adjusting the pitch and timbre of the audio, paragraph [0073]). In regard to claim 5, Nguyen discloses determining that the first audio data associated with the audio stream produced by the client device of the first participant is to be modified comprises receiving a command from the client device of the first participant (a participant in the virtual conference requests that their accent is modified, paragraph [0066]). In regard to claim 6, Nguyen discloses the command comprises data indicating an audio effect to be applied by the AI model (accent conversion, paragraph [0066]). In regard to claim 7, Nguyen discloses determining that the first audio data associated with the audio stream produced by the client device of the first participant is to be modified comprises receiving a command from a client device of a second participant of the plurality of participants (a participant in the virtual conference requests that a second participant’s accent is converted, paragraph [0066]). In regard to claim 8, Nguyen discloses causing the second audio data associated with the modified audio stream to be provided during the virtual meeting in place of the first audio data comprises causing, for a subset of the plurality of participants, the second audio data to be provided in place of the first audio data (multiple participants in the virtual conference request accent conversion to a particular accent, paragraph [0116]). In regard to claim 9, Nguyen discloses a system (Fig. 8, 800), comprising: a memory (memory 820); and a processing device (processor 810), coupled to the memory, configured to perform operations comprising: causing a virtual meeting user interface (UI) to be presented during a virtual meeting between a plurality of participants, the virtual meeting UI providing first audio data associated with an audio stream produced by a client device of a first participant of the plurality of participants (see Fig. 5A, a GUI 500 for virtual conferences is displayed, the GUI 500 including multiple participants 502 and 504, paragraph [0088]; the virtual conference associated with an audio stream including audio data from a client device, paragraphs [0104-0106]); determining that the first audio data associated with the audio stream produced by the client device of the first participant is to be modified during the virtual meeting (a request to convert audio from a participant in the virtual conference is received, paragraph [0107]); generating, using an artificial intelligence (AI) model and using the audio stream produced by the client device of the first participant as input to the AI model, a modified audio stream to improve understandability of the first audio data by one or more participants of the plurality of participants (an accent conversion (AC) model generates from a received audio stream including speech in a source accent to a second audio stream including speech in a target accent, paragraph [0112]; the AC model comprising one or more machine learning models, paragraphs [0013-0014]; allowing participants to more easily understand each other, paragraphs [0016-0017]); and causing second audio data associated with the modified audio stream to be provided during the virtual meeting in place of the first audio data (the second audio stream is transmitted to client devices in the virtual conference, paragraph [0115]). Nguyen does not expressly disclose the modification to improve understandability is applied to first audio data comprising speech disrupted by a speech disorder and that the modified audio stream comprises speech with a removed speech disorder. Karishma discloses a method for correcting speech disrupted by a speech disorder to generate a modified audio stream comprising speech with a removed speech disorder (correction of dysarthric speech, see Abstract) using an AI model, wherein the Al model is trained on a plurality of items of training data to correct speech disrupted by a speech disorder (a feed forward neural network is trained using training data, section III, first paragraph and section III-B), and wherein an item of training data comprises sample audio data comprising speech disrupted by a speech disorder and target audio data that corresponds to the sample audio data and comprises speech without the speech disorder (input training data comprises dysarthric speech and the corresponding target data comprises normal speech, section III, first paragraph and section III-B). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Nguyen to generate a modified audio stream comprising speech with a removed speech disorder using a trained AI model trained using speech disrupted by a speech disorder and target audio comprising speech without the disorder, because the method objectively enhances the dysarthric speech more effectively than alternative methods, as taught by Karishma (sections IV and V). In regard to claim 10, Nguyen does not disclose the item of training data comprises speech disrupted by the speech disorder and target audio data that comprises speech without the speech disorder. Karishma discloses the item of training data comprises: the sample audio data comprising the speech disrupted by a speech disorder (input dysarthric speech, section III-B); and the target audio data that comprises speech without the speech disorder (target normal speech, section III-B), wherein the sample audio data and the target audio data are associated with the same speaker or with different speakers having similar vocal qualities (the speech signals comprise female dysarthric speech signals matched with female non-dysarthric speech signals and male dysarthric speech signals matched with male non-dysarthric speech signals, section III; where dynamic time warping is performed on the spectrogram of the speech signals to determine a best match between any two provided sequences, section III-A). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use training data wherein the sample audio data and the target audio data are associated with the same speaker or with different speakers having similar vocal qualities, because the method objectively enhances the dysarthric speech more effectively than alternative methods, as taught by Karishma (sections IV and V). In regard to claim 11, Nguyen does not disclose the speech comprises a speech disorder. Malik discloses the speech disorder comprises at least one of verbal apraxia; cluttering; aphasia; stuttering; or a speech sound disorder (stutters, etc., paragraphs [0014], [0030], and [0032]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to remove speech disorders comprising at least one of the above speech disorders, because it would allow the participant’s speech impairment to be corrected and remove impairments that are not intended as part of speech when the participant speaks, as suggested by Malik (paragraphs [0013-0014]). In regard to claim 12, Nguyen discloses generating the modified audio stream comprises using the AI model to perform at least one of: increase a pitch of the audio stream; or change a timbre of the audio stream (the AC model is trained to allow a user to select a desired voice as well as an accent, paragraph [0114]; where performing voice conversion (VC) comprises adjusting the pitch and timbre of the audio, paragraph [0073]). In regard to claim 13, Nguyen discloses determining that the first audio data associated with the audio stream produced by the client device of the first participant is to be modified comprises receiving a command from the client device of the first participant (a participant in the virtual conference requests that their accent is modified, paragraph [0066]). In regard to claim 14, Nguyen discloses the command comprises data indicating an audio effect to be applied by the AI model (accent conversion, paragraph [0066]). In regard to claim 15, Nguyen discloses determining that the first audio data associated with the audio stream produced by the client device of the first participant is to be modified comprises receiving a command from a client device of a second participant of the plurality of participants (a participant in the virtual conference requests that a second participant’s accent is converted, paragraph [0066]). In regard to claim 16, Nguyen discloses causing the second audio data associated with the modified audio stream to be provided during the virtual meeting in place of the first audio data comprises causing, for a subset of the plurality of participants, the second audio data to be provided in place of the first audio data (multiple participants in the virtual conference request accent conversion to a particular accent, paragraph [0116]). In regard to claim 17, Nguyen discloses a method (Fig. 7, 700), comprising: causing a virtual meeting user interface (UI) to be presented during a virtual meeting between a plurality of participants, the virtual meeting UI providing a plurality of at a plurality of time periods during the virtual meeting, wherein each first audio data of the plurality of first audio data is associated with an audio stream produced by a client device of a respective participant of the plurality of participants (see Fig. 5A, a GUI 500 for virtual conferences is displayed, the GUI 500 including multiple participants 502 and 504, paragraph [0088]; the virtual conference associated with an audio stream including audio data from a plurality of client devices, paragraphs [0104-0106]); determining that the plurality of first audio data are to be modified during the virtual meeting (a request to convert audio to a target accent for each of a plurality of participants is received, paragraphs [0107-0108]); generating, using a plurality of artificial intelligence (AI) models and using the audio streams of the plurality of participants as input to the AI models, a plurality of modified audio streams (accent conversion (AC) models generate from received audio streams including speech in sources accents, second audio streams including speech in a target accent, paragraphs [0090] and [0012]; the AC model comprising one or more machine learning models, paragraphs [0013-0014]), wherein each modified audio stream is associated with a participant of the plurality of participants (any participants with a different accent will be converted, paragraph [0090]), and the respective modified audio streams improve understandability of the respective first audio data by one or more participants of the plurality of participants (the conversion allows participants to more easily understand each other, paragraphs [0016-0017]); and causing a plurality of second audio data associated with the plurality of modified audio streams to be provided during the virtual meeting in place of the plurality of first audio data (the second audio streams are transmitted to client devices in the virtual conference, paragraphs [0090] and [0115]). Nguyen does not expressly disclose the modification to improve understandability is applied to first audio data comprising speech disrupted by a speech disorder and that the modified audio stream comprises speech with a removed speech disorder. Karishma discloses a method for correcting speech disrupted by a speech disorder to generate a modified audio stream comprising speech with a removed speech disorder (correction of dysarthric speech, see Abstract) using an AI model, wherein the Al model is trained on a plurality of items of training data to correct speech disrupted by a speech disorder (a feed forward neural network is trained using training data, section III, first paragraph and section III-B), and wherein an item of training data comprises sample audio data comprising speech disrupted by a speech disorder and target audio data that corresponds to the sample audio data and comprises speech without the speech disorder (input training data comprises dysarthric speech and the corresponding target data comprises normal speech, section III, first paragraph and section III-B). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Nguyen to generate a modified audio stream comprising speech with a removed speech disorder using a trained AI model trained using speech disrupted by a speech disorder and target audio comprising speech without the disorder, because the method objectively enhances the dysarthric speech more effectively than alternative methods, as taught by Karishma (sections IV and V). In regard to claim 18, Nguyen discloses the plurality of AI models comprises a first AI model and a second AI model (one or more AC processes comprising the trained AC model, paragraph [0065]); the first AI model applies an audio effect to a first audio stream of the audio streams (an AC process is generated for each participant requesting an accent conversion, paragraphs [0068-0069]); and the second AI model applies the same audio effect to a second audio stream of the audio streams (an AC process is generated for each participant requesting an accent conversion, paragraphs [0068-0069]). In regard to claim 19, Nguyen discloses the plurality of AI models comprises a first AI model and a second AI model (one or more AC processes comprising the trained AC model, paragraph [0065]); the first AI model applies a first audio effect to a first audio stream of the audio streams (an AC process is generated for each participant requesting an accent conversion, paragraphs [0068-0069]); and the second AI model applies a second audio effect to a second audio stream of the audio streams, wherein the second audio effect is different from the first audio effect (an AC process for each source-target accent pair is generated, paragraphs [0068-0069]). In regard to claim 20, Nguyen discloses determining that the plurality of first audio data is to be modified comprises receiving a command from a client device of a first participant of the plurality of participants (a participant in the virtual conference requests accent conversion for a plurality of participants, paragraph [0066]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Biadsy et al. and Wang et al. disclose additional methods for correcting speech signals comprising a speech disorder. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIAN LOUIS ALBERTALLI whose telephone number is (571)272-7616. The examiner can normally be reached M-F 8AM-3PM, 4PM-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached at 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. BLA 7/21/26 /BRIAN L ALBERTALLI/Primary Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Show 5 earlier events
Feb 10, 2026
Response Filed
Mar 19, 2026
Final Rejection mailed — §103
Jun 03, 2026
Interview Requested
Jun 10, 2026
Applicant Interview (Telephonic)
Jun 10, 2026
Examiner Interview Summary
Jun 18, 2026
Request for Continued Examination
Jun 22, 2026
Response after Non-Final Action
Jul 23, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706095
METHODS AND SYSTEMS FOR REDUCING LATENCY IN AUTOMATED ASSISTANT INTERACTIONS
2y 1m to grant Granted Aug 11, 2026
Patent 12700399
METHOD AND APPARATUS FOR TRAINING ENCODER
2y 3m to grant Granted Aug 04, 2026
Patent 12673585
VIBRATION SENSING STEERING WHEEL TO OPTIMIZE VOICE COMMAND ACCURACY
3y 7m to grant Granted Jul 07, 2026
Patent 12658189
METHOD FOR RESPONDING TO CONTROL VOICE, DEVICE, AND STORAGE MEDIUM
2y 9m to grant Granted Jun 16, 2026
Patent 12646517
VIRTUAL REALITY HEADSET AND ARTIFICIAL INTELLIGENCE VIRTUAL ASSISTANT INTEGRATION FOR ADDRESSING A LANGUAGE BARRIER WITH A CUSTOMER
2y 0m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
98%
With Interview (+16.6%)
2y 9m (~4m remaining)
Median Time to Grant
High
PTA Risk
Based on 862 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month