Prosecution Insights
Last updated: August 15, 2026
Application No. 18/788,269

METHODS FOR REAL-TIME ACCENT CONVERSION AND SYSTEMS THEREOF

Final Rejection §103
Filed
Jul 30, 2024
Priority
May 06, 2021 — provisional 63/185,345 +2 more
Examiner
BLANKENAGEL, BRYAN S
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Sanas AI Inc.
OA Round
2 (Final)
67%
Grant Probability
Favorable
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
261 granted / 389 resolved
+5.1% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
22 currently pending
Career history
411
Total Applications
across all art units

Statute-Specific Performance

§101
24.9%
-15.1% vs TC avg
§103
50.5%
+10.5% vs TC avg
§102
12.1%
-27.9% vs TC avg
§112
7.5%
-32.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 389 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments filed 06/22/2026 have been fully considered but they are not persuasive. Regarding arguments on pages 7-8 of the Remarks, Examiner notes that in the references Office Action from 09/18/2023, the limitation in question is not fully quoted. The limitation includes details about the training that are not taught by Dirac, but are also not claimed in the instant application. Since the limitation of the current claims remain broader than the referenced limitation, Dirac is still relied upon to teach them. Examiner further notes that regarding the use of machine learning, Dirac teaches that the accent translation models are updated and refined using machine learning by using new audio samples. Therefore, interpreting the accent translation models of Dirac as the machine-learning algorithm of the claims, Dirac teaches that the models are updated or trained using the new audio samples, the audio samples corresponding to each of the different accent sample sets corresponding to the different speakers. For example, Dirac col. 6 line 56 – col. 7 line 6 teaches different speakers with different accents, as well as multiple speakers for each accent. Regarding arguments on pages 9-10 of the Remarks, Examiner notes that while Peng does teach the alignment of transcriptions, Peng para [0026] also teaches “In some examples, the alignment can include aligning speech frames, or snippets, between the expected phonetic transcription and the actual phonetic transcription. In some examples, alignment can include identifying speech frames, or snippets, based on context, e.g., surrounding phonemes, and aligning the speech frames based on the context.” Therefore, the alignment of speech frames is taught. Solving a problem in a different manner does not constitute teaching away, as the claims do not expressly forbid the conversion of speech to text. Further, the use of speech to text in this case is limited to only the alignment step of the claims, as the other limitations are taught by the Chang and Dirac references. Indeed, para [0032] of the Specification refers to the use of a transcription when performing the alignment and classification. Therefore, the combination of references is proper. While Dirac may not mention conversion of speech to text, this is not the same as teaching away. Dirac similarly does not teach alignment. Therefore, using the Peng reference to teach alignment using transcription does not destroy the combination in any way. Regarding arguments on pages 10-11 of the Remarks, Examiner notes that Peng solving a problem not contemplated by Dirac has no bearing on the combination of references. Rather, the combination is to show that if a person of ordinary skill in the art wanted to perform the steps of the instant application, that the references were available and could be combined. For example, one of ordinary skill in the art would understand that speech-to-text can be performed on speech, and would result in a transcription that could be used for alignment of the speech frames. In the combination of references, the alignment of Peng would be used on the captured speech content of different speakers having the same accent as taught by Dirac Fig. 1. Regarding the classification, Examiner notes that classifying the accent of a speaker extends to classifying the frames of speech as belonging to the speaker with the accent. Regarding arguments on pages 11-12 of the Remarks, Examiner notes that the claims do not require that the first and second phonemes be different. By modifying the phoneme formant stream to shape the formant patterns, the pronunciation of the phonemes would be changed. One of ordinary skill in the art would understand that the concepts of dialect, accent, phonemes, and pronunciation are all interconnected, and for example, a person with a different accent would pronounce a word differently than someone with a different accent. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4, 6-7, 15-17, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang et al. (US 2014/0187210 A1), hereinafter referred to as Chang, in view of Dirac et al. (US 10,163,451 B1), hereinafter referred to as Dirac. Regarding claim 1, Chang teaches: A system, comprising memory having instructions stored thereon and one or more processors coupled to the memory (Fig. 10 elements 1020, 1030, para [0069], where the components of the system are used) and configured to execute the instructions to: apply a first machine-learning algorithm to second speech content comprising a set of phonemes associated with a first pronunciation of the second speech content to generate an output, wherein the first machine-learning algorithm is trained with fourth audio data comprising first speech content captured from a first plurality of different speakers having a first accent (Fig. 8, para [0063], where a speaker configures their voice to have an accent removed, and para [0031-32], where the phonemes are used to determine the dialect of the speaker); synthesize, using the output and a second machine-learning algorithm trained with first audio data comprising the first accent and second audio data comprising a second accent, third audio data representative of the second speech content having the second accent (Fig. 8, para [0040], [0054], [0064], where another speaker chooses to apply voice modifications to other speakers for synthesis in another accent); convert the third audio data into a synthesized version of the second speech content having the second accent (para [0035], where the synthesizer converts the spectrogram to a time-domain digital speech signal); and output the synthesized version of the second speech content via a digital communication application executed by the system (Fig. 3 element 340, para [0035], where the synthesized speech signal is output, and para [0018], [0022], where applications are used). Chang does not teach: wherein the first machine-learning algorithm is trained with first speech content from a first plurality of speakers having a first accent, using the output and a second machine-learning algorithm trained with first audio data comprising the first accent and second audio data comprising a second accent, Dirac teaches: wherein the first machine-learning algorithm is trained with first speech content from a first plurality of speakers having a first accent (Fig. 1 element 131-134, col. 3 lines 36-60, where sample sets for different accents are used, and col. 7 lines 52-60, where machine learning is used to update and refine the accent translation models) using the output and a second machine-learning algorithm trained with first audio data comprising the first accent and second audio data comprising a second accent (Fig. 3 element 131-132, 321-322, col. 5 line 63 - col. 6 line 16, where sample sets for different accents are used, and accent translation models use both accents, and col. 7 lines 52-60, where machine learning is used to update and refine the accent translation models), It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chang by using the machine learning of Dirac (Dirac col. 7 lines 52-60) in the accent modification of Chang (Chang para [0063-64]), in order to refine the accent translation models continually using new audio samples (Dirac col. 7 lines 52-60). Regarding claim 2, Chang in view of Dirac teaches: The system of claim 1, wherein the digital communication application comprises a video-based communication platform (Chang para [0065], where a video conference call is performed). Regarding claim 3, Chang in view of Dirac teaches: The system of claim 1, further comprising a hardware microphone, wherein the instructions comprise a virtual microphone that is executable by the one or more processors to obtain the second speech content from the hardware microphone after the second speech content is captured by the hardware microphone from a user of the system (Dirac Fig. 6 elements 621, 631, col. 8 lines 16-32, where a hardware microphone captures audio, and the providing of the audio to the accent translation components is considered use of a virtual microphone). Regarding claim 4, Chang in view of Dirac teaches: The system of claim 3, wherein the virtual microphone is further executable by the one or more processors to apply the first machine-learning algorithm, synthesize the third audio data, convert the third audio data, and route the synthesized version of the second speech content to the digital communication application (Dirac Fig. 6 elements 631, 622, 632, col. 8 lines 16-32, where the accent translation components translate the input audio and provide to be output to the second party device). Regarding claim 6, Chang in view of Dirac teaches: The system of claim 1, wherein the one or more processors are further configured to execute the instructions to map at least a first non-text linguistic representation of a first phoneme of the set of phonemes to a second non-text linguistic representation of a second phoneme associated with a second pronunciation of the second speech content, wherein the synthesized version of the second speech content further comprises the second phoneme and the first and second phonemes are different phonemes (Chang para [0031], where phonemes are distinguished by a unique pattern or signature in a spectrogram, and para [0035], where the phoneme formant stream is modified to shape formant patterns to match phonemes in another dialect). Regarding claim 7, Chang in view of Dirac teaches: The system of claim 6, wherein the one or more processors are further configured to execute the instructions to map one or more frames in the output to one or more corresponding frames in the second non-text linguistic representation (Chang para [0031], [0035], where a spectrogram represents the spectra of the frames, and where the locations of the phonemes are modified or shaped to form the modified spectrogram). Regarding claim 15, Chang teaches: One or more non-transitory computer-readable media comprising instructions that, when executed by at least one processor (Fig. 10, elements 1020, 1030, para [0069], [0073], where the components of the system are used), cause the at least one processor to: apply a first machine-learning algorithm to first speech content comprising first phonemes associated with a first pronunciation of the first speech content to derive a non-text linguistic representation of the first phonemes (Fig. 8, para [0063], where a speaker configures their voice to have an accent removed, and para [0031-32], where the phonemes are used to determine the dialect of the speaker, and are identified using formants); synthesize, based on the non-text linguistic representation of the first phonemes (para [0035], where the formants in the spectrogram are modified) and using a second machine-learning algorithm trained with first audio data comprising a first accent and second audio data comprising a second accent, third audio data representative of the first speech content having the second accent (Fig. 8, para [0040], [0054], [0064], where another speaker chooses to apply voice modifications to other speakers for synthesis in another accent), wherein the synthesizing comprises mapping at least a first non-text linguistic representation of a first phoneme of the first phonemes associated with the first pronunciation of the first speech content to a second non-text linguistic representation of a second phoneme of an updated set of phonemes associated with a second pronunciation of the first speech content that is different from the first pronunciation of the first speech content (para [0031], where phonemes are distinguished by a unique pattern or signature in a spectrogram, and para [0035], where the phoneme formant stream is modified to shape formant patterns to match phonemes in another dialect); convert the third audio data into a synthesized version of the first speech content having the second accent and comprising the updated set of phonemes (para [0035], where the synthesizer converts the spectrogram to a time-domain digital speech signal); and output the synthesized version of the first speech content via a digital communication application via which the first speech content was received (Fig. 3 element 340, para [0035], where the synthesized speech signal is output, and para [0018], [0022], where applications are used). Chang does not teach: a first machine-learning algorithm; using a second machine-learning algorithm trained with first audio data comprising a first accent and second audio data comprising a second accent, Dirac teaches: a first machine-learning algorithm (Fig. 1 element 131-134, col. 3 lines 36-60, where sample sets for different accents are used, and col. 7 lines 52-60, where machine learning is used to update and refine the accent translation models); using a second machine-learning algorithm trained with first audio data comprising a first accent and second audio data comprising a second accent (Fig. 3 element 131-132, 321-322, col. 5 line 63 - col. 6 line 16, where sample sets for different accents are used, and accent translation models use both accents, and col. 7 lines 52-60, where machine learning is used to update and refine the accent translation models), It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chang by using the machine learning of Dirac (Dirac col. 7 lines 52-60) in the accent modification of Chang (Chang para [0063-64]), in order to refine the accent translation models continually using new audio samples (Dirac col. 7 lines 52-60). Regarding claim 16, Chang in view of Dirac teaches: The non-transitory computer-readable media of claim 15, wherein the digital communication application comprises a video-based communication platform (Chang para [0065], where a video conference call is performed) and the instructions, when executed by the at least one processor, further cause the at least one processor to: obtain by a virtual microphone the first speech content from a hardware microphone after the first speech content is captured by the hardware microphone (Dirac Fig. 6 elements 621, 631, col. 8 lines 16-32, where a hardware microphone captures audio, and the providing of the audio to the accent translation components is considered use of a virtual microphone); and route by the virtual microphone the synthesized version of the first speech content to the digital communication application (Dirac Fig. 6 elements 631, 622, 632, col. 8 lines 16-32, where the accent translation components translate the input audio and provide to be output to the second party device). Regarding claim 17, Chang in view of Dirac teaches: The non-transitory computer-readable media of claim 16, wherein the instructions, when executed by the at least one processor, are further configured to execute the virtual microphone to apply the first machine-learning algorithm, synthesize the third audio data, convert the third audio data, and route the synthesized version of the first speech content to the digital communication application (Dirac Fig. 6 elements 631, 622, 632, col. 8 lines 16-32, where the accent translation components translate the input audio and provide to be output to the second party device). Regarding claim 19, Chang in view of Dirac teaches: The non-transitory computer-readable media of claim 15, wherein the first speech content further comprises a set of prosodic features, the instructions, when executed by the at least one processor further causes the at least one processor to synthesize the third audio data and the set of prosodic features, and the synthesized version of the first speech content has the set of prosodic features (Chang para [0035], where the synthesized voice removes the accent while still sounding like the speaker, and para [0046], where tone and tempo are options for modification). Regarding claim 20, Chang in view of Dirac teaches: The non-transitory computer-readable media of claim 15, wherein the instructions, when executed by the at least one processor further causes the at least one processor to transmit the synthesized version of the first speech content to a computing device via the digital communication application and one or more networks (Chang Fig. 8, para [0064], where the modified audio is transmitted to a speaker in a conference call). Claim(s) 5, 8-12, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang, in view of Dirac, and further in view of Peng et al. (US 2015/0170642 A1), hereinafter referred to as Peng. Regarding claim 5, Chang in view of Dirac teaches: The system of claim 1, wherein the first audio data corresponds to a second plurality of speakers having the first accent and the second audio data corresponds to a single speaker having the second accent (Dirac Fig. 1 element 131-134, col. 3 lines 36-60, where sample sets for different accents using various individuals are used, and Chang para [0032], where the dialect of a speaker is known ahead of time). Chang in view of Dirac does not teach: wherein the one or more processors are further configured to execute the instructions to align and classify each of a plurality of frames of the first speech content corresponding to respective ones of a first plurality of speakers to train the first machine-learning algorithm, Peng teaches: wherein the one or more processors are further configured to execute the instructions to align and classify each of a plurality of frames of the first speech content corresponding to respective ones of a first plurality of speakers to train the first machine-learning algorithm (para [0025-26], where alignment of frames is performed, and para [0030], where the accent group identifier is determined), It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chang in view of Dirac by using the training of Peng (Peng para [0018]) on the system of Chang in view of Dirac (Chang para [0063-64]) to associate actual pronunciations to an expected pronunciation and perform replacements of substitutions before processing an utterance (Peng para [0003]). Regarding claim 8, Chang teaches: A method implemented by one or more computing devices and comprising: applying a first machine-learning algorithm to second speech content comprising a first set of phonemes associated with a first pronunciation of the second speech content (Fig. 8, para [0063], where a speaker configures their voice to have an accent removed, and para [0031-32], where the phonemes are used to determine the dialect of the speaker), synthesizing, based on an output of the first machine-learning algorithm and using a second machine-learning-algorithm trained with first audio data comprising the first accent and second audio data comprising a second accent, third audio data representative of the second speech content having the second accent (Fig. 8, para [0040], [0054], [0064], where another speaker chooses to apply voice modifications to other speakers for synthesis in another accent); converting the third audio data into a synthesized version of the second speech content having the second accent (para [0035], where the synthesizer converts the spectrogram to a time-domain digital speech signal); and outputting the synthesized version of the second speech content via a digital communication application executed by the one or more computing devices (Fig. 3 element 340, para [0035], where the synthesized speech signal is output, and para [0018], [0022], where applications are used). Chang does not teach: a first machine-learning algorithm; using a second machine-learning-algorithm trained with first audio data comprising the first accent and second audio data comprising a second accent, wherein the first machine-learning algorithm is trained based on an alignment and classification of frames of first speech content corresponding to respective speakers having a first accent, Dirac teaches: a first machine-learning algorithm; using a second machine-learning-algorithm trained with first audio data comprising the first accent and second audio data comprising a second accent (Fig. 3 element 131-132, 321-322, col. 5 line 63 - col. 6 line 16, where sample sets for different accents are used, and accent translation models use both accents, and col. 7 lines 52-60, where machine learning is used to update and refine the accent translation models), It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chang by using the machine learning of Dirac (Dirac col. 7 lines 52-60) in the accent modification of Chang (Chang para [0063-64]), in order to refine the accent translation models continually using new audio samples (Dirac col. 7 lines 52-60). Peng teaches: wherein the first machine-learning algorithm is trained based on an alignment and classification of frames of first speech content corresponding to respective speakers having a first accent (para [0025-26], where alignment of frames is performed, and para [0030], where the accent group identifier is determined); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chang in view of Dirac by using the training of Peng (Peng para [0018]) on the system of Chang in view of Dirac (Chang para [0063-64]) to associate actual pronunciations to an expected pronunciation and perform replacements of substitutions before processing an utterance (Peng para [0003]). Regarding claim 9, Chang in view of Dirac and Peng teaches: The method of claim 8, wherein the digital communication application comprises a video-based communication platform (Chang para [0065], where a video conference call is performed). Regarding claim 10, Chang in view of Dirac and Peng teaches: The method of claim 8, further comprising: obtaining by a virtual microphone the second speech content from a hardware microphone of the one or more computing devices after the second speech content is captured by the hardware microphone from a user (Dirac Fig. 6 elements 621, 631, col. 8 lines 16-32, where a hardware microphone captures audio, and the providing of the audio to the accent translation components is considered use of a virtual microphone); and routing by the virtual microphone the synthesized version of the second speech content to the digital communication application (Dirac Fig. 6 elements 631, 622, 632, col. 8 lines 16-32, where the accent translation components translate the input audio and provide to be output to the second party device). Regarding claim 11, Chang in view of Dirac and Peng teaches: The method of claim 8, wherein the first audio data corresponds to a second plurality of speakers having the first accent and the second audio data corresponds to a single speaker having the second accent (Dirac Fig. 1 element 131-134, col. 3 lines 36-60, where sample sets for different accents using various individuals are used, and Chang para [0032], where the dialect of a speaker is known ahead of time). Regarding claim 12, Chang in view of Dirac and Peng teaches: The method of claim 8, further comprising mapping at least a first non-text linguistic representation of a first phoneme of the first set of phonemes to a second non-text linguistic representation of a second phoneme of a second set of phonemes associated with a second pronunciation of the second speech content to facilitate the synthesizing, wherein the synthesized version of the second speech content comprises the second set of phonemes (Chang para [0031], where phonemes are distinguished by a unique pattern or signature in a spectrogram, and para [0035], where the phoneme formant stream is modified to shape formant patterns to match phonemes in another dialect). Regarding claim 18, Chang in view of Dirac teaches: The non-transitory computer-readable media of claim 15, wherein the first audio data corresponds to a second plurality of speakers having the first accent and the second audio data corresponds to a single speaker having the second accent (Dirac Fig. 1 element 131-134, col. 3 lines 36-60, where sample sets for different accents using various individuals are used, and Chang para [0032], where the dialect of a speaker is known ahead of time). Chang in view of Dirac does not teach: wherein the instructions, when executed by the at least one processor, further cause the at least one processor to align and classify each of a plurality of frames of the first speech content corresponding to respective ones of a first plurality of speakers to train the first machine-learning algorithm, Peng teaches: wherein the instructions, when executed by the at least one processor, further cause the at least one processor to align and classify each of a plurality of frames of the first speech content corresponding to respective ones of a first plurality of speakers to train the first machine-learning algorithm (para [0025-26], where alignment of frames is performed, and para [0030], where the accent group identifier is determined), It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Chang in view of Dirac by using the training of Peng (Peng para [0018]) on the system of Chang in view of Dirac (Chang para [0063-64]) to associate actual pronunciations to an expected pronunciation and perform replacements of substitutions before processing an utterance (Peng para [0003]). Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang, in view of Dirac, and Peng, and further in view of Feinauer et al. (US 2020/0193971 A1), hereinafter referred to as Feinauer. Regarding claim 13, Chang in view of Dirac and Peng teaches: The method of claim 8, further comprising converting the synthesized third audio data into a synthesized version of third speech content having the second accent between 50-700 ms after receiving the third speech content having the first accent, wherein the synthesized version of the third speech content has the second accent (Chang para [0031], where accents consist of phonemes that are distinguished by a unique pattern of signature in a spectrogram). Chang in view of Dirac and Peng does not teach: between 50-700 ms; Feinauer teaches: between 50-700 ms (col. 22 line 60 - col. 23 line 9, where an example of 100 ms is determined as sounding like real time) Chang in view of Dirac and Peng teaches accent conversion in real time (Chang para [0015]). However, claim 13 recites that the conversion is performed between 50-700 ms after receiving the content. Feinauer teaches a lag duration that sounds like real time to a user, with an example of 100 ms (Feinauer col. 22 line 60 - col. 23 line 9). The cited section of Feinauer gives the value of 100 ms only as an example, recognizing that any value that is perceived as real time would suffice, and selection of such would be within the level of ordinary skill in the art, and would lead to predictable results. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have substituted the specific time value of Feinauer in place of the real time designation of Chang in view of Dirac and Peng, where the result of the substitution would predictably result in a non-noticeable delay. Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chang, in view of Dirac, and Peng, and further in view of Ganapathiraju et al. (US 2014/0025379 A1), hereinafter referred to as Ganapathiraju. Regarding claim 14, Chang in view of Dirac and Peng teaches: The method of claim 8, wherein the first machine-learning algorithm comprises a non-text learned linguistic representation for the first accent (Chang para [0031], where accents consist of phonemes that are distinguished by a unique pattern of signature in a spectrogram) and the method further comprises: aligning and classifying each of the frames according to phoneme sounds of the first speech content to train the first machine-learning algorithm (Peng para [0025-26], where alignment of frames is performed, and para [0030], where the accent group identifier is determined); and detecting, for each of another plurality of frames in the second speech content, a respective phoneme sound based on the non-text learned linguistic representation (Chang para [0031], where phonemes are distinguished by a unique pattern or signature in a spectrogram, and para [0035], where the phoneme formant stream is modified to shape formant patterns to match phonemes in another dialect). Chang in view of Dirac and Peng does not teach: monophone and triphone sounds; Ganapathiraju teaches: monophone and triphone sounds (para [0023], where monophones and triphones are both used); Chang in view of Dirac and Peng teaches processing using phonemes (Chang para [0031]). However, claim 1 recites that monophones and triphones are used. Ganapathiraju teaches using monophones and triphones (Ganapathiraju para [0023]). Para [0023] recognizes that phonemes can be modeled in isolation or in context of other phonemes, both being within the level of ordinary skill in the art, and predictable in usage. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have substituted the monophones and triphones of Ganapathiraju in place of the phonemes of Chang in view of Dirac and Peng, where the result of the substitution would predictably allow for processing of individual phonemes or phonemes in context of other phonemes. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 2021/0097974 A1 para [0048] teaches a one-to-many mapping model to learn to synthesize with appropriate pronunciations in different contexts. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRYAN S BLANKENAGEL whose telephone number is (571)270-0685. The examiner can normally be reached 8:00am-5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /BRYAN S BLANKENAGEL/Primary Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Jul 30, 2024
Application Filed
Mar 25, 2026
Non-Final Rejection mailed — §103
Jun 22, 2026
Response Filed
Jul 21, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694883
PERIODIC-COMBINED-ENVELOPE-SEQUENCE GENERATION DEVICE, PERIODIC-COMBINED-ENVELOPE-SEQUENCE GENERATION METHOD, PERIODIC-COMBINED-ENVELOPE-SEQUENCE GENERATION PROGRAM AND RECORDING MEDIUM
2y 9m to grant Granted Jul 28, 2026
Patent 12682908
AUDIO PACKET LOSS COMPENSATION METHOD AND APPARATUS AND ELECTRONIC DEVICE
3y 0m to grant Granted Jul 14, 2026
Patent 12682912
Pseudo Real-time Content-Aware Auditory Cleansing
2y 9m to grant Granted Jul 14, 2026
Patent 12676159
HARMONIC TRANSPOSITION IN AN AUDIO CODING METHOD AND SYSTEM
1y 9m to grant Granted Jul 07, 2026
Patent 12670895
INVERTIBLE NEURAL NETWORK TO SYNTHESIZE AUDIO SIGNALS
7y 0m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
67%
Grant Probability
99%
With Interview (+33.3%)
2y 8m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 389 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month