Prosecution Insights
Last updated: October 02, 2026
Application No. 18/599,018

SPEECH MODIFICATION USING ACCENT EMBEDDINGS

Final Rejection §103
Filed
Mar 07, 2024
Priority
Mar 09, 2023 — provisional 63/451,040 +2 more
Examiner
BECKER, TYLER JUSTIN
Art Unit
2657
Tech Center
2600 — Communications
Assignee
Sri International
OA Round
2 (Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
1m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
20 granted / 27 resolved
+12.1% vs TC avg
Moderate +13% lift
Without
With
+13.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
12 currently pending
Career history
46
Total Applications
across all art units

Statute-Specific Performance

§101
17.7%
-22.3% vs TC avg
§103
55.2%
+15.2% vs TC avg
§102
12.8%
-27.2% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 27 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The amendments filed March 30th, 2026 have been entered. Claims 1, 7, 12, and 17 have been amended. Claims 1-20 are pending and have been examined. Response to Arguments Applicant’s arguments with respect to claim(s) 1-11 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant’s amendments and arguments, see page 10 of the applicant's remarks, filed March 30th, 2026, with respect to the rejection of claims 12-20 under 35 U.S.C. 102 have been fully considered and are persuasive. The rejection of January 2nd, 2026 has been withdrawn. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-4 and 7-9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Garman et al. (US Pat. Pub. No. 2021/0256961 A1 hereinafter Garman) in view of Golman et al. (US Pat. Pub. No. 20230223011 A1 hereinafter Golman). Regarding claim 1, Garman discloses a method comprising: obtaining a dataset of a plurality of sample speech clips (Garman, Fig. 3, 302; [0045]: "For synthesis, the inputs may include: the text to be spoken 302, the language of the text 304, identification of the speaker 306, and the output accent 308 (which may be the same as the language of the text)."); generating a plurality of sequence embeddings based on the plurality of sample speech clips (Garman, Fig. 3, 320; [0033]: "With embedding, each input item may be replaced by a vector of values. Accordingly, phonemes 214 may be embedded in feature vector 220, while prosody 218 may be embedded in feature vector 222."); initializing a plurality of speaker embeddings; initializing a plurality of accent embeddings; updating the plurality of speaker embeddings based on the plurality of sample speech clips; updating the plurality of accent embeddings based on the plurality of sample speech clips (Garman, Fig. 3, 324 and 326; [0052]: "The inputs for the phonemes 316, accents 308, and speakers 306 may be fed into an Embedding layer to generated embeddings 320, 324, 326."); generating a plurality of augmented embeddings based on the plurality of sequence embeddings, the plurality of speaker embeddings, and the plurality of accent embeddings (Garman, Fig. 3, 228 and Output; [0040]: "DRNN 228 may take a sequence of embedded phonemes 220, prosodic values 222, language identifiers 224, and speaker identifiers 226 as input, as described above."; [0041]: "The output of DRNN 228 may be a set of acoustic features 234, as described above. DRNN 228 may be language independent. When trained, DRNN 228 may be a single universal speech model and may encode all the information necessary to produce speech for all of the trained languages and speakers."); and generating a plurality of synthetic speech clips based on the plurality of augmented embeddings (Garman, Fig. 3, 314; [0046]: "Many of the components of system 300 are similar to those shown in FIG. 2. In embodiments, differences may include text conversion block 310 instead of ASR & aligner 212, the output of each frame may be used as input to a later frame after some time delay, and the output goes to a decoder 312 to produce speech."). However, Garman fails to expressly recite obtaining a dataset of a plurality of sample speech clips that include audio data; generating a plurality of sequence embeddings based on the plurality of sample speech clips that include the audio data; updating the plurality of speaker embeddings based on the plurality of sample speech clips that include the audio data; and updating the plurality of accent embeddings based on the plurality of sample speech clips that include the audio data. Golman teaches obtaining a dataset of a plurality of sample speech clips that include audio data (Golman, [0043]: “FIG. 2 is a schematic showing features 216 that can be extracted from a speech audio signal 202, according to some example embodiments of the present disclosure. The speech audio signal 202 may include waveforms 214.”; [0063]: “During inference, an output (target) accent can be selected from accents available on training stage. During the training stage, datasets of different voices and accents can be used. Any of the datasets can be validated for appropriate sound quality and then used for output target voice and accent.”); generating a plurality of sequence embeddings based on the plurality of sample speech clips that include the audio data (Golman, [0048]: “The linguistic features extraction module 304 may extract, from a time frame of the speech audio signal 202, linguistic features 210. In some embodiments, linguistic features 210 may include hidden features of an Automatic Speech Recognition (ASR) neural network with additional custom training and transformations or phonemes belonging to a phoneme set for a predetermined language.”); updating the plurality of speaker embeddings based on the plurality of sample speech clips that include the audio data; and updating the plurality of accent embeddings based on the plurality of sample speech clips that include the audio data (Golman, [0043]: “The features 216 may include acoustic features and linguistic features 210. The acoustic features may include pitch 206 (or main frequency (F0)), energy 208 (signal amplitude), and voice activity detection (VAD) 212.”). Garman and Golman are analogous arts because they each belong to the same field of speech processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the parametric speech synthesis system of Garman, to incorporate the teachings of Golman to extract embeddings from input audio data. This allows the system to convert accents in real-time speech scenarios (Golman, [0003]), thus improving the user’s experience when conversing with speakers with strong accents. Regarding claim 2, the rejection of claim 1 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. Garman, in view of Golman, further discloses wherein a sample speech clip of the plurality of sample speech clips is labeled with a text transcript and an accent identifier, and wherein updating the plurality of accent embeddings comprises updating an accent embedding of the plurality of accent embeddings based on the text transcript and the accent identifier (Garman, Fig. 3, 228 and Output; [0040]: "DRNN 228 may take a sequence of embedded phonemes 220, prosodic values 222, language identifiers 224, and speaker identifiers 226 as input, as described above."; Garman, [0058]: "For example, as shown in FIG. 2, phonemes may be embedded 220. In embodiments, the embedding may be performed so that similar phonemes may generate similar embedding vectors. Likewise, the differences between the phonemes may also be reflected in the embedding vectors. Embedding bag layers are similar to embedding layers, but may accept zero or more inputs and may generate one or more outputs. Embedding bag functions may compute, for example, sums or means of several embedding vectors. In embodiments, embedding bag functions may be used with prosodic elements 218, while embedding functions may be used with phonemes 214, languages 208, and speaker elements 210."). Regarding claim 3, the rejection of claim 1 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. Garman, in view of Golman, further discloses wherein a sample speech clip of the plurality of sample speech clips is labeled with a text transcript and a speaker identifier, and wherein updating the plurality of speaker embeddings comprises updating a speaker embedding of the plurality of speaker embeddings based on the text transcript and the speaker identifier (Garman, Fig. 3, 228 and Output; [0040]: "DRNN 228 may take a sequence of embedded phonemes 220, prosodic values 222, language identifiers 224, and speaker identifiers 226 as input, as described above."; Garman, [0058]: "For example, as shown in FIG. 2, phonemes may be embedded 220. In embodiments, the embedding may be performed so that similar phonemes may generate similar embedding vectors. Likewise, the differences between the phonemes may also be reflected in the embedding vectors. Embedding bag layers are similar to embedding layers, but may accept zero or more inputs and may generate one or more outputs. Embedding bag functions may compute, for example, sums or means of several embedding vectors. In embodiments, embedding bag functions may be used with prosodic elements 218, while embedding functions may be used with phonemes 214, languages 208, and speaker elements 210."). Regarding claim 4, the rejection of claim 1 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. Garman, in view of Golman, further discloses wherein generating the plurality of augmented embeddings comprises summing the plurality of sequence embeddings with the plurality of speaker embeddings and the plurality of accent embeddings (Garman, Fig. 3, 228 and Output; [0040]: "DRNN 228 may take a sequence of embedded phonemes 220, prosodic values 222, language identifiers 224, and speaker identifiers 226 as input, as described above."; [0041]: "The output of DRNN 228 may be a set of acoustic features 234, as described above. DRNN 228 may be language independent. When trained, DRNN 228 may be a single universal speech model and may encode all the information necessary to produce speech for all of the trained languages and speakers."). Regarding claim 7, Garman discloses a computing system comprising processing circuitry and memory for executing a machine learning system (Garman, [0010]: “a method for text-to-speech conversion may be implemented in a computer system comprising a processor, memory accessible by the processor, and computer program instructions stored in the memory and executable by the processor”), the machine learning system configured to: obtain a dataset of a plurality of sample speech clips (Garman, Fig. 3, 302; [0045]: "For synthesis, the inputs may include: the text to be spoken 302, the language of the text 304, identification of the speaker 306, and the output accent 308 (which may be the same as the language of the text)."); generate a plurality of sequence embeddings based on the plurality of sample speech clips (Garman, Fig. 3, 320; [0033]: "With embedding, each input item may be replaced by a vector of values. Accordingly, phonemes 214 may be embedded in feature vector 220, while prosody 218 may be embedded in feature vector 222."); initialize a plurality of speaker embeddings and a plurality of accent embeddings; update the plurality of speaker embeddings based on the plurality of sample speech clips; update the plurality of accent embeddings based on the plurality of sample speech clips (Garman, Fig. 3, 324 and 326; [0052]: "The inputs for the phonemes 316, accents 308, and speakers 306 may be fed into an Embedding layer to generated embeddings 320, 324, 326."); generate a plurality of augmented embeddings based on the plurality of sequence embeddings, the plurality of speaker embeddings, and the plurality of accent embeddings (Garman, Fig. 3, 228 and Output; [0040]: "DRNN 228 may take a sequence of embedded phonemes 220, prosodic values 222, language identifiers 224, and speaker identifiers 226 as input, as described above."; [0041]: "The output of DRNN 228 may be a set of acoustic features 234, as described above. DRNN 228 may be language independent. When trained, DRNN 228 may be a single universal speech model and may encode all the information necessary to produce speech for all of the trained languages and speakers."); and generate a plurality of synthetic speech clips based on the plurality of augmented embeddings (Garman, Fig. 3, 314; [0046]: "Many of the components of system 300 are similar to those shown in FIG. 2. In embodiments, differences may include text conversion block 310 instead of ASR & aligner 212, the output of each frame may be used as input to a later frame after some time delay, and the output goes to a decoder 312 to produce speech."). However, Garman fails to expressly recite obtain a dataset of a plurality of sample speech clips that include audio data; generate a plurality of sequence embeddings based on the plurality of sample speech clips that include the audio data; update the plurality of speaker embeddings based on the plurality of sample speech clips that include the audio data; and update the plurality of accent embeddings based on the plurality of sample speech clips that include the audio data. Golman teaches obtain a dataset of a plurality of sample speech clips that include audio data (Golman, [0043]: “FIG. 2 is a schematic showing features 216 that can be extracted from a speech audio signal 202, according to some example embodiments of the present disclosure. The speech audio signal 202 may include waveforms 214.”; [0063]: “During inference, an output (target) accent can be selected from accents available on training stage. During the training stage, datasets of different voices and accents can be used. Any of the datasets can be validated for appropriate sound quality and then used for output target voice and accent.”); generate a plurality of sequence embeddings based on the plurality of sample speech clips that include the audio data (Golman, [0048]: “The linguistic features extraction module 304 may extract, from a time frame of the speech audio signal 202, linguistic features 210. In some embodiments, linguistic features 210 may include hidden features of an Automatic Speech Recognition (ASR) neural network with additional custom training and transformations or phonemes belonging to a phoneme set for a predetermined language.”); update the plurality of speaker embeddings based on the plurality of sample speech clips that include the audio data; and update the plurality of accent embeddings based on the plurality of sample speech clips that include the audio data (Golman, [0043]: “The features 216 may include acoustic features and linguistic features 210. The acoustic features may include pitch 206 (or main frequency (F0)), energy 208 (signal amplitude), and voice activity detection (VAD) 212.”). Garman and Golman are analogous arts because they each belong to the same field of speech processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the parametric speech synthesis system of Garman, to incorporate the teachings of Golman to extract embeddings from input audio data. This allows the system to convert accents in real-time speech scenarios (Golman, [0003]), thus improving the user’s experience when conversing with speakers with strong accents. Regarding claim 8, the rejection of claim 7 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. Garman, in view of Golman, further discloses wherein a sample speech clip of the plurality of sample speech clips is labeled with a text transcript and an accent identifier, and wherein to update the plurality of accent embeddings, the machine learning system is configured to: update an accent embedding of the plurality of accent embeddings based on the text transcript and the accent identifier (Garman, Fig. 3, 228 and Output; [0040]: "DRNN 228 may take a sequence of embedded phonemes 220, prosodic values 222, language identifiers 224, and speaker identifiers 226 as input, as described above."; Garman, [0058]: "For example, as shown in FIG. 2, phonemes may be embedded 220. In embodiments, the embedding may be performed so that similar phonemes may generate similar embedding vectors. Likewise, the differences between the phonemes may also be reflected in the embedding vectors. Embedding bag layers are similar to embedding layers, but may accept zero or more inputs and may generate one or more outputs. Embedding bag functions may compute, for example, sums or means of several embedding vectors. In embodiments, embedding bag functions may be used with prosodic elements 218, while embedding functions may be used with phonemes 214, languages 208, and speaker elements 210."). Regarding claim 9, the rejection of claim 7 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. Garman, in view of Golman, further discloses wherein to generate the plurality of augmented embeddings, the machine learning system is configured to: sum the plurality of sequence embeddings with the plurality of speaker embeddings and the plurality of accent embeddings (Garman, Fig. 3, 228 and Output; [0040]: "DRNN 228 may take a sequence of embedded phonemes 220, prosodic values 222, language identifiers 224, and speaker identifiers 226 as input, as described above."; [0041]: "The output of DRNN 228 may be a set of acoustic features 234, as described above. DRNN 228 may be language independent. When trained, DRNN 228 may be a single universal speech model and may encode all the information necessary to produce speech for all of the trained languages and speakers."). Claim(s) 5 and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Garman, in view of Golman, as applied to claims 1-4 and 7-9 above, and further in view of Finkelstein et al. (US Pat. Pub. No. 2023/0018384 A1 hereinafter Finkelstein). Regarding claim 5, the rejection of claim 1 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. However, Garman, in view of Golman, fails to expressly recite wherein the plurality of synthetic speech clips comprises a first synthetic speech clip associated with a speaker and a first accent and a second synthetic speech clip associated with the speaker and a second accent, and wherein the method further comprises: aligning a first set of frames associated with the first synthetic speech clip with a second set of frames associated with the second synthetic speech clip; and generating an instance of training data based on the alignment of the first set of frames associated with the first synthetic speech clip and the second set of frames associated with the second synthetic speech clip. Finkelstein teaches wherein the plurality of synthetic speech clips comprises a first synthetic speech clip associated with a speaker and a first accent and a second synthetic speech clip associated with the speaker and a second accent (Finkelstein, [0025]: "Accordingly, a given text input can produce synthesized speech in a given language across various different accents/dialects and/or different speaking styles, as well as produce synthesized speech across different languages."), and wherein the method further comprises: aligning a first set of frames associated with the first synthetic speech clip with a second set of frames associated with the second synthetic speech clip (Finkelstein, [0031]: "examples herein are directed toward the trained voice cloning system 200 generating training synthesized speech representations 202 that clone the voice of a target speaker in a target accent/dialect (e.g., second accent/dialect)."; [0005]: "The operations may further include sampling, from the training synthesized speech representation, a sequence of fixed-length reference frames providing reference prosodic features that represent the prosody captured by the training synthesized speech representation."); and generating an instance of training data based on the alignment of the first set of frames associated with the first synthetic speech clip and the second set of frames associated with the second synthetic speech clip (Finkelstein, [0031]: "examples herein are directed toward the trained voice cloning system 200 generating training synthesized speech representations 202 that clone the voice of a target speaker in a target accent/dialect (e.g., second accent/dialect)."). Garman, Golman, and Finkelstein are analogous arts because they each belong to the same field of speech processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the parametric speech synthesis system of Garman, as modified by the real time accent correction method of Golman, to incorporate the teachings of Finkelstein to generate a training data set based on synthesized speech. Generating a synthetic training data set allows the system to create training data that would be challenging to collect naturally (Finkelstein, [0025]). This ensures a large and varied training set is available for training speech processing systems. Regarding claim 10, the rejection of claim 7 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. However, Garman, in view of Golman, fails to expressly recite wherein the plurality of synthetic speech clip comprises a first synthetic speech clip associated with a speaker and a first accent and a second synthetic speech clip associated with the speaker and a second accent, and wherein the machine learning system is further configured to: align a first set of frames associated with the first synthetic speech clip with a second set of frames associated with the second synthetic speech clip; and generate an instance of training data based on the alignment of the first set of frames associated with the first synthetic speech clip and the second set of frames associated with the second synthetic speech clip. Finkelstein teaches wherein the plurality of synthetic speech clip comprises a first synthetic speech clip associated with a speaker and a first accent and a second synthetic speech clip associated with the speaker and a second accent (Finkelstein, [0025]: "Accordingly, a given text input can produce synthesized speech in a given language across various different accents/dialects and/or different speaking styles, as well as produce synthesized speech across different languages."), and wherein the machine learning system is further configured to: align a first set of frames associated with the first synthetic speech clip with a second set of frames associated with the second synthetic speech clip (Finkelstein, [0031]: "examples herein are directed toward the trained voice cloning system 200 generating training synthesized speech representations 202 that clone the voice of a target speaker in a target accent/dialect (e.g., second accent/dialect)."; [0005]: "The operations may further include sampling, from the training synthesized speech representation, a sequence of fixed-length reference frames providing reference prosodic features that represent the prosody captured by the training synthesized speech representation."); and generate an instance of training data based on the alignment of the first set of frames associated with the first synthetic speech clip and the second set of frames associated with the second synthetic speech clip (Finkelstein, [0031]: "examples herein are directed toward the trained voice cloning system 200 generating training synthesized speech representations 202 that clone the voice of a target speaker in a target accent/dialect (e.g., second accent/dialect)."). Garman, Golman, and Finkelstein are analogous arts because they each belong to the same field of speech processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the parametric speech synthesis system of Garman, as modified by the real time accent correction method of Golman, to incorporate the teachings of Finkelstein to generate a training data set based on synthesized speech. Generating a synthetic training data set allows the system to create training data that would be challenging to collect naturally (Finkelstein, [0025]). This ensures a large and varied training set is available for training speech processing systems. Claim(s) 6 and 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Garman, in view of Golman, as applied to claims 1-4 and 7-9 above, and further in view of Shakeri. Regarding claim 6, the rejection of claim 1 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. However, Garman, in view of Golman, fails to expressly recite providing the plurality of synthetic speech clips to an autoencoder model; and training the autoencoder model to modify speech of an audio waveform based on the plurality of synthetic speech clips. Shakeri teaches providing the plurality of synthetic speech clips to an autoencoder model; and training the autoencoder model to modify speech of an audio waveform based on the plurality of synthetic speech clips (Shakeri, Col. 15, lines 26-28: "As shown in FIG. 6, acoustic feature encoder 605 and acoustic feature decoder 606 are trained on a voice conversion task using one or more training examples 601."). Garman, Golman, and Shakeri are analogous arts because they each belong to the same field of speech processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the parametric speech synthesis system of Garman, as modified by the real time accent correction method of Golman, to incorporate the teachings of Shakeri to use the training data set to train a voice conversion autoencoder model. Training the autoencoder allows the system to properly convert a voice into another target voice (Shakeri, Col. 15, lines 28-34). This ensures the system can successfully convert voices for the user. Regarding claim 11, the rejection of claim 7 is incorporated. Garman, in view of Golman, discloses all of the elements of the current invention as stated above. However, Garman, in view of Golman, fails to expressly recite wherein the machine learning system is further configured to: provide the plurality of synthetic speech clips to an autoencoder model; and train the autoencoder model to modify speech of an audio waveform based on the plurality of synthetic speech clips. Shakeri teaches wherein the machine learning system is further configured to: provide the plurality of synthetic speech clips to an autoencoder model; and train the autoencoder model to modify speech of an audio waveform based on the plurality of synthetic speech clips (Shakeri, Col. 15, lines 26-28: "As shown in FIG. 6, acoustic feature encoder 605 and acoustic feature decoder 606 are trained on a voice conversion task using one or more training examples 601."). Garman, Golman, and Shakeri are analogous arts because they each belong to the same field of speech processing. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the parametric speech synthesis system of Garman, as modified by the real time accent correction method of Golman, to incorporate the teachings of Shakeri to use the training data set to train a voice conversion autoencoder model. Training the autoencoder allows the system to properly convert a voice into another target voice (Shakeri, Col. 15, lines 28-34). This ensures the system can successfully convert voices for the user. Allowable Subject Matter Claims 12-20 allowed. The following is a statement of reasons for the indication of allowable subject matter: Regarding claim 12, the prior art search did not yield any references that disclosed the limitation “decomposing the audio waveform into first one or more magnitude spectral slices associated with a source accent and an original phase that corresponds to the first one or more magnitude spectral slices;” before the effective filing date of the claimed invention, nor any references that would have been obvious to one of ordinary skill in the art to have combined to yield the abovementioned limitation before the effective filing date of the claimed invention. As such, especially in combination with the other limitations of claim 12, the claim is allowable. Regarding claim 17, the claim recites limitations comparable to claim 12, and thus are allowable for the same reasons as stated above with regards to claim 12. Regarding claims 13-16 and 18-20, each of these claims depends on one of the independent claims above, or on another dependent claim, and therefore incorporates all limitations therefrom. As such, claims 13-16 and 18-20 are allowable for at least the same reasons as stated above with regards to claims 1 and 17. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Zhao et al. (US Pat. Pub. No. 2023/0335107 A1) discloses a system for reference-free foreign accent conversion. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER J BECKER whose telephone number is (703)756-1271. The examiner can normally be reached M-Th, 7:15am-5:45pm PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TYLER BECKER/ Examiner, Art Unit 2657 /DANIEL C WASHBURN/ Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Show 2 earlier events
Mar 10, 2026
Interview Requested
Mar 25, 2026
Examiner Interview Summary
Mar 25, 2026
Applicant Interview (Telephonic)
Mar 30, 2026
Response Filed
Jul 28, 2026
Final Rejection mailed — §103
Sep 03, 2026
Interview Requested
Sep 15, 2026
Examiner Interview Summary
Sep 15, 2026
Applicant Interview (Telephonic)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731597
METHOD AND SYSTEM FOR CONSTRUCTING LEARNING DATABASE USING VOICE PERSONAL INFORMATION PROTECTION TECHNOLOGY
2y 8m to grant Granted Sep 08, 2026
Patent 12718808
SUBTITLE GENERATION METHOD, APPARATUS, ELECTRONIC DEVICE, STORAGE MEDIUM AND PROGRAM
2y 8m to grant Granted Aug 25, 2026
Patent 12694228
REAL-TIME USER COMMUNICATION SENTIMENT DETECTION FOR DYNAMIC ANOMALY DETECTION AND MITIGATION
3y 4m to grant Granted Jul 28, 2026
Patent 12682113
SYSTEMS, METHODS, AND APPARATUSES FOR GENERATING STRUCTURED DATA FROM UNSTRUCTURED DATA USING NATURAL LANGUAGE PROCESSING TO GENERATE A SECURE MEDICAL DASHBOARD
3y 0m to grant Granted Jul 14, 2026
Patent 12651592
SYSTEM, METHOD, AND COMPUTER PROGRAM FOR REAL-TIME LANGUAGE TRANSLATION USING GENERATIVE ARTIFICIAL INTELLIGENCE
3y 0m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
87%
With Interview (+13.1%)
2y 8m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 27 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month