Prosecution Insights
Last updated: September 05, 2026
Application No. 18/951,397

END-TO-END TEXT-TO-SPEECH CONVERSION

Non-Final OA §103§DOUBLEPATENT
Filed
Nov 18, 2024
Priority
Mar 29, 2017 — GR 20170100126 +7 more
Examiner
SHIN, SEONG-AH A
Art Unit
Tech Center
Assignee
Google LLC
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
330 granted / 421 resolved
+18.4% vs TC avg
Strong +22% interview lift
Without
With
+21.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
20 currently pending
Career history
444
Total Applications
across all art units

Statute-Specific Performance

§101
23.3%
-16.7% vs TC avg
§103
46.7%
+6.7% vs TC avg
§102
14.0%
-26.0% vs TC avg
§112
6.8%
-33.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 421 resolved cases

Office Action

§103 §DOUBLEPATENT
CTNF 18/951,397 CTNF 90194 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. 12-151 AIA 26-51 12-51 Status of Claims Claims 2-21 pending in this application. Claim 1 is canceled. Double Patenting 08-33 AIA The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the claims at issue are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg , 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman , 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi , 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum , 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel , 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington , 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the reference application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO internet Web site contains terminal disclaimer forms which may be used. Please visit http://www.uspto.gov/forms/. The filing date of the application will determine what form should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to http://www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claims 2-21 are rejected on the ground of nonstatutory double patenting over claims 1-20 of U.S. Patent No. 11,107,457. Although the claims at issue are not identical, they are not patentably distinct from each other because adding inherent and/or unnecessary limitations/step and rearranging the claims would be within the level of one of ordinary skill in the art. It is well settled that the insertion of an element, e.g. “receiving a respective embedding of the character in the sequence, processing the respective embedding of the character in the sequence to generate a respective transformed embedding of the character, and processing the respective transformed embedding of the character in the sequence to generate a respective encoded representation of the character”, and its function is an obvious expedient if the remaining elements perform the same function as before. In re Karlson, 136 USPQ 184 (CCPA 1963) . Also note Ex parte Rainu , 168 USPQ 375 (Bd. App. 1969). Insertion of a reference element or step whose function is not needed would be obvious to one of ordinary skill in the art. Instant Application No. 18/951.397 U.S. Patent No. 11,107,457 2.A computer-implemented method for generating, from a representation of a text input, a spectrogram of a verbal utterance of the text input using a text-to-speech conversion system, the method comprising: processing, using an encoder neural network of the text-to-speech conversion system , the representation of the text input to generate a respective encoded representation of each of a plurality of components in the representation of the text input; generating, using a decoder neural network of the text-to-speech conversion system, multiple frames of a spectrogram based on the encoded representations; processing, using a post-processing neural network of the text-to-speech conversion system, the spectrogram to generate a waveform synthesizer input; and generating a waveform of the verbal utterance of the text input based on the waveform synthesizer input. 3. The method of claim 2, wherein the encoder neural network comprises an encoder pre-net neural network and an encoder CBHG neural network, and wherein processing, using the encoder neural network of the text-to-speech conversion system, the representation of the text input to generate a respective encoded representation of each of the plurality of components in the representation of the text input comprises: receiving, using the encoder pre-net neural network, a respective embedding of each component in the plurality of the components, processing, using the encoder pre-net neural network, the respective embedding of each component in the plurality of components to generate a respective transformed embedding of the component, and processing, using the encoder CBHG neural network, a respective transformed embedding of each component in the plurality of components to generate a respective encoded representation of the component. 4. The method of claim 3, wherein the encoder CBHG neural network comprises a bank of 1-D convolutional filters, followed by a highway network, and followed by a bidirectional recurrent neural network. 5. The method of claim 4, wherein the bidirectional recurrent neural network is a gated recurrent unit neural network. 6. The method of claim 4, wherein the encoder CBHG neural network includes a residual connection between the transformed embeddings and outputs of the bank of 1-D convolutional filters. 7. The method of claim 4, wherein the bank of 1-D convolutional filters includes a max pooling along time layer with stride one. 8. The method of claim 2, further comprising receiving a sequence of decoder inputs, wherein generating, using the decoder neural network of the text-to-speech conversion system, multiple frames of the spectrogram based on the encoded representations comprises :for each decoder input in the sequence of decoder inputs, processing, using the decoder neural network of the text-to-speech conversion system, the decoder input and the encoded representations to generate multiple frames of the spectrogram, and wherein a first decoder input in the sequence of decoder inputs is a predetermined initial frame. 9. The method of claim 2, wherein the spectrogram is a mel-scale spectrogram. 10. The method of claim 2 , wherein generating the waveform of the verbal utterance of the text input based on the waveform synthesizer input comprises: processing, using a waveform synthesizer of the text-to-speech conversion system, the waveform synthesizer input to generate the waveform of the verbal utterance of the text input. 11. The method of claim 10, further comprising: generating speech using the waveform; and providing the generated speech for playback. 12. The method of claim 10, wherein the waveform synthesizer is a trainable spectrogram to waveform inverter. 13. The method of claim 10, wherein the waveform synthesizer is a vocoder. 14. The method of claim 9, wherein the waveform synthesizer input is a linear- scale spectrogram of the verbal utterance of the text input. 1. A computer-implemented method for generating, from a sequence of characters in a particular natural language, a spectrogram of a verbal utterance of the sequence of characters in the particular natural language using a text-to-speech conversion system, the method comprising: processing, using an encoder neural network of the text-to-speech conversion system, the sequence of characters to generate a respective encoded representation of each of the characters in the sequence , comprising, for each character in the sequence of characters: receiving a respective embedding of the character in the sequence, processing the respective embedding of the character in the sequence to generate a respective transformed embedding of the character, and processing the respective transformed embedding of the character in the sequence to generate a respective encoded representation of the character; receiving a sequence of decoder inputs; for each decoder input in the sequence of decoder inputs, processing, using a decoder neural network of the text-to-speech conversion system, the decoder input and the encoded representations to generate multiple frames of the spectrogram ; and generating a waveform from the spectrogram of the verbal utterance of the sequence of characters in the particular natural language . 2. The method of claim 1, wherein the encoder neural network comprises an encoder pre-net neural network and an encoder CBHG neural network, and wherein for each character in the sequence of characters, receiving the respective embedding of the character in the sequence comprises receiving, using the encoder pre-net neural network, the respective embedding of the character in the sequence, processing the respective embedding of the character in the sequence to generate the respective transformed embedding of the character comprises processing, using the encoder pre-net neural network, the respective embedding of the character in the sequence to generate the respective transformed embedding of the character, and processing the respective transformed embedding of the character in the sequence to generate the respective encoded representation of the character comprises processing, using the encoder CBHG neural network, the respective transformed embedding of the character in the sequence to generate the respective encoded representation of the character. The method of claim 2, wherein the encoder CBHG neural network comprises a bank of 1-D convolutional filters, followed by a highway network, and followed by a bidirectional recurrent neural network. 4. The method of claim 3, wherein the bidirectional recurrent neural network is a gated recurrent unit neural network. 5. The method of claim 3, wherein the encoder CBHG includes a residual connection between the transformed embeddings and outputs of the bank of 1-D convolutional filters. 6. The method of claim 3, wherein the bank of 1-D convolutional filters includes a max pooling along time layer with stride one. 7. The method of claim 1, wherein a first decoder input in the sequence is a predetermined initial frame. 8. The method of claim 1, wherein the spectrogram is a compressed spectrogram. 9. The method of claim 8, wherein the compressed spectrogram is a mel-scale spectrogram. 10. The method of claim 8, further comprising: processing the compressed spectrogram to generate a waveform synthesizer input; and processing, using a waveform synthesizer of the text-to-speech conversion system, the waveform synthesizer input to generate the waveform of the verbal utterance of the input sequence of characters in the particular natural language. 11. The method of claim 1, further comprising: generating speech using the waveform; and providing the generated speech for playback. 12. The method of claim 10, wherein the waveform synthesizer is a trainable spectrogram to waveform inverter. 13. The method of claim 10, wherein the waveform synthesizer is a vocoder. 14. The method of claim 10, wherein the waveform synthesizer input is a linear-scale spectrogram of the verbal utterance of the input sequence of characters in the particular natural language. Claims 2-21 are rejected on the ground of nonstatutory double patenting over claims 1-20 of U.S. Patent No. 12,190,860. Although the claims at issue are not identical, they are not patentably distinct from each other because adding inherent and/or unnecessary limitations/step and rearranging the claims would be within the level of one of ordinary skill in the art. It is well settled that the insertion of an element, e.g. “wherein the encoder neural network and the decoder neural network have been trained on training data that maps representations of text inputs to mel-scale spectrograms”, and its function is an obvious expedient if the remaining elements perform the same function as before. In re Karlson, 136 USPQ 184 (CCPA 1963) . Also note Ex parte Rainu , 168 USPQ 375 (Bd. App. 1969). Insertion of a reference element or step whose function is not needed would be obvious to one of ordinary skill in the art. Instant Application No. 18/951.397 U.S. Patent No. 12,190,860 2.A computer-implemented method for generating, from a representation of a text input, a spectrogram of a verbal utterance of the text input using a text-to-speech conversion system, the method comprising: processing, using an encoder neural network of the text-to-speech conversion system , the representation of the text input to generate a respective encoded representation of each of a plurality of components in the representation of the text input; generating, using a decoder neural network of the text-to-speech conversion system, multiple frames of a spectrogram based on the encoded representations; processing, using a post-processing neural network of the text-to-speech conversion system, the spectrogram to generate a waveform synthesizer input; and generating a waveform of the verbal utterance of the text input based on the waveform synthesizer input. 3. The method of claim 2, wherein the encoder neural network comprises an encoder pre-net neural network and an encoder CBHG neural network, and wherein processing, using the encoder neural network of the text-to-speech conversion system, the representation of the text input to generate a respective encoded representation of each of the plurality of components in the representation of the text input comprises: receiving, using the encoder pre-net neural network, a respective embedding of each component in the plurality of the components, processing, using the encoder pre-net neural network, the respective embedding of each component in the plurality of components to generate a respective transformed embedding of the component, and processing, using the encoder CBHG neural network, a respective transformed embedding of each component in the plurality of components to generate a respective encoded representation of the component. 4. The method of claim 3, wherein the encoder CBHG neural network comprises a bank of 1-D convolutional filters, followed by a highway network, and followed by a bidirectional recurrent neural network. 5. The method of claim 4, wherein the bidirectional recurrent neural network is a gated recurrent unit neural network. 6. The method of claim 4, wherein the encoder CBHG neural network includes a residual connection between the transformed embeddings and outputs of the bank of 1-D convolutional filters. 7. The method of claim 4, wherein the bank of 1-D convolutional filters includes a max pooling along time layer with stride one. 8. The method of claim 2, further comprising receiving a sequence of decoder inputs, wherein generating, using the decoder neural network of the text-to-speech conversion system, multiple frames of the spectrogram based on the encoded representations comprises :for each decoder input in the sequence of decoder inputs, processing, using the decoder neural network of the text-to-speech conversion system, the decoder input and the encoded representations to generate multiple frames of the spectrogram, and wherein a first decoder input in the sequence of decoder inputs is a predetermined initial frame. 9. The method of claim 2, wherein the spectrogram is a mel-scale spectrogram. 10. The method of claim 2 , wherein generating the waveform of the verbal utterance of the text input based on the waveform synthesizer input comprises: processing, using a waveform synthesizer of the text-to-speech conversion system, the waveform synthesizer input to generate the waveform of the verbal utterance of the text input. 11. The method of claim 10, further comprising: generating speech using the waveform; and providing the generated speech for playback. 12. The method of claim 10, wherein the waveform synthesizer is a trainable spectrogram to waveform inverter. 13. The method of claim 10, wherein the waveform synthesizer is a vocoder. 14. The method of claim 9, wherein the waveform synthesizer input is a linear- scale spectrogram of the verbal utterance of the text input. 1. A computer-implemented method for generating, from a representation of a text input, a mel-scale spectrogram of a verbal utterance of the text input using a text-to-speech conversion system, the method comprising: processing, using an encoder neural network of the text-to-speech conversion system, the representation of the text input to generate a respective encoded representation of each of a plurality of components in the representation of the text input; receiving a sequence of decoder inputs; and for each decoder input in the sequence of decoder inputs, processing, using a decoder neural network of the text-to-speech conversion system, the decoder input and the encoded representations to generate multiple frames of the mel-scale spectrogram, wherein the encoder neural network and the decoder neural network have been trained on training data that maps representations of text inputs to mel-scale spectrograms. 2. The method of claim 1, wherein the encoder neural network comprises an encoder pre-net neural network and an encoder CBHG neural network, and wherein processing, using the encoder neural network of the text-to-speech conversion system, the representation of the text input to generate a respective encoded representation of each of the plurality of components in the representation of the text input comprises: receiving, using the encoder pre-net neural network, a respective embedding of each component in the plurality of the components, processing, using the encoder pre-net neural network, the respective embedding of each component in the plurality of components to generate a respective transformed embedding of the component, and processing, using the encoder CBHG neural network, a respective transformed embedding of each component in the plurality of components to generate a respective encoded representation of the component. 3. The method of claim 2, wherein the encoder CBHG neural network comprises a bank of 1-D convolutional filters, followed by a highway network, and followed by a bidirectional recurrent neural network. 4. The method of claim 3, wherein the bidirectional recurrent neural network is a gated recurrent unit neural network. 5. The method of claim 3, wherein the encoder CBHG neural network includes a residual connection between the transformed embeddings and outputs of the bank of 1-D convolutional filters. 6. The method of claim 3, wherein the bank of 1-D convolutional filters includes a max pooling along time layer with stride one. 7. The method of claim 1, wherein a first decoder input in the sequence of decoder inputs is a predetermined initial frame. 8. The method of claim 1, further comprising: processing the mel-scale spectrogram to generate a waveform synthesizer input; and processing, using a waveform synthesizer of the text-to-speech conversion system, the waveform synthesizer input to generate a waveform of the verbal utterance of the text input. 9. The method of claim 1, further comprising generating a waveform from the mel-scale spectrogram of the verbal utterance of the text input. 10. The method of claim 9, further comprising: generating speech using the waveform; and providing the generated speech for playback . 11. The method of claim 8, wherein the waveform synthesizer is a trainable spectrogram to waveform inverter. 12. The method of claim 8, wherein the waveform synthesizer is a vocoder. 13. The method of claim 8, wherein the waveform synthesizer input is a linear-scale spectrogram of the verbal utterance of the text input. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-103 AIA The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. 07-23-aia AIA The factual inquiries set forth in Graham v. John Deere Co. , 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 07-20-02-aia AIA This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. 07-21-fti Claim s 2, 8-15, 19-21 are rejected under pre-AIA 35 U.S.C. 103(a) as being unpatentable over Sotelo et al., (“CHAR2WAV: END-TO-END SPEECH SYNTHESIS”, hereinafter Sotelo, Mar-10-2017) in view of Pollet et al., (US Pub. 2009/0048841, hereinafter Pollet) . Regarding claim 2, Sotelo discloses a computer-implemented method for generating, from a representation of a text input, a spectrogram of a verbal utterance of the text input using a text-to-speech conversion system, the method comprising: processing, using an encoder neural network of the text-to-speech conversion system, the representation of the text input to generate a respective encoded representation of each of a plurality of components in the representation of the text input ([Abstract] a speech synthesis, Char2Wav, has two components: “a reader and a neural vocoder. The reader is an encoder-decoder model with attention. The encoder is a bidirectional recurrent neural network that accepts text or phonemes as inputs, while the decoder is a recurrent neural network (RNN) with attention that produces vocoder acoustic features”); generating, using a decoder neural network of the text-to-speech conversion system, multiple frames of a spectrogram based on the encoded representations (pp. 2, section 3 and Fig. 1, An attention-based recurrent sequence generator is a RNN that receives and processes text input in order to generate vocoder feature frames; pp. 3, Fig. 2 shows samples generated by our model and their corresponding alignments to the text); processing, using a post-processing neural network of the text-to-speech conversion system, the [spectrogram] to generate a waveform synthesizer input; and generating a waveform of the verbal utterance of the text input based on the waveform synthesizer input (pp. 2, section 3 and Fig. 1, mapping from a sequence of vocoder features to corresponding audio samples to generate audio a waveform). Sotelo does not explicitly teach the bracketed limitation, however, Pollet does explicitly teach including the bracketed limitation: processing, using a post-processing neural network of the text-to-speech conversion system, the [spectrogram] to generate a waveform synthesizer input; and generating a waveform of the verbal utterance of the text input based on the waveform synthesizer input ([0020] generating speech waveform from speech segments which may be based on time and/or frequency domain; Figs. 8 and 10, [0100]-[0102] an example of hybrid output speech synthesis. The top pane displays the spectrogram and the bottom pane displays the output speech signal). Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate a system and method of end-to-end speech synthesis as taught by Sotelo with applying the method of input spectrogram to generate the output speech signal as taught by Pollet to provide one or more parameters in the segment sequencing and track generation which are adapted in small steps so that speech synthesis gradually improves in quality at each iteration (Pollet, [0100]). Regarding claim 8, Sotelo in view of Pollet discloses the method of claim 2, and Sotelo further discloses: receiving a sequence of decoder inputs, wherein generating, using the decoder neural network of the text-to-speech conversion system, multiple frames of the spectrogram based on the encoded representations comprises: for each decoder input in the sequence of decoder inputs, processing, using the decoder neural network of the text-to-speech conversion system, the decoder input and the encoded representations to generate multiple frames of the spectrogram (pp. 2, section 3 and Fig. 1, An attention-based recurrent sequence generator is a RNN that receives and processes text input in order to generate vocoder feature frames; pp. 3, Fig. 2 shows samples generated by our model and their corresponding alignments to the text), and wherein a first decoder input in the sequence of decoder inputs is a predetermined initial frame (Fig. 1, pp. 2, section 3.1, decoder input is a frame). Sotelo does not explicitly teach, however, Pollet does explicitly teach: [spectrogram] (Figs. 8 and 10, [0100]-[0102] an example of hybrid output speech synthesis. The top pane displays the spectrogram and the bottom pane displays the output speech signal). Therefore, it would have been obvious to one of ordinary skill before the effective filing date of the claimed invention to incorporate a system and method of end-to-end speech synthesis as taught by Sotelo with applying the method of input spectrogram to generate the output speech signal as taught by Pollet to provide one or more parameters in the segment sequencing and track generation which are adapted in small steps so that speech synthesis gradually improves in quality at each iteration (Pollet, [0100]). Regarding claim 9, Sotelo in view of Pollet discloses the method of claim 2, and Pollet further discloses: wherein the spectrogram is a mel-scale spectrogram ([0044] Spectrum information represented in a specific form such as MEL-LSP's, MFCC's, MEL-CEPs, harmonic components, etc.). Regarding claim 10, Sotelo in view of Pollet discloses the method of claim 2, and Sotelo further discloses: wherein generating the waveform of the verbal utterance of the text input based on the waveform synthesizer input comprises: processing, using a waveform synthesizer of the text-to-speech conversion system, the waveform synthesizer input to generate the waveform of the verbal utterance of the text input (pp. 2, section 3 and Fig. 1, an attention-based recurrent sequence generator is a RNN that receives and processes text input in order to generate Audio waveform; mapping from a sequence of vocoder features to corresponding audio samples to generate audio a waveform; pp. 3, Fig. 2 shows samples generated by our model and their corresponding alignments to the text). Regarding claim 11, Sotelo in view of Pollet discloses the method of claim 10, and Sotelo further discloses: generating speech using the waveform; and providing the generated speech for playback (pp. 2, Fig. 1, generate audio waveform). Regarding claim 12, Sotelo in view of Pollet discloses the method of claim 10, and Sotelo further discloses: wherein the waveform synthesizer is a trainable spectrogram to waveform inverter (pp. 3, sections 4 and 5, training detail). Regarding claim 13, Sotelo in view of Pollet discloses the method of claim 10, and Sotelo further discloses: wherein the waveform synthesizer is a vocoder (pp. 2, section 3 and Fig. 1, An attention-based recurrent sequence generator is an RNN that receives and processes text input in order to generate vocoder feature frames). Regarding claim 14, Sotelo in view of Pollet discloses the method of claim 9, and Pollet further discloses: wherein the waveform synthesizer input is a linear- scale spectrogram of the verbal utterance of the text input ([0090] “The template vectors for augmenting the boundary models can be represented by a full vector, a piece-wise linear approximation”). Regarding claims 15, Claim 15 is the corresponding system claim to method claim 2. Therefore, claim 15 is rejected using the same rationale as applied to claim 2, above. Regarding claims 19-21, Claims 19-21 are the corresponding medium claims to method claims 2, 9 and 11. Therefore, claims 19-21 are rejected using the same rationale as applied to claims 2, 9 and 11 above . Allowable Subject Matter Claims 3-7 and 16-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and if rewritten or amended to overcome the rejection(s) under provisional double patent, set forth in this Office action. Conclusion 07-96 AIA The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see attached form PTO-892 . Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEONG-AH A. SHIN whose telephone number is (571)272-5933. The examiner can normally be reached 9 AM-3PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. Seong-ah A. Shin Primary Examiner Art Unit 2659 /SEONG-AH A SHIN/Primary Examiner, Art Unit 2659 Application/Control Number: 18/951,397 Page 2 Art Unit: 2659 Application/Control Number: 18/951,397 Page 3 Art Unit: 2659 Application/Control Number: 18/951,397 Page 4 Art Unit: 2659 Application/Control Number: 18/951,397 Page 5 Art Unit: 2659 Application/Control Number: 18/951,397 Page 6 Art Unit: 2659 Application/Control Number: 18/951,397 Page 7 Art Unit: 2659 Application/Control Number: 18/951,397 Page 8 Art Unit: 2659 Application/Control Number: 18/951,397 Page 9 Art Unit: 2659 Application/Control Number: 18/951,397 Page 10 Art Unit: 2659 Application/Control Number: 18/951,397 Page 11 Art Unit: 2659 Application/Control Number: 18/951,397 Page 12 Art Unit: 2659 Application/Control Number: 18/951,397 Page 13 Art Unit: 2659 Application/Control Number: 18/951,397 Page 14 Art Unit: 2659 Application/Control Number: 18/951,397 Page 15 Art Unit: 2659 Application/Control Number: 18/951,397 Page 16 Art Unit: 2659 Application/Control Number: 18/951,397 Page 17 Art Unit: 2659 Application/Control Number: 18/951,397 Page 18 Art Unit: 2659 Application/Control Number: 18/951,397 Page 19 Art Unit: 2659 Application/Control Number: 18/951,397 Page 20 Art Unit: 2659 Application/Control Number: 18/951,397 Page 21 Art Unit: 2659 Application/Control Number: 18/951,397 Page 22 Art Unit: 2659 Application/Control Number: 18/951,397 Page 23 Art Unit: 2659 Application/Control Number: 18/951,397 Page 24 Art Unit: 2659 Application/Control Number: 18/951,397 Page 25 Art Unit: 2659 Application/Control Number: 18/951,397 Page 26 Art Unit: 2659
Read full office action

Prosecution Timeline

Nov 18, 2024
Application Filed
Jun 05, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT
Aug 04, 2026
Applicant Interview (Telephonic)
Aug 09, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725610
VOICE RECOGNITION SYSTEM, SERVER, DISPLAY APPARATUS AND CONTROL METHODS THEREOF
3y 11m to grant Granted Sep 01, 2026
Patent 12725609
Hotwording by Degree
2y 2m to grant Granted Sep 01, 2026
Patent 12694220
METHOD AND SYSTEM FOR PERSONALIZED EMBEDDING SEARCH ENGINE
3y 5m to grant Granted Jul 28, 2026
Patent 12682898
KEY PHRASE SPOTTING
2y 3m to grant Granted Jul 14, 2026
Patent 12670904
SELECTING AN AUTOMATED ASSISTANT AS THE PRIMARY AUTOMATED ASSISTANT FOR A DEVICE BASED ON DETERMINED AFFINITY SCORES FOR CANDIDATE AUTOMATED ASSISTANTS
3y 6m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+21.6%)
2y 7m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 421 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month