Prosecution Insights
Last updated: August 17, 2026
Application No. 18/951,356

ON-DEVICE TEXT-TO-SPEECH MODEL PERSONALIZATION

Non-Final OA §101§103
Filed
Nov 18, 2024
Examiner
WEAVER, ADAM MICHAEL
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Qualcomm Incorporated
OA Round
1 (Non-Final)
87%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 87% — above average
87%
Career Allowance Rate
13 granted / 15 resolved
+24.7% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
20 currently pending
Career history
49
Total Applications
across all art units

Statute-Specific Performance

§101
34.5%
-5.5% vs TC avg
§103
44.1%
+4.1% vs TC avg
§102
17.0%
-23.0% vs TC avg
§112
2.3%
-37.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 15 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim(s) 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Step 1: Independent claims 1, 17, and 20 recite a system, method, and computer-readable medium (CRM), respectively. These claims therefore invoke a statutory category (machine and process) in Step 1 of the Subject Matter Eligibility Test. Step 2A, Prong One: Independent claims 1, 17, and 20, under their broadest reasonable interpretation, recite a system, method, and CRM of processing user’s audio signals (speech), checking the speech signals against a transcription of the same speech signal to get a confidence value, and checking the speech signals against a generated version of the same speech signals to get a loss value or checking the transcription against a lexical diversity criteria. This is an abstract idea in the form of certain methods of organizing human activity (i.e. mental processes such as observation, evaluation, judgement, and opinion), as well as various mathematical operations (i.e. calculating a confidence value, calculating a loss value). The steps of receiving audio signals and processing the audio signals, the comparison of the audio signals to their transcripts, and the comparison of the audio signals to a TTS output could be performed by a human using pen and paper or by purely mental reasoning. Step 2A, Prong Two: The claims do not integrate the judicial exception into a practical application. The recitation of “an automatic speech recognition (ASR)” and “a personalized text-to-speech (TTS)” are generic instructions to perform the abstract idea on/using a computer and do not impose a meaningful limit on the judicial exception. The ASR and personalized TTS are recited at high-levels of generality and are merely used as tools to perform the abstract idea faster and more efficiently. The audio signal gathering and analysis steps required to perform the method themselves do not add a meaningful limitation to the method. Mere data gathering and analysis do not provide an inventive concept. There is no improvement to the functioning of the ASR, the personalized TTS, the functioning of the computer itself, or to any other technology or technical field. Step 2B: The claims do not include any additional elements that amount to significantly more than the judicial exception. The only additional elements beyond the abstract idea are the ASR and the personalized TTS, which perform generic computational functions such as receiving, analyzing, and outputting data. Such elements are well-understood, routine, and conventional within the field. Accordingly, claims 1, 17, and 20 are directed to an abstract idea and do not include significantly more than the abstract idea itself. With respect to claim 2, the claim relates to the adequacy of the criteria checks relating to the fitness of the speech samples being used for the personalized TTS model. This is insignificant extra-solution activity; pre-solutional activities do not provide an inventive concept. The only additional element is the “personalized TTS model”, which is a generic instruction to perform the abstract idea on/using a computer and does not impose a meaningful limit on the judicial exception. No additional elements are present. With respect to claim 3, the claim relates to, following the criteria checks, the audio samples including audio samples that were chosen because they exceeded a confidence threshold, exceeded a loss threshold, or satisfied a lexicon diversity criteria. This is insignificant extra-solution activity; mere data gathering does not provide an inventive concept. No additional elements are present. With respect to claim 4 and claim 5, the claims relate to checking a signal-to-noise ratio (SNR) of the audio signals against an SNR threshold, associating the speech samples with that SNR value that exceeds the threshold, measuring the SNR value of the sample, and comparing it to the SNR threshold. Calculating an SNR value is a purely mathematical operation and could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present. With respect to claim 6, the claim relates to the processors being configured to adapt a personalized TTS model. This is insignificant extra-solution activity; pre-solutional activities do not provide an inventive concept. The only additional element is the “personalized TTS model”, which is a generic instruction to perform the abstract idea on/using a computer and does not impose a meaningful limit on the judicial exception. No additional elements are present. With respect to claim 7, the claim relates to using a trigger to begin adapting the personalized TTS model. This trigger is described generally with respect to the claim and therefore could be anything. This could be performed by a human using pen and paper or by purely mental reasoning. The only additional element is the “personalized TTS model”, which is a generic instruction to perform the abstract idea on/using a computer and does not impose a meaningful limit on the judicial exception. No additional elements are present. With respect to claim 8, the claim relates to the trigger condition causing a transition of a device into a different mode. This is described generally and could be performed by a human, save for the recitation of generic computer components. The only additional element is the “personalized TTS model”, which is a generic instruction to perform the abstract idea on/using a computer and does not impose a meaningful limit on the judicial exception. No additional elements are present. With respect to claim 9, the claim relates to performing ASR operations on the audio signal to generate a transcript and a confidence value and comparing the confidence value to the threshold. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. The only additional element is the “ASR”, which is a generic instruction to perform the abstract idea on/using a computer and does not impose a meaningful limit on the judicial exception. No additional elements are present. With respect to claim 10, the claim relates to inputting the ASR transcription into a personalized TTS model to obtain an output, calculating a loss value based on the comparison between the transcription and the output, and comparing that loss value to a threshold. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. The only additional elements are the “ASR” and the “personalized TTS model”, which are generic instructions to perform the abstract idea on/using a computer and do not impose a meaningful limit on the judicial exception. No additional elements are present. With respect to claims 11-13, the claims relate to comparing the ASR transcription to a reference, wherein that reference could be a vocabulary or another ASR transcription, and determining whether or not the original transcription satisfies the lexicon diversity criteria. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. The only additional elements are the “ASR” and the “personalized TTS model”, which are generic instructions to perform the abstract idea on/using a computer and do not impose a meaningful limit on the judicial exception. No additional elements are present. With respect to claim 14, the claim relates to utilizing microphones to capture the audio signals. This is insignificant extra-solution activity; mere data gathering and pre-solution activity does not provide an inventive concept. No additional elements are present. With respect to claim 15 and claim 16, the claims relate to the processors being integrated into any of a mobile phone, a table computer, a wearable electronic device, a camera device, or a vehicle. This is insignificant extra-solution activity; mere data gathering and pre-solution activity does not provide an inventive concept. No additional elements are present. With respect to claim 18, the claim relates to performing noise reduction, filtering the audio signals based on speech being identified, and filtering the audio signals to remove those that include non-user speech and do not include the user speech. Noise reduction is a purely mathematical operation and could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present. With respect to claim 19, the claim relates to an output being generated by the personalized TTS model. This is insignificant extra-solution activity; generic data outputting does not provide an inventive concept. No additional elements are present. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 6-7, 9, and 11-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mamkina et al. (US Patent No. 10,755,709), hereinafter referred to as Mamkina, in view of Ye et al. (US Patent No. 11,587,569), hereinafter referred to as Ye. Regarding claim 1, Mamkina discloses a device comprising: a memory configured to store a set of speech samples (Mamkina Fig. 10 reference character 1006 and 1008); and one or more processors coupled to the memory (Mamkina Fig. 10 reference characters 1004 and 290), wherein the one or more processors are configured to: obtain, during normal operation of the device, one or more audio signals that include user speech (Mamkina Fig. 9A reference character 902); and perform a sequence of sample criteria checks on the speech samples associated with the one or more audio signals, wherein the sequence of sample criteria checks include: a check whether a confidence value associated with an automatic speech recognition (ASR) transcription of a sample exceeds a transcription confidence threshold (Mamkina Fig. 9A reference characters 906-914 and Fig. 9B reference character 916). However, Mamkina fails to disclose and a check whether a loss value associated with a personalized text-to-speech (TTS) output of the sample exceeds a loss threshold, the ASR transcription satisfies a lexicon diversity criterion, or both. Ye teaches a system and method for generating and using text-to-speech (TTS) data for improved speech recognition models. Ye teaches and a check whether a loss value associated with a personalized text-to-speech (TTS) output of the sample exceeds a loss threshold, the ASR transcription satisfies a lexicon diversity criterion, or both ("In some embodiments, the new TTS training is obtained from a multi-speaker neural TTS system for a keyword that is underrepresented in the baseline training data. In some embodiments, the new TTS training data is used for pronunciation learning and normalization of keyword dependent confidence scores in keyword spotting (KWS) applications," Ye col. 4 lines 61-67). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Ye’s teaching of comparing an ASR transcription against a lexical reference. Comparing a transcription of speech against that of a lexical reference, in this case a document of keywords that might be commonly uttered by a user, works to ensure that the speech captured matches with words listed in the lexical reference. This facilitates the learning of the user’s speech patterns and habits, ensuring a more holistic training of the model. Regarding claim 2, Mamkina, in view of Ye, discloses all of the limitations of claim 1. However, Mamkina fails to disclose wherein the sequence of sample criteria checks is related to fitness of the speech samples for use in adapting a personalized TTS model at the device. Ye teaches wherein the sequence of sample criteria checks is related to fitness of the speech samples for use in adapting a personalized TTS model at the device ("In some instances, the new TTS training data is speech that is personalized to a particular speaker in terms of acoustic features of the particular speaker and/or in terms of content found in speech typically spoken by the particular speaker. Thus, wherein there is limited adaptation data (e.g., natural speech data) available for the particular speaker, a main model can undergo efficient and effective rapid speaker adaptation utilizing the personalized speech generated from a neural network language model (NNLM) generator and a neural TTS system," Ye col. 3 lines 24-34). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Ye’s teaching of utilizing criteria checks to ensure that speech samples are adequate to train a TTS model. In this, the comparison of a transcription of speech against that of a lexical reference, here being a document of keywords that might be commonly uttered by a user, works to ensure that the speech captured matches with words listed in the lexical reference. This facilitates the learning of the user’s speech patterns and habits, ensuring a more holistic training of the model. Regarding claim 3, Mamkina, in view of Ye, discloses all of the limitations of claim 1. Mamkina further discloses wherein, after performance of the sequence of sample criteria checks, the set of speech samples includes one or more speech samples that are associated with a corresponding confidence value that exceeds the transcription confidence threshold (Mamkina Fig. 9A reference characters 906-914 and Fig. 9B reference character 916). However, Mamkina fails to disclose and that are associated with a corresponding loss value that exceeds the loss threshold or a corresponding ASR transcript that satisfies the lexicon diversity criterion. Ye teaches and that are associated with a corresponding loss value that exceeds the loss threshold or a corresponding ASR transcript that satisfies the lexicon diversity criterion ("In some embodiments, the new TTS training is obtained from a multi-speaker neural TTS system for a keyword that is underrepresented in the baseline training data. In some embodiments, the new TTS training data is used for pronunciation learning and normalization of keyword dependent confidence scores in keyword spotting (KWS) applications," Ye col. 4 lines 61-67). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Ye’s teaching of including speech samples that exceeded a comparison score of an ASR transcription against a lexical reference. Comparing a transcription of speech against that of a lexical reference, in this case a document of keywords that might be commonly uttered by a user, works to ensure that the speech captured matches with words listed in the lexical reference. This facilitates the learning of the user’s speech patterns and habits, ensuring a more holistic training of the model. Regarding claim 6, Mamkina, in view of Ye, discloses all of the limitations of claim 1. However, Mamkina fails to disclose wherein the one or more processors are further configured to adapt a personalized TTS model based on the set of speech samples. Ye teaches wherein the one or more processors are further configured to adapt a personalized TTS model based on the set of speech samples ("Disclosed embodiments also include systems, methods, and devices that can be used to facilitate improved techniques for generating TTS data to modify speech recognition models, such as in utilizing personalized speech synthesis for rapid speaker adaption," Ye col. 2 lines 44-48). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Ye’s teaching of using the speech samples to create a personalized TTS model. This would help the model to facilitate enhanced user engagement and improved accessibility and comprehension. It would make it easier for the user to understand what the TTS model is outputting, and also it would help to further train the model itself to become more consistent and more accurate. Regarding claim 7, Mamkina, in view of Ye, discloses all of the limitations of claim 6. Mamkina further discloses wherein the one or more processors are further configured to adapt the personalized TTS model based on detection of a trigger condition associated with the device ("The device 110, using a wakeword detection module 220, then processes audio data corresponding to the input audio 11 to determine if a keyword (such as a wakeword) is detected in the audio data. Following detection of a wakeword, the speech-controlled device 110 sends audio data 111, corresponding to the utterance, to a server 120 that includes an ASR module 250. The audio data 111 may be output from an acoustic front end (AFE) 256 located on the device 110 prior to transmission, or the audio data 111 may be in a different form for processing by a remote AFE 256, such as the AFE 256 located with the ASR module 250," Mamkina col. 6 lines 29-40). Regarding claim 9, Mamkina, in view of Ye, discloses all of the limitations of claim 1. Mamkina further discloses wherein the one or more processors are further configured to: perform one or more ASR operations on the sample to generate the ASR transcription and the confidence value, wherein the ASR transcription includes text data that represents the user speech included in the sample, and wherein the confidence value indicates a confidence that the text data matches the user speech (Mamkina Fig. 9A reference character 906 and 914); and compare the confidence value to the transcription confidence threshold (Mamkina Fig. 9A reference character 914 and Fig. 9B reference character 916). Regarding claim 11, Mamkina, in view of Ye, discloses all of the limitations of claim 1. Mamkina further discloses wherein the one or more processors are further configured to: compare the ASR transcription to a reference (Mamkina Fig. 8 reference characters 250 and 807). However, Mamkina fails to disclose and determine whether the ASR transcription satisfies the lexicon diversity criterion based on the comparison. Ye teaches and determine whether the ASR transcription satisfies the lexicon diversity criterion based on the comparison ("Next, the computing system determines that the keyword received in the preceding act (act 420) may be a keyword that is determined to be underrepresented in the baseline training data (act 430). The baseline training data is data used to train the model identified in the first act (act 410) of method 400. (See the description for FIG. 5 for further information on determining that a keyword is underrepresented)," Ye col. 9 lines 6-14). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Ye’s teaching of comparing an ASR transcription against a lexical reference. Comparing a transcription of speech against that of a lexical reference, in this case a document of keywords that might be commonly uttered by a user, works to ensure that the speech captured matches with words listed in the lexical reference. This facilitates the learning of the user’s speech patterns and habits, ensuring a more holistic training of the model. Regarding claim 12, Mamkina, in view of Ye, discloses all of the limitations of claim 11. However, Mamkina fails to disclose wherein the reference includes a vocabulary associated with initial training of a personalized TTS model. Ye teaches wherein the reference includes a vocabulary associated with initial training of a personalized TTS model ("Referring now to FIG. 5, a data chart 501 is shown that illustrates a plurality of keywords (510, 520) mapped according to a scale of confidence scores 530. The confidence scores 530 for particular words can be used as a basis to determine a likelihood that a KWS system will be able to accurately recognize the particular keywords when detected by the KWS systems," Ye col. 9 lines 31-37). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Ye’s teaching of including a vocabulary or a document of keywords associated with training a TTS model. Comparing a transcription of speech against that of a lexical reference, in this case a document of keywords that might be commonly uttered by a user, works to ensure that the speech captured matches with words listed in the lexical reference. This facilitates the learning of the user’s speech patterns and habits, ensuring a more holistic training of the model. Regarding claim 13, Mamkina, in view of Ye, discloses all of the limitations of claim 11. However, Mamkina fails to disclose wherein the reference includes at least a portion of one or more ASR transcriptions of one or more of the set of speech samples. Ye teaches wherein the reference includes at least a portion of one or more ASR transcriptions of one or more of the set of speech samples ("The computing system obtains transcripts of live queries with a designated keyword (act 610). In some embodiments, the new TTS training data corresponds to the new TTS training data 160 of FIG. 1," Ye col. 11 lines 3-6). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Ye’s teaching of including a reference of other ASR transcriptions. Comparing a current transcription of speech against that of past transcriptions works to ensure that the speech captured matches with that of historical speech patterns and habits. This facilitates the more comprehensive learning of the user’s speech patterns and habits, ensuring a more holistic training of the model. Regarding claim 14, Mamkina, in view of Ye, discloses all of the limitations of claim 1. Mamkina further discloses further comprising one or more microphones coupled to the one or more processors and configured to capture the one or more audio signals (Mamkina Fig. 1 reference character 103). Regarding claim 15, Mamkina, in view of Ye, discloses all of the limitations of claim 1. Mamkina further discloses wherein the one or more processors are integrated in at least one of a mobile phone (Mamkina Fig. 12 reference character 110b), a tablet computer device (Mamkina Fig. 12 reference character 110d), a wearable electronic device (Mamkina Fig. 12 reference character 110c), or a camera device ("The device 110 may additionally include an image or video capture component, such as the camera 115," Mamkina col. 36 lines 10-12), and wherein the mobile phone, the tablet computer device, the wearable electronic device, or the camera device is configured to perform the sequence of sample criteria checks ("As illustrated in FIG. 12, multiple devices (120, 110a-110e, 1202, 1204) may contain components of the system 100 and the devices may be connected over a network 199," Mamkina col. 36 lines 60-62). Regarding claim 16, Mamkina, in view of Ye, discloses all of the limitations of claim 1. Mamkina further discloses wherein the one or more processors are integrated in a vehicle (Mamkina Fig. 12 reference character 110e) that is configured to perform the sequence of sample criteria checks ("As illustrated in FIG. 12, multiple devices (120, 110a-110e, 1202, 1204) may contain components of the system 100 and the devices may be connected over a network 199," Mamkina col. 36 lines 60-62). As to claim 17, method claim 17 and system claim 1 are related as system and method of using same, with each claimed element’s function corresponding to the system step. Accordingly, claim 17 is similarly rejected under the same rationale as applied above with respect to the system claim. Regarding claim 18, Mamkina, in view of Ye, discloses all of the limitations of claim 17. Mamkina further discloses further comprising, prior to performing the sequence of sample criteria checks: performing one or more noise reduction operations on the speech samples ("The AFE 256 may reduce noise in the audio data 111 and divide the digitized audio data 111 into frames representing time intervals for which the AFE 256 determines a number of values (i.e., features) representing qualities of the audio data 111, along with a set of those values (i.e., a feature vector or audio feature vector) representing features/qualities of the audio data 111 within each frame," Mamkina col. 8 lines 30-36); performing a filtering process on the speech samples, wherein the filtering process includes: performing user identification on the speech samples to identify the user speech and non-user speech ("Further, audio frames during the utterance that do not include speech may be filtered out by the VAD detector 610 and thus not considered by the ASR feature extraction 606 and/or user recognition feature extraction 608," Mamkina col. 24 lines 32-36); and filtering the speech samples to remove one or more samples that include the non-user speech and do not include the user speech; or a combination thereof ("Further, audio frames during the utterance that do not include speech may be filtered out by the VAD detector 610 and thus not considered by the ASR feature extraction 606 and/or user recognition feature extraction 608," Mamkina col. 24 lines 32-36). Regarding claim 19, Mamkina, in view of Ye, discloses all of the limitations of claim 17. However, Mamkina fails to disclose wherein the personalized TTS output is generated by a personalized TTS model at the device that is configured to mimic pronunciation of one or more test users. Ye teaches wherein the personalized TTS output is generated by a personalized TTS model at the device that is configured to mimic pronunciation of one or more test users ("In some instances, the new TTS training data is speech that is personalized to a particular speaker in terms of acoustic features of the particular speaker and/or in terms of content found in speech typically spoken by the particular speaker. Thus, wherein there is limited adaptation data (e.g., natural speech data) available for the particular speaker, a main model can undergo efficient and effective rapid speaker adaptation utilizing the personalized speech generated from a neural network language model (NNLM) generator and a neural TTS system," Ye col. 3 lines 24-34). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Ye’s teaching of using personalized TTS model to generate output. This is an obvious and natural next step that would occur upon creating and training a personalized TTS model. Output generation is the goal of creating the model. As to claim 20, computer-readable medium (CRM) claim 20 and system claim 1 are related as system and CRM of using same, with each claimed element’s function corresponding to the system step. Accordingly, claim 20 is similarly rejected under the same rationale as applied above with respect to the system claim. Claim(s) 4-5, 8, and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mamkina, in view of Ye, and further in view of Sung et al. (US Patent Application Publication No. 2022/0301542). Regarding claim 4, Mamkina, in view of Ye, discloses all of the limitations of claim 1. Mamkina further discloses wherein the sequence of sample criteria checks further includes a check whether a signal-to-noise ratio (SNR) value, and wherein, after performance of the sequence of sample criteria checks, each speech sample of the set of speech samples is associated with a corresponding SNR value ("The device 110 may use various techniques to determine whether audio data includes speech. Some embodiments may apply voice activity detection (VAD) techniques. Such techniques may determine whether speech is present in input audio based on various quantitative aspects of the input audio, such as a spectral slope between one or more frames of the input audio; energy levels of the input audio in one or more spectral bands; signal-to-noise ratios of the input audio in one or more spectral bands; or other quantitative aspects," Mamkina col. 6 lines 50-58). However, Mamkina fails to disclose associated with the sample exceeds an SNR threshold, that exceeds the SNR threshold. Sung teaches a system and method for generating a personalized text-to-speech model. Sung teaches associated with the sample exceeds an SNR threshold, that exceeds the SNR threshold ("According to an embodiment, the electronic device 430 may verify the consistency between the additionally recorded data 1010 and the existing training data 1030 based on a signal-to-noise ratio (SNR). When the difference between an SNR of the additionally recorded data 1010 and an SNR of the existing training data 1030 is less than or equal to a threshold value, the electronic device 430 may determine the additionally recorded data 1010 and the existing training data 1030 to be consistent," Sung para [0116]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Sung’s teaching of utilizing an SNR threshold. Utilizing an SNR threshold in order to check whether or not a speech sample should be input to the next step is a common and well-known technique within the art. It allows a quick check to ensure that the speech sample can be heard and used correctly and accurately as training data. Regarding claim 5, Mamkina, in view of Ye, and further in view of Sung, discloses all of the limitations of claim 4. However, Mamkina fails to disclose wherein the one or more processors are further configured to: measure the SNR value associated with the sample; and compare the SNR value to the SNR threshold. Sung teaches wherein the one or more processors are further configured to: measure the SNR value associated with the sample ("According to an embodiment, the electronic device 430 may verify the consistency between the additionally recorded data 1010 and the existing training data 1030 based on a signal-to-noise ratio (SNR). When the difference between an SNR of the additionally recorded data 1010 and an SNR of the existing training data 1030 is less than or equal to a threshold value, the electronic device 430 may determine the additionally recorded data 1010 and the existing training data 1030 to be consistent," Sung para [0116]); and compare the SNR value to the SNR threshold ("According to an embodiment, the electronic device 430 may verify the consistency between the additionally recorded data 1010 and the existing training data 1030 based on a signal-to-noise ratio (SNR). When the difference between an SNR of the additionally recorded data 1010 and an SNR of the existing training data 1030 is less than or equal to a threshold value, the electronic device 430 may determine the additionally recorded data 1010 and the existing training data 1030 to be consistent," Sung para [0116]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Sung’s teaching of utilizing an SNR threshold. Utilizing an SNR threshold in order to check whether or not a speech sample should be input to the next step is a common and well-known technique within the art. It allows a quick check to ensure that the speech sample can be heard and used correctly and accurately as training data. Regarding claim 8, Mamkina, in view of Ye, discloses all of the limitations of claim 7. However, Mamkina fails to disclose wherein the trigger condition includes transition of the device to a sleep mode, detection of a target time of day, receipt of a user input associated with adapting the personalized TTS model, operation of the device in a low power operating mode for a threshold time period, detection of the device being connected to an external power source, or a combination thereof. Sung teaches wherein the trigger condition includes transition of the device to a sleep mode, detection of a target time of day, receipt of a user input associated with adapting the personalized TTS model, operation of the device in a low power operating mode for a threshold time period, detection of the device being connected to an external power source, or a combination thereof ("The sensor module 1476 may detect an operational state (e.g., power or temperature) of the electronic device 1401 or an environmental state (e.g., a state of a user) external to the electronic device 1401, and generate an electric signal or data value corresponding to the detected state," Sung para [0147]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Sung’s teaching of providing a transition and a detection of a device in an operational state involving the power of the device. It is commonplace within the art to provide detection as to what state the power of the device is in, and this would have been an obvious inclusion here. Regarding claim 10, Mamkina, in view of Ye, discloses all of the limitations of claim 1. However, Mamkina fails to disclose wherein the one or more processors are further configured to: provide the ASR transcription to a personalized TTS model to generate the personalized TTS output of the sample; generate the loss value based on a comparison of the personalized TTS output to the sample; and compare the loss value to the loss threshold. Sung teaches wherein the one or more processors are further configured to: provide the ASR transcription to a personalized TTS model to generate the personalized TTS output of the sample ("According to an embodiment, the electronic device 430 may generate an intermediate result based on the intermediate model 570 in operation 513. The intermediate result may include a sound source corresponding to a text generated using the intermediate model 570, and may include a numerical value indicating the difference between the generated sound source and the sound source recorded from the same text from the user," Sung para [0086] and "According to an embodiment, the electronic device 430 may generate a sound source 853 corresponding to a text 851 using an intermediate model 830 (e.g., the intermediate model 570 of FIG. 5) in operation 812. The text 851 may be a text corresponding to a recorded sound source 855 (e.g., a sound source included in the training data 650 of FIG. 6)," Sung para [0100]); generate the loss value based on a comparison of the personalized TTS output to the sample (Sung Fig. 8 reference character 813); and compare the loss value to the loss threshold ("The comparison factor may indicate the spectral distance between the generated sound source 853 and the recorded sound source 855. The spectral distance may be a Euclidean distance calculated by extracting a mel-cepstrum from the two sound sources and aligning frames through dynamic time warping. The decrease in the spectral distance may indirectly indicate a decrease in the difference between the generated sound source 853 and the recorded sound source 855. Thus, when the comparison factor decreases as training progresses the user may verify that a speech model approaches the tone or accent of the target speaker. However, when the comparison factor no longer decreases despite the progress of the training, the sound source generated by the speech model trained up to that point may correspond to the best the speech model can simulate the target speaker. Thus, in such a case, it may be a factor that ends the training," Sung para [0101]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mamkina’s teaching of recognizing a user using a speech recognition system by including Sung’s teaching of utilizing the personalized TTS model to generate an output, comparing that output to the original sample, finding a loss value, and comparing that loss value to a threshold loss value. This is a common and well-known practice within the art to cause the convergence of models. Comparing the model’s output to a objective and correct example and using a loss threshold to ensure that the model is performing well enough is the most direct way to test the model’s performance. This would have been an obvious inclusion in this case. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: US Patent Application Publication No. 2025/0349282 US Patent Application Publication No. 2023/0267925 US Patent Application Publication No. 2022/0310058 Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADAM MICHAEL WEAVER whose telephone number is (571)272-7062. The examiner can normally be reached Monday-Friday, 8AM-5PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ADAM MICHAEL WEAVER/ Examiner, Art Unit 2658 /RICHEMOND DORVIL/ Supervisory Patent Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Nov 18, 2024
Application Filed
Jun 24, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12664978
FEDERATED KNOWLEDGE DISTILLATION ON AN ENCODER OF A GLOBAL ASR MODEL AND/OR AN ENCODER OF A CLIENT ASR MODEL
3y 6m to grant Granted Jun 23, 2026
Patent 12657219
INFORMATION PROCESSING DEVICE, COMPUTER PROGRAM PRODUCT, AND INFORMATION PROCESSING METHOD
2y 3m to grant Granted Jun 16, 2026
Patent 12651117
METHODS AND SYSTEMS FOR VERIFICATION OF PLANT PROCEDURES' COMPLIANCE TO WRITING MANUALS
4y 0m to grant Granted Jun 09, 2026
Patent 12651266
SYSTEMS AND METHODS FOR RANKING CALL INTENT PROBABILITY
2y 3m to grant Granted Jun 09, 2026
Patent 12639355
IDENTIFYING HALLUCINATIONS IN LARGE LANGUAGE MODEL OUTPUT
2y 9m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
87%
Grant Probability
99%
With Interview (+33.3%)
2y 6m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 15 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month