Prosecution Insights
Last updated: August 17, 2026
Application No. 18/511,369

ELECTRONIC DEVICE INCLUDING TEXT TO SPEECH MODEL AND METHOD FOR CONTROLLING THE SAME

Final Rejection §103
Filed
Nov 16, 2023
Priority
Nov 16, 2022 — RE 10-2022-0153726 +2 more
Examiner
ARMSTRONG, ANGELA A
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Korea Advanced Institute of Science and Technology
OA Round
2 (Final)
74%
Grant Probability
Favorable
3-4
OA Rounds
1y 0m
Est. Remaining
82%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
484 granted / 654 resolved
+12.0% vs TC avg
Moderate +8% lift
Without
With
+8.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 10m
Avg Prosecution
34 currently pending
Career history
678
Total Applications
across all art units

Statute-Specific Performance

§101
21.6%
-18.4% vs TC avg
§103
44.8%
+4.8% vs TC avg
§102
13.5%
-26.5% vs TC avg
§112
7.7%
-32.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 654 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This Office Action is in response to the amendment filed April 16, 2026. Claims 1-5, 8, 10-15, and 18-20 have been amended. Claims 6 and 16 have been amended. Claims 1-5, 7-15, and 17-20 remain pending. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claims 1-3, 10-13, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Shin et al (US Patent Application Publication NO. 2019/0385592), hereinafter Shin, in view of Wang et al (CN 103559537 -English Translation), hereinafter Wang. Regarding claims 1 and 11, Shin teaches a device (with memory and processor) and method [Fig 1, Fig 2, para 0041] comprising instructions and steps to: obtain audio data based on a text to speech (TTS) model (310) wherein the audio data corresponds to an input text [para 0133 -- the speech synthesis model 310 may learn text data of the first training data and output speech data (first speech data) based on speech corresponding to the text], based on identifying that the audio data comprises an error, identify an error part of the audio data which comprises the identified error [para 00136-0137; 0142-0143; 0147 -- The speech recognition model 320 may learn speech recognition based on the first speech data. The first speech data may be data obtained from the speech synthesis model 310. The first speech recognition result may be text data; controller 330 may control to change the parameter of the speech synthesis model 310 based on the first speech recognition result obtained from the speech recognition model 320 ]. Shin fails to specifically teach identify an activity of each of the plurality of nodes related to the error part, and modify a subset of the plurality of nodes based on the identified activity of the subset, wherein the subset comprises at least one node. IN a similar field of endeavor, Wang teaches template matching based on error back propagation, providing for determining an output layer node error, a hidden node error from a model with a plurality of neurons, nodes and weights, operating on the subset of nodes that are the ‘error nodes’; calculating an error rate for the nodes according to the template; and adjusting weights of nodes having an error rate (note, which is a subset of the node network) greater than or equal to a threshold value [para 0010; 0030-0039]. One having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the node error detection and weight modifications suggested by Wang, in the system of Shin, and the results would have been predictable and would provide an improved TTS model and thereby increase system performance and the user’s experience, Regarding claims 2 and 12, the combination of Shin and Wang teaches reducing a weight related to the at least one node [Wang’s node weight adjustments – para 0030-0039]. Regarding claims 3 and 13, the combination of Shin and Wang teaches replacing the subset with at least one pre-stored node and, wherein the at least one pre-stored node is stored in the memory and corresponds to text corresponding to the error part [Wang para 0035 – invalid nodes deleted..where replacing invalid nodes with known optimum nodes is an obvious step requiring only routine skill in the art]. Regarding claims 10 and 19, the combination of Shin and Wang teaches control the communication module to transmit to a server information related to the error part and the modification of the subset, receive, through the communication module, a modified TTS model from the server, and update the TTS model stored in the memory based on the modified TTS mode [Shin’s AI server processing including learning processor, model storage and communication unit for data transmission and receival –para 0042; 0065-0067]. Regarding claim 20, Shin teaches a device (with memory and processor) and method [Fig 1, Fig 2, para 0041] comprising instructions and steps to: obtain audio data based on a text to speech (TTS) model (310) wherein the audio data corresponds to an input text [para 0133 -- the speech synthesis model 310 may learn text data of the first training data and output speech data (first speech data) based on speech corresponding to the text], based on identifying that the audio data comprises an error, identify an error part of the audio data which comprises the identified error [para 00136-0137; 0142-0143; 0147 -- The speech recognition model 320 may learn speech recognition based on the first speech data. The first speech data may be data obtained from the speech synthesis model 310. The first speech recognition result may be text data; controller 330 may control to change the parameter of the speech synthesis model 310 based on the first speech recognition result obtained from the speech recognition model 320 ]. Shin fails to specifically teach identify an activity of each of the plurality of nodes related to the error part, and reduce a weight related to a subset of the plurality of nodes based on the identified activity of the subset. IN a similar field of endeavor, Wang teaches template matching based on error back propagation, providing for determining an output layer node error, a hidden node error from a model with a plurality of neurons, nodes and weights, operating on the subset of nodes that are the ‘error nodes’; calculating an error rate for the nodes according to the template; and adjusting weights of nodes having an error rate (note, which is a subset of the node network) greater than or equal to a threshold value [para 0010; 0030-0039]. One having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the node error detection and weight modifications suggested by Wang, in the system of Shin, and the results would have been predictable and would provide an improved TTS model and thereby increase system performance and the user’s experience, Claims 4-5, 7, 14-15 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Shin in view of Wang and further in view of Won et al (KR 10-2386635 – English translation), hereinafter Won. Regarding claims 4 and 14, Shin and Wang fail to specifically teach identifying that the audio data includes at least one phoneme having a length equal to or greater than a preset length, identify a part of the voice signal corresponding to the at least one phoneme as the error part. Won teaches the evaluation items of speech synthesis data may include the presence of errors for each analysis unit included in the speech synthesis data, the type of error, the degree of error, and the location of the analysis unit determined as an error in the sentence corresponding to the speech synthesis data [para 0058]. One having ordinary skill in the art would have recognized the advantages of implementing the speech synthesis evaluation techniques suggested by Won, in the system of Shin/Wang, and the results would be predictable so as to determine specific locations and identifications of errors to more clearly determine which portions of the model should be adjusted, so as to provide an improved TTS model and thereby increase system performance and the user’s experience. Regarding claims 5 and 15, the combination of Shin, Wang and Won teaches based identifying that the audio data includes a waveform part having an abnormal waveform, identify the waveform part as the error part [Won’s speech synthesis evaluation – para 0058]. Regarding claims 7 and 17, the combination of Shin, Wang and Won teaches display the input text on the display, and identify the error part based on a user input received through the display, wherein the user input comprises selection of a portion of the input text [Won’s user interface and checkbox features – para 0077]. Claims 8-9 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Shin in view of Wang and further in view of Yun et al (KR 10-2018-0039371 – English translation), hereinafter Yun. Regarding claims 8-9 and 18, Shin and Wang fail to teach identifying a sentence structure of the input text; obtaining, based on the sentence structure, at least one character string; obtaining a character string audio data resulting from inputting the at least one character string into the TTS model; and identifying, based on the character string audio data, whether the error part has been modified, wherein the obtaining at least one character string comprises changing a text before or after a portion of the input text corresponding to the error part. Yun teaches performing speech synthesis by receiving text as an input from a user; when determining that there is an error in text as the speech recognition, providing and displaying recognition error content and an error correction request to and on an interface unit; and when an agent unit is provided with a response to an error correction request, reflecting the response, correcting the error, and then providing error-corrected text [para 0051-0054]. One having ordinary skill in the art at the time of the invention would have recognized the advantages of implementing the correction of errors using user interactions as suggested by Yun, in the system of Shin/Wang, so as to allow for correcting an error of an automatic interpretation through interaction with a user, and the results would have been predictable and provided a more user-friendly system and thereby enhance and improve the user’s experience. Response to Arguments Applicant's arguments filed 04/26/2026 have been fully considered but they are not persuasive. Commencing on pp 10 of the response, applicants replicate the sections of the Shin reference that were applied to the claim 1; bottom of pp 10, applicants contend that “Shin, however, does not disclose or suggest to ‘identify an error part of the audio data which comprises the identified error’”; examiner argues that, in the comparison of “eat” and “eaaaaaat”, clearly, the system of Shin is noting/marking the section of data that contains the error; furthermore, examiner notes, that Shin does not limit the error to “a word” – looking at para 142-143 of Shin, reflecting back to para 0141 – Shin teaches that the error can be as large as a sentence – repeating, clearly, that Shin tracks/marks a ‘section’ of the text data that contains the error. On page 11 of the response, applicants argue that “Wang, however, does not disclose or suggest to ‘modify a subset of the plurality of nodes based on the identified activity of the subset, wherein the subset comprises at least one node’”; examiner notes that the arguments are toward the amended claim language; see mapping above, in the rejection, and that the correction-of-nodes technique in Wang, operates on the subsection of nodes, that are invalid. Examiner disagrees with applicants contention that the correction occurs “across the entire network” – in fact, examiner notes that, yes, Wang operates on the “portion of the network that contains the errors”; hence, the claimed ‘subset’ is taught by Wang. The ‘back propagation’ is to reset the remaining network, according to the ‘error corrected nodes’; ie, the weights on the error section are corrected so as not to create an error, and then the whole network is automatically activated and responds accordingly to the ‘corrected subsection of error nodes’. Lastly, in response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANGELA A ARMSTRONG whose telephone number is (571)272-7598. The examiner can normally be reached M,T,TH,F 11:30-8:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. ANGELA A. ARMSTRONG Primary Examiner Art Unit 2659 /ANGELA A ARMSTRONG/Primary Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Nov 16, 2023
Application Filed
Jan 16, 2026
Non-Final Rejection mailed — §103
Apr 16, 2026
Response Filed
Jul 01, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688849
CONVERSATIONAL DIGITAL ASSISTANT
2y 10m to grant Granted Jul 21, 2026
Patent 12682187
UNIFIED NATURAL LANGUAGE MODEL WITH SEGMENTED AND AGGREGATE ATTENTION
3y 8m to grant Granted Jul 14, 2026
Patent 12682158
LABEL INDUCTION
3y 8m to grant Granted Jul 14, 2026
Patent 12640146
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD AND RECORDING MEDIUM
3y 6m to grant Granted May 26, 2026
Patent 12640140
ELECTRONIC APPARATUS AND CONTROLLING METHOD THEREOF
3y 2m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
74%
Grant Probability
82%
With Interview (+8.4%)
3y 10m (~1y 0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 654 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month