Prosecution Insights
Last updated: August 18, 2026
Application No. 18/789,219

METHODS AND SYSTEMS FOR TRAINING AN ARTIFICIAL INTELLIGENCE (AI) TOTAL DURATION-AWARE MODEL TO CONTROL THE TOTAL DURATION OF SPEECH UTTERANCES BY A TEXT-TO-SPEECH (TTS) COMPUTING SYTEM

Final Rejection §102§103
Filed
Jul 30, 2024
Priority
Jun 05, 2024 — provisional 63/656,238
Examiner
COLUCCI, MICHAEL C
Art Unit
2655
Tech Center
2600 — Communications
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
76%
Grant Probability
Favorable
3-4
OA Rounds
1y 1m
Est. Remaining
91%
With Interview

Examiner Intelligence

Grants 76% — above average
76%
Career Allowance Rate
765 granted / 1009 resolved
+13.8% vs TC avg
Strong +15% interview lift
Without
With
+15.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
34 currently pending
Career history
1050
Total Applications
across all art units

Statute-Specific Performance

§101
14.1%
-25.9% vs TC avg
§103
61.2%
+21.2% vs TC avg
§102
8.7%
-31.3% vs TC avg
§112
4.8%
-35.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1009 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Response to Arguments Applicant's arguments with respect to claims 1 and 15 have been considered but are moot in view of the new ground(s) of rejection. Applicant’s arguments are directed to the amended subject matter; new citations from existing prior art are provided below. Note: The claims are not directed towards patent ineligible subject matter under 35 U.S.C. 101 Step 1: IS THE CLAIM DIRECTED TO A PROCESS, MACHINE, MANUFACTURE OR COMPOSITION OF MATTER? Yes Step 2A.1: IS THE CLAIM DIRECTED TO A LAW OF NATURE, A NATURAL PHENOMENON (PRODUCT OF NATURE) OR AN ABSTRACT IDEA? No Step 2A.2: DOES THE CLAIM RECITE ADDITIONAL ELEMENTS THAT INTEGRATE THE JUDICIAL EXCEPTION INTO A PRACTICAL APPLICATION? Yes, if the claims are alternatively construed to be abstract in step 2A1. The claims seek to improve the output of predicted frames supported by the specification, and reflected by the claims e.g. in spec: 0031 0032 In other words, the claims enable the invention to increase accuracy of predicted frame durations. Supported by the following: In Finjan Inc. v. Blue Coat Systems, Inc., 879 F.3d 1299, 125 USPQ2d 1282 (Fed. Cir. 2018), the claimed invention was a method of virus scanning that scans an application program, generates a security profile identifying any potentially suspicious code in the program, and links the security profile to the application program. 879 F.3d at 1303-04, 125 USPQ2d at 1285-86. The Federal Circuit noted that the recited virus screening was an abstract idea, and that merely performing virus screening on a computer does not render the claim eligible. 879 F.3d at 1304, 125 USPQ2d at 1286. The court then continued with its analysis under part one of the Alice/Mayo test by reviewing the patent’s specification, which described the claimed security profile as identifying both hostile and potentially hostile operations. The court noted that the security profile thus enables the invention to protect the user against both previously unknown viruses and “obfuscated code,” as compared to traditional virus scanning, which only recognized the presence of previously-identified viruses. The security profile also enables more flexible virus filtering and greater user customization. 879 F.3d at 1304, 125 USPQ2d at 1286. The court identified these benefits as improving computer functionality, and verified that the claims recite additional elements (e.g., specific steps of using the security profile in a particular way) that reflect this improvement. Accordingly, the court held the claims eligible as not being directed to the recited abstract idea. 879 F.3d at 1304-05, 125 USPQ2d at 1286-87. This analysis is equivalent to the Office’s analysis of determining that the additional elements integrate the judicial exception into a practical application at Step 2A Prong Two, and thus that the claims were not directed to the judicial exception (Step 2A: NO). Examples of claims that improve technology and are not directed to a judicial exception include: Enfish, LLC v. Microsoft Corp., 822 F.3d 1327, 1339, 118 USPQ2d 1684, 1691-92 (Fed. Cir. 2016) (claims to a self-referential table for a computer database were directed to an improvement in computer capabilities and not directed to an abstract idea); McRO, Inc. v. Bandai Namco Games Am. Inc., 837 F.3d 1299, 1315, 120 USPQ2d 1091, 1102-03 (Fed. Cir. 2016) (claims to automatic lip synchronization and facial expression animation were directed to an improvement in computer-related technology and not directed to an abstract idea); Visual Memory LLC v. NVIDIA Corp., 867 F.3d 1253,1259-60, 123 USPQ2d 1712, 1717 (Fed. Cir. 2017) (claims to an enhanced computer memory system were directed to an improvement in computer capabilities and not an abstract idea); Finjan Inc. v. Blue Coat Systems, Inc., 879 F.3d 1299, 125 USPQ2d 1282 (Fed. Cir. 2018) (claims to virus scanning were found to be an improvement in computer technology and not directed to an abstract idea); SRI Int’l, Inc. v. Cisco Systems, Inc., 930 F.3d 1295, 1303 (Fed. Cir. 2019) (claims to detecting suspicious activity by using network monitors and analyzing network packets were found to be an improvement in computer network technology and not directed to an abstract idea). Additional examples are provided in MPEP § 2106.05(a). Regarding the December 5th 2025 Memo in light of September 26, 2025 Appeals Review Panel Decision in Ex parte Desjardins, Appeal 2024-000567 for Application 16/319,040, in deciding if a recited abstract idea does or does not direct the entire claim to an abstract idea, when a claim is considered as a whole: Paragraph 21 of the Specification, which the Appellant cites, identifies improvements in training the machine learning model itself. Of course, such an assertion in the Specification alone is insufficient to support a patent eligibility determination, absent a subsequent determination that the claim itself reflects the disclosed improvement. See MPEP § 2106.05(a) (citing Intellectual Ventures I LLC v. Symantec Corp., 838 F.3d 1307, 1316 (Fed. Cir. 2016)). Here, however, we are persuaded that the claims reflect such an improvement. For example, one improvement identified in the 8 Appeal2024-000567 Application 16/319,040 Specification is to "effectively learn new tasks in succession whilst protecting knowledge about previous tasks." Spec. ,r 21. The Specification also recites that the claimed improvement allows artificial intelligence (AI) systems to "us[e] less of their storage capacity" and enables "reduced system complexity." Id. When evaluating the claim as a whole, we discern at least the following limitation of independent claim 1 that reflects the improvement: "adjust the first values of the plurality of parameters to optimize performance of the machine learning model on the second machine learning task while protecting performance of the machine learning model on the first machine learning task." We are persuaded that constitutes an improvement to how the machine learning model itself operates, and not, for example, the identified mathematical calculation. Under a charitable view, the overbroad reasoning of the original panel below is perhaps understandable given the confusing nature of existing § 101 jurisprudence, but troubling, because this case highlights what is at stake. Categorically excluding AI innovations from patent protection in the United States jeopardizes America's leadership in this critical emerging technology. Yet, under the panel's reasoning, many AI innovations are potentially unpatentable-even if they are adequately described and nonobvious-because the panel essentially equated any machine learning with an unpatentable "algorithm" and the remaining additional elements as "generic computer components," without adequate explanation. Dec. 24. Examiners and panels should not evaluate claims at such a high level of generality. Specifically, Ex Parte Desjardins explained the following: Enfish ranks among the Federal Circuit's leading cases on the eligibility of technological improvements. In particular, Enfish recognized that “[m]uch of the advancement made in computer technology consists of improvements to software that, by their very nature, may not be defined by particular physical features but rather by logical structures and processes.” 822 F.3d at 1339. Moreover, because “[s]oftware can make non-abstract improvements to computer technology, just as hardware improvements can,” the Federal Circuit held that the eligibility determinations should turn on whether “the claims are directed to an improvement to computer functionality versus being directed to an abstract idea.” Id. at 1336. (Desjardins, page 8). Further in Ex Parte Desjardins, Appeal No. 2024-000567 (PTAB September 26, 2025, Appeals Review Panel Decision) (precedential), the claimed invention was a method of training a machine learning model on a series of tasks. The Appeals Review Panel (ARP) overall credited benefits including reduced storage, reduced system complexity and streamlining, and preservation of performance attributes associated with earlier tasks during subsequent computational tasks as technological improvements that were disclosed in the patent application specification. Specifically, the ARP upheld the Step 2A Prong One finding that the claims recited an abstract idea (i.e., mathematical concept). In Step 2A Prong Two, the ARP then determined that the specification identified improvements as to how the machine learning model itself operates, including training a machine learning model to learn new tasks while protecting knowledge about previous tasks to overcome the problem of “catastrophic forgetting” encountered in continual learning systems. Importantly, the ARP evaluated the claims as a whole in discerning at least the limitation “adjust the first values of the plurality of parameters to optimize performance of the machine learning model on the second machine learning task while protecting performance of the machine learning model on the first machine learning task” reflected the improvement disclosed in the specification. Accordingly, the claims as a whole integrated what would otherwise be a judicial exception instead into a practical application at Step 2A Prong Two, and therefore the claims were The claim itself does not need to explicitly recite the improvement described in the specification (e.g., “thereby increasing the bandwidth of the channel”). See, e.g., Ex Parte Desjardins, Appeal No. 2024-000567 (PTAB September 26, 2025, Appeals Review Panel Decision) (precedential), in which the specification identified the improvement to machine learning technology by explaining how the machine learning model is trained to learn new tasks while protecting knowledge about previous tasks to overcome the problem of “catastrophic forgetting,” and that the claims reflected the improvement identified in the specification. Indeed, enumerated improvements identified in the Desjardins specification included disclosures of the effective learning of new tasks in succession in connection with specifically protecting knowledge concerning previously accomplished tasks; allowing the system to reduce use of storage capacity; and the enablement of reduced complexity in the system. Such improvements were tantamount to how the machine learning model itself would function in operation and therefore not subsumed in the identified mathematical calculation. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 6, 10-14, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20240161730 A1 Elias; Isaac et al. (hereinafter Elias) in view of US 20250149023 A1 KIM; Tae Woo et al. (hereinafter KIM). Re claim 1, Elias teaches 1. A method for training an AI duration model to control the duration of speech utterances by a text-to-speech computing system when converting text into speech, the method comprising: (neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) providing training data to the AI duration model, the training data including a plurality of phonemes derived from a string of text, corresponding actual frame durations for each of the phonemes, and a target output speech time duration; (phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) … the predicted frame durations being generated such that a summation of frame durations for the plurality of phonemes is approximately equal to the target output speech time duration (input duration equals output duration 0060-0061… including in Mel time-freq input equals output 0052 as well i.e. time vs freq, wherein a time-frequency representation and temporal structure are directed to Mel) calculating a loss with a loss function to quantify a difference of at least the predicted frame durations and the actual frame durations, and using the loss to train the AI duration model by adjusting parameters of the AI duration model that are used to generate the predicted frame durations. (predicting frames 0054, calculating the loss difference between prediction and reference/actual as in fig. 2 and using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) …wherein the Al duration model is configured to generate, based on phonemes and corresponding frame durations, an audio representation having a temporal structure defined by the frame durations, and to convert the audio representation into a time-domain waveform having a duration corresponding to the target output speech time duration. (input duration equals output duration 0060-0061 including in Mel time-freq input equals output 0052 as well i.e. time vs freq, wherein a time-frequency representation and temporal structure are directed to Mel, and further thereof expressly converted into time domaon 0064) However, while Elias teaches neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: the AI duration model masking actual frame durations for a subset of the plurality of phonemes…; (KIM phoneme duration model specific to masked speech 0030 and 0043 with fig. 2) the AI duration model generating predicted frame durations for the masked actual frame durations of the subset of the plurality of phonemes; and (KIM prediction output based on phoneme duration model specific to masked speech 0030 and 0043 with fig. 2) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias to incorporate the above claim limitations as taught by KIM to allow for simple substation of one known element for another to yield predictable results, such as the masked speech in place of Elias speech, to allow the model to learn to infer durations from contextual information rather than relying on direct, unmasked acoustic cues, leading to more natural and controllable speech synthesis, to provide this option for more exact and smooth synthesis outputs. Re claim 6, Elias teaches 6. The method according to Claim 1, the method further comprising parsing the string of text into a plurality of phonemes. (using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) Re claim 10, Elias teaches 10. The method according to Claim 1, where the loss is calculated using cross-entropy loss. (using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) Re claim 11, Elias teaches 11. The method according to Claim 1, the method further comprising generating one or more audio representations based on the phonemes, frame time durations for the phonemes, and the target output speech time duration. (phoneme based, duration and frame based, and output speech dependent thereof…predicting frames 0054, calculating the loss difference between prediction and reference/actual as in fig. 2 and using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) Re claim 12, Elias teaches 12. The method according to Claim 11, the audio representations being one or more Mel spectrograms. (utilizing Mel frequency spectrograms for predicting frames 0054, calculating the loss difference between prediction and reference/actual as in fig. 2 and using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) Re claim 13, Elias teaches 13. The method according to Claim 11, the method further comprising converting the audio representations into an output waveform. (synthesized output by predicting frames 0054, calculating the loss difference between prediction and reference/actual as in fig. 2 and using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) Re claim 14, Elias teaches 14. The method according to Claim 13, where the output waveform is a time-domain signal. (using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) Re claim 20, Elias teaches 20. The method of Claim 15, wherein the AI duration model was previously trained with training data including a plurality of phonemes derived from a string of text, corresponding actual frame durations for each of the phonemes, and a target output speech time duration, wherein the training of the AI duration model included: (predicting frames 0054, calculating the loss difference between prediction and reference/actual as in fig. 2 and using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) calculating a loss with a loss function to quantify a difference of at least the predicted frame durations and the actual frame durations, and using the loss to train the AI duration model by adjusting parameters of the AI duration model that are used to generate the predicted frame durations. (based on predicting frames 0054, calculating the loss difference between prediction and reference/actual as in fig. 2 and using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) However, while Elias teaches neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: the AI duration model masking actual frame durations for a subset of the plurality of phonemes; (KIM phoneme duration model specific to masked speech 0030 and 0043 with fig. 2) the AI duration model generating predicted frame durations for the masked actual frame durations of the subset of the plurality of phonemes; and (KIM prediction output based on phoneme duration model specific to masked speech 0030 and 0043 with fig. 2) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias to incorporate the above claim limitations as taught by KIM to allow for simple substation of one known element for another to yield predictable results, such as the masked speech in place of Elias speech, to allow the model to learn to infer durations from contextual information rather than relying on direct, unmasked acoustic cues, leading to more natural and controllable speech synthesis, to provide this option for more exact and smooth synthesis outputs. Claims 2-5 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20240161730 A1 Elias; Isaac et al. (hereinafter Elias) in view of US 20250149023 A1 KIM; Tae Woo et al. (hereinafter KIM) and further in view of US 11704507 B1 Fantinuoli; Claudio (hereinafter Fantinuoli) Re claims 2, while Elias teaches synthesizing and synchronizing speech features to output speech, as well as neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: 2. The method according to Claim 1, where the target output speech time duration is approximately equal to a time duration for an initial speech. (Fantinuoli approximating the input speech to the output speech by constraining speech features such as word and time ratios, speed, and latency thresholds between conversions e.g. translations col 10 lines 26 to col 11 line 16) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias in view of KIM to incorporate the above claim limitations as taught by Fantinuoli to allow for combining prior art elements according to a known methods to yield predictable results, such as using latency reduction during transformations from one speech to another such as in TTS modeling, wherein latency is reduced from source or target by synchronizing durations and rate of speech, which improves awareness of when pauses or removal of stop-words/pauses are needed to synch output speech to a more natural sounding flow of conversation or general speaking style dependent on the user or scenario, in combination with compression for non-complex synchronization. Re claims 3, while Elias teaches synthesizing and synchronizing speech features to output speech, as well as neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: 3. The method according to Claim 2, where the string of text and the speech generated from the string of text are of a first language, and the initial speech is of a second language, such that the target output speech time duration for the speech of the first language is approximately equal to the time duration for the initial speech of the second language. (Fantinuoli approximating the input speech to the output speech by constraining speech features such as word and time ratios, speed, and latency thresholds between conversions e.g. translations col 10 lines 26 to col 11 line 16) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias in view of KIM to incorporate the above claim limitations as taught by Fantinuoli to allow for combining prior art elements according to a known methods to yield predictable results, such as using latency reduction during translations between languages, from one speech to another such as in TTS modeling, wherein latency is reduced from source or target by synchronizing durations and rate of speech, which improves awareness of when pauses or removal of stop-words/pauses are needed to synch output speech to a more natural sounding flow of conversation or general speaking style dependent on the user or scenario, in combination with compression for non-complex synchronization. Re claims 4, while Elias teaches synthesizing and synchronizing speech features to output speech, as well as neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: 4. The method according to Claim 1, where the target output speech time duration is greater than a time duration for an initial speech, where speech at the target output speech time duration is a speed-up version of the initial speech. (Fantinuoli approximating the input speech to the output speech by constraining speech features such as word and time ratios, speed, and latency thresholds between conversions e.g. translations col 10 lines 26 to col 11 line 16) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias in view of KIM to incorporate the above claim limitations as taught by Fantinuoli to allow for combining prior art elements according to a known methods to yield predictable results, such as using latency reduction and speed increase/slow-down via thresholds, during transformations from one speech to another such as in TTS modeling, wherein latency is reduced from source or target by synchronizing durations and rate of speech, which improves awareness of when pauses or removal of stop-words/pauses are needed to synch output speech to a more natural sounding flow of conversation or general speaking style dependent on the user or scenario, in combination with compression for non-complex synchronization. Re claims 5, while Elias teaches synthesizing and synchronizing speech features to output speech, as well as neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: 5. The method according to Claim 1, where the target output speech time duration is less than a time duration for an initial speech, where speech at the target output speech time duration is a slowed-down version of the initial speech. (Fantinuoli approximating the input speech to the output speech by constraining speech features such as word and time ratios, speed, and latency thresholds between conversions e.g. translations col 10 lines 26 to col 11 line 16) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias in view of KIM to incorporate the above claim limitations as taught by Fantinuoli to allow for combining prior art elements according to a known methods to yield predictable results, such as using latency reduction and speed increase/slow-down via thresholds, during transformations from one speech to another such as in TTS modeling, wherein latency is reduced from source or target by synchronizing durations and rate of speech, which improves awareness of when pauses or removal of stop-words/pauses are needed to synch output speech to a more natural sounding flow of conversation or general speaking style dependent on the user or scenario, in combination with compression for non-complex synchronization. Claims 7-9 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20240161730 A1 Elias; Isaac et al. (hereinafter Elias) in view of US 20250149023 A1 KIM; Tae Woo et al. (hereinafter KIM) and further in view of US 20240038212 A1 Shih; Kevin et al. (hereinafter Shih) Re claim 7, while Elias teaches synthesizing and synchronizing speech features to output speech, as well as neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: 7. The method according to Claim 1, where the AI model masks the actual frame durations for the subset of the plurality of phonemes non-sequentially. (Shih mask phoneme duration processing 0047 with non-sequential or sequential task performance, random sampling, and K-means inclusive of mean-square 0098 and 0101) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias in view of KIM to incorporate the above claim limitations as taught by Shih to allow for combining prior art elements according to a known methods to yield predictable results, such as using well-known mathematical or processing techniques for significant improvements in training efficiency, phonetic alignment, and generation quality, particularly in Non-Autoregressive, regressive, or Text-to-Speech systems without requiring strict, pre-aligned, or frame-level-only operations, and instead rely on semantic or cluster-based information when prioritized over duration when alignment is not needed or already took place if further improvement is achievable for signal quality e.g. more natural sounding. Re claim 8, while Elias teaches synthesizing and synchronizing speech features to output speech, as well as neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: 8. The method according to Claim 1, where the AI model masks the actual frame durations for the subset of the plurality of phonemes randomly. (Shih mask phoneme duration processing 0047 with non-sequential or sequential task performance, random sampling, and K-means inclusive of mean-square 0098 and 0101) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias in view of KIM to incorporate the above claim limitations as taught by Shih to allow for combining prior art elements according to a known methods to yield predictable results, such as using well-known mathematical or processing techniques for significant improvements in training efficiency, phonetic alignment, and generation quality, particularly in Non-Autoregressive, regressive, or Text-to-Speech systems without requiring strict, pre-aligned, or frame-level-only operations, and instead rely on semantic or cluster-based information when prioritized over duration when alignment is not needed or already took place if further improvement is achievable for signal quality e.g. more natural sounding. Re claim 9, while Elias teaches synthesizing and synchronizing speech features to output speech, as well as neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: 9. The method according to Claim 1, where the loss is calculated using mean-squared error loss. (Shih mask phoneme duration processing 0047 with non-sequential or sequential task performance, random sampling, and K-means inclusive of mean-square 0098 and 0101) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias in view of KIM to incorporate the above claim limitations as taught by Shih to allow for combining prior art elements according to a known methods to yield predictable results, such as using well-known mathematical or processing techniques for significant improvements in training efficiency, phonetic alignment, and generation quality, particularly in Non-Autoregressive, regressive, or Text-to-Speech systems without requiring strict, pre-aligned, or frame-level-only operations, and instead rely on semantic or cluster-based information when prioritized over duration when alignment is not needed or already took place if further improvement is achievable for signal quality e.g. more natural sounding. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim 15 is rejected under 35 U.S.C. 102(a)(1) as being anticipated US 20240161730 A1 Elias; Isaac et al. (hereinafter Elias). Re claim 15, Elias teaches 15. A method for using an AI duration model to control the generation of speech utterances by a text-to-speech computing system when converting text into speech, the method comprising: (neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) obtaining an AI duration model trained to generate phonemes and frame durations for the phonemes based on inputs comprising text to be converted into speech and a target output speech time duration; (using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) identifying the text to be converted into speech; (TTS… using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) identifying the target output speech time duration; (TTS output thereof as speech… using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) providing the text and the target output speech time duration to the AI duration model, wherein the AI duration model tokenizes the text into a plurality of phonemes and predicts a frame duration for each phoneme in the plurality of phonemes based on the target output speech time duration, such that the summation of the frame durations for the plurality of phonemes is approximately equal to the target output speech time duration; and (predicting frames 0054, calculating the loss difference between prediction and reference/actual as in fig. 2 and using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) generating output based on the phonemes and predicted frame duration for each phoneme. (synthesized output, predicting frames 0054, calculating the loss difference between prediction and reference/actual as in fig. 2 and using a method such as cross-entropy 0051 using phonemes from text sequence 0004 and 0010 such as by parsing 0049 to supply training data, based on a neural network modeling with TTS to produce synthesized speech in the time-domain 0064 with fig. 1) …the output including an audio representation of the text, the audio representation comprising a time-frequency representation having a temporal structure defined by the predicted frame durations; and (input duration equals output duration 0060-0061… including in Mel time-freq input equals output 0052 as well i.e. time vs freq, wherein a time-frequency representation and temporal structure are directed to Mel) converting the audio representation into an output waveform comprising a time- domain signal having a duration corresponding to the target output speech time duration. (input duration equals output duration 0060-0061 including in Mel time-freq input equals output 0052 as well i.e. time vs freq, wherein a time-frequency representation and temporal structure are directed to Mel, and further thereof expressly converted into time domaon 0064) Claim 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over US 20240161730 A1 Elias; Isaac et al. (hereinafter Elias) in view of US 11704507 B1 Fantinuoli; Claudio (hereinafter Fantinuoli). Re claim 16, while Elias teaches synthesizing and synchronizing speech features to output speech, as well as neural network model driven duration training for predicted frames based on loss calculations for phonemes in frames, it fails to teach: 16. The method according to Claim 15, wherein the method includes using the AI duration model to translate a first speech segment in a first language having a first speech segment duration into a second speech segment in a second language having a second speech segment duration, such that the target speech time duration is approximately equal to the first speech segment duration. (Fantinuoli approximating the input speech to the output speech by constraining speech features such as word and time ratios, speed, and latency thresholds between conversions e.g. translations col 10 lines 26 to col 11 line 16) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the system of Elias to incorporate the above claim limitations as taught by Fantinuoli to allow for combining prior art elements according to a known methods to yield predictable results, such as using latency reduction during translations between languages, from one speech to another such as in TTS modeling, wherein latency is reduced from source or target by synchronizing durations and rate of speech, which improves awareness of when pauses or removal of stop-words/pauses are needed to synch output speech to a more natural sounding flow of conversation or general speaking style dependent on the user or scenario, in combination with compression for non-complex synchronization. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 10741169 B1 Trueba; Jaime Lorenzo et al. Phoneme analysis Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL COLUCCI whose telephone number is (571)270-1847. The examiner can normally be reached on M-F 9 AM - 7 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL COLUCCI/Primary Examiner, Art Unit 2655 (571)-270-1847 Examiner FAX: (571)-270-2847 Michael.Colucci@uspto.gov
Read full office action

Prosecution Timeline

Jul 30, 2024
Application Filed
Feb 18, 2026
Non-Final Rejection mailed — §102, §103
Apr 27, 2026
Examiner Interview Summary
Apr 27, 2026
Applicant Interview (Telephonic)
Apr 30, 2026
Response Filed
Jun 04, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706085
SYSTEMS AND METHODS FOR CONTINUAL LEARNING FOR END TO-END AUTOMATIC SPEECH RECOGNITION
2y 5m to grant Granted Aug 11, 2026
Patent 12688851
VOICE BASED ACTIVATION DETECTION
2y 4m to grant Granted Jul 21, 2026
Patent 12682896
SYSTEM AND METHOD FOR THE GENERATION OF WORKLISTS FROM INTERACTION RECORDINGS
2y 7m to grant Granted Jul 14, 2026
Patent 12664976
QUERY REPLAY FOR PERSONALIZED RESPONSES IN AN LLM POWERED ASSISTANT
2y 6m to grant Granted Jun 23, 2026
Patent 12664977
FLY PARAMETER COMPRESSION AND DECOMPRESSION TO FACILITATE FORWARD AND/OR BACK PROPAGATION AT CLIENTS DURING FEDERATED LEARNING
2y 1m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
76%
Grant Probability
91%
With Interview (+15.2%)
3y 1m (~1y 1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 1009 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month