Prosecution Insights
Last updated: October 02, 2026
Application No. 18/828,759

PITCH CONTROL ALGORITHM

Final Rejection §102§103
Filed
Sep 09, 2024
Examiner
WEAVER, ADAM MICHAEL
Art Unit
2658
Tech Center
2600 — Communications
Assignee
Sony Group Corporation
OA Round
2 (Final)
88%
Grant Probability
Favorable
3-4
OA Rounds
5m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 88% — above average
88%
Career Allowance Rate
14 granted / 16 resolved
+25.5% vs TC avg
Strong +31% interview lift
Without
With
+31.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
21 currently pending
Career history
53
Total Applications
across all art units

Statute-Specific Performance

§101
29.6%
-10.4% vs TC avg
§103
52.6%
+12.6% vs TC avg
§102
14.6%
-25.4% vs TC avg
§112
1.6%
-38.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 16 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Amendment filed 06/10/2026 has been entered. Objections to the Specification and Drawings have been withdrawn. Claims 1-16 remain pending in the application. Terminal Disclaimer The terminal disclaimer filed on 06/10/2026 disclaiming the terminal portion of any patent granted on this application which would extend beyond the expiration date of 18/828,759 has been reviewed and was approved on 06/17/2026. The terminal disclaimer has been recorded. Response to Arguments Applicant’s arguments, filed 6/10/2026, with respect to the Double Patenting rejection, have been fully considered. Therefore, this rejection has been withdrawn. With respect to the 35 U.S.C. 101 abstract idea rejection, see pages 8-12, the Applicant’s argument has been fully considered and is persuasive. Therefore, this rejection has been withdrawn. With respect to the 35 U.S.C. 102 rejection, on pages 12-14, of claims 13-14 under Calapodescu et al. (US Patent Application Publication No. 2023/0215421), hereinafter referred to as Calapodescu, the Applicant asserts that Calapodescu fails to disclose or suggest “send[ing] the input Htext through a series of one-dimensional convolution layers to a fully connected layer to produce an output in which, for each of a plurality of tokens, an estimate of an associated token-wise pitch is provided,” and therefore Calapodescu fails to disclose the features of amended claim 13 because Calapodescu relies on an iterative process to sequentially predict prosodic attributes for each token rather than sending an input text through a series of 1D convolutional layers and a fully connected layer to produce an output including an estimate of an associated token-wise pitch for each of a plurality of tokens. The Applicant further asserts that Calapodescu does not “convert a representation incorporating the estimate of the associated token-wise pitch into a speech representation to modify intermediate representations of phonemes for generation of an audio output.” In response to Applicant’s argument that Calapodescu does not disclose or suggest “send[ing] the input Htext through a series of one-dimensional convolution layers to a fully connected layer to produce an output in which, for each of a plurality of tokens, an estimate of an associated token-wise pitch is provided”, Calapodescu Fig. 4 very clearly illustrates taking in a text with reference character 406, embedding it, and then sending it to prosodic attribute predictors 1 and 2, reference characters 404a and 404b respectively, which contain multiple one-dimensional convolutional layers before a full connected layer, to then predictor prosodic attributes, in this case pitch. This figure succinctly discloses all of the claimed limitations. Calapodescu also discloses “convert a representation incorporating the estimate of the associated token-wise pitch into a speech representation to modify intermediate representations of phonemes for generation of an audio output” through Calapodescu para [0069]: “The decoder 420 embodied in an FFTr stack 420 and an FC layer 422 is trained to generate speech representations of the input text that include prosody associated with the annotations 408,” and Calapodescu para [0070]: “An output of the model 400 generated via the decoder 420, such as a speech signal (e.g., a spectrogram), can be provided to a speech synthesizer 430, such as but not limited to a decoder (e.g., with a vocoder), for generating speech from the speech signal.” Both of these excerpts show that a speech representation is created based upon the estimated pitches, and that the pitches are modified to generate a final audio output through the prosody annotations. Thus, Applicant’s arguments regarding the 35 U.S.C. 102 rejection are not persuasive. With respect to the 35 U.S.C. 103 rejection, on pages 14-15, of claims 1-12 and 15-16 under Calapodescu in view of Shih et al. (US Patent No. 12,488,778), hereinafter referred to as Shih, the Applicant asserts that the cited art, alone or combined, fails to disclose or suggest “group the phoneme-level pitch values into N bins, wherein each bin of the N bins has a same number of phoneme-level pitch values as other bins of the N bins.” They further assert that nothing in Shih discloses or suggests that each bin contains the same number of speech characteristics, let alone the same number of phoneme-level pitch values. Applicant’s arguments with respect to the rejection(s) of claim(s) 1-12 and 15-16 under Calapodescu, in view of Shih have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Calapodescu, in view of Shih, and further in view of Oplustil Gallegos. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 13-14 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Calapodescu et al. (US Patent Application Publication No. 2023/0215421), hereinafter referred to as Calapodescu. Regarding claim 13, Calapodescu discloses a device, comprising: a non-transitory computer-readable medium text at least in part by by causing the at least one processor system to send (Calapodescu Fig. 4 reference character 404a shows the prosodic attribute predictor, which predicts pitch and “An embedding layer 411 embeds tokens in the input text 406 and provides the embedded tokens to the encoder," Calapodescu para [0069])text through a series of one-dimensional convolution layers (Calapodescu Fig. 4 reference characters 410a and 412a are two one-dimensional convolutional layers in series) a plurality of tokens, an estimate of an associated token-wise pitch is provided (Calapodescu Fig. 4 reference character 416a shows the fully connected layer, where 410a and 412a are concatenated, then the output is sent to reference character 418, another one-dimensional convolutional layer, where the output then is the predicted pitch); and convert a representation incorporating the estimate of the associated token-wise pitch into a speech representation to modify intermediate representations of phonemes for generation of an audio output (“The decoder 420 embodied in an FFTr stack 420 and an FC layer 422 is trained to generate speech representations of the input text that include prosody associated with the annotations 408,” Calapodescu para [0069], prosody annotations 408 are the modifications to the prosody that are generated in the output, and “An output of the model 400 generated via the decoder 420, such as a speech signal (e.g., a spectrogram), can be provided to a speech synthesizer 430, such as but not limited to a decoder (e.g., with a vocoder), for generating speech from the speech signal,” Calapodescu para [0070]). Regarding claim 14, Calapodescu discloses all of the limitations of claim 13. Calapodescu further discloses wherein the one-dimensional convolutional layers are concatenated (Calapodescu Fig. 4 reference character 416a shows the fully connected layer, where 410a and 412a are concatenated). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-12 and 15-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Calapodescu, in view of Shih et al. (US Patent No. 12,488,778), hereinafter referred to as Shih, and further in view of Oplustil Gallegos et al. (US Patent Application Publication No. 2024/0087558), hereinafter referred to as Oplustil Gallegos. Regarding claim 1, Calapodescu discloses an apparatus comprising: at least one processor system configured to: receive text ("The encoder is trained to process an input text 406 including annotations 408 indicating prosodic features in the input text," Calapodescu para [0067]); and convert a latent representation of the text incorporating the respective encoded pitch target into a speech representation to modify intermediate representations of phonemes for generation of an audio output (“The decoder 420 embodied in an FFTr stack 420 and an FC layer 422 is trained to generate speech representations of the input text that include prosody associated with the annotations 408,” Calapodescu para [0069], prosody annotations 408 are the modifications to the prosody that are generated in the output, and “An output of the model 400 generated via the decoder 420, such as a speech signal (e.g., a spectrogram), can be provided to a speech synthesizer 430, such as but not limited to a decoder (e.g., with a vocoder), for generating speech from the speech signal,” Calapodescu para [0070]). However, Calapodescu fails to disclose convert the text toa plurality of phonemes, wherein each phoneme of the plurality of phonemes has wherein each bin of the N bins has a same number of phoneme-level pitch values as other bins of the N bins; to obtain a respective encoded pitch target for that respective bin such that the phoneme-level pitch values in the respective bin are represented by the respectiveencoded pitch target. Shih teaches a method to implement a generative text-to-speech model. Shih teaches convert the text to plural phonemes, each having a respective phoneme-level pitch value (Shih Fig. 4 reference character 410 partitions text into phonemes 412); to obtain a respective encoded pitch target for that respective bin such that the phoneme-level pitch values in the respective bin are represented by the respectiveencoded pitch target (Shih Fig. 4 reference character 420 associated a phoneme feature vector 422 with the phonemes); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Calapodescu’s disclosure of a neural text-to-speech model with prosody control by including Shih’s teaching of converting the text to phonemes, grouping the phonemes together, and encoding/embedding them into a vector. Converting a textual element into its individual phonemes is a well-known technique within the text-to-speech system art, as it allows for individual acoustic parameters to be discovered, analyzed, and modified. Grouping the phonemes together allows for like phonemes with similar initial acoustic parameters to be analyzed together, thereby simplifying the amount of data that needs to be analyzed. Encoding or embedding the phonemes and groups of phonemes into vector format allows for numerical representation of the data, which further simplifies it and allows for input into neural models. This also allows for easier increase of dimensionality of the data, as neural networks tend to increase the dimensionality of data since multi-dimensionality of data provides additional degrees of freedom to fit complex functions (Shih col. 3 lines 28-33). Oplustil Gallegos teaches a method and system for modifying speech generated by a text-to-speech synthesizer. Oplustil Gallegos teaches group the phoneme-level pitch values into N bins, wherein each bin of the N bins has a same number of phoneme-level pitch values as other bins of the N bins (“This vector can then be simplified or “coarse grained” by binning/scaling/grouping the pitch values into N bins or groups according to pitch/frequency in 281. In an example, N is three which provides integer values {0, 1, 2}. These bins may be obtained by calculating the min and max for the entire dataset in 271 and 273 and splitting that range into three equally sized bins. Using these pitch/frequency bins the average phoneme pitch vector is turned into an integer vector containing the values 0 (lowest frequency bin) to 2 (highest frequency bin)”, Oplustil Gallegos para [0271]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Calapodescu’s disclosure of a neural text-to-speech model with prosody control by including Oplustil Gallegos’ teaching of utilizing binning to group pitch values into bins according to pitch. Binning is a common and well-known technique within data processing utilized to simplify data and improve the reliability of the data and its analysis. This would have been an obvious inclusion. Regarding claim 2, Calapodescu, in view of Shih and further in view of Oplustil Gallegos, discloses all of the limitations of claim 1. Calapodescu further discloses wherein the processor system is configured to: predict pitch from an input Htext, which is identical to decoder input minus any pitch information ("An embedding layer 411 embeds tokens in the input text 406 and provides the embedded tokens to the encoder," Calapodescu para [0069] and Calapodescu Fig. 4 reference character 404a shows the prosodic attribute predictor, which predicts pitch). Regarding claim 3, Calapodescu, in view of Shih and further in view of Oplustil Gallegos, discloses all of the limitations of claim 2. Calapodescu further discloses wherein the processor system is configured to: send the input Htext to a first one-dimensional convolution layer in series with a second one-dimensional convolution layer (Calapodescu Fig. 4 reference characters 410a and 412a are two one-dimensional convolutional layers in series). Regarding claim 4, Calapodescu, in view of Shih and further in view of Oplustil Gallegos, discloses all of the limitations of claim 3. Calapodescu further discloses wherein the one-dimensional convolutional layers are concatenated (Calapodescu Fig. 4 reference character 416a shows the fully connected layer, where 410a and 412a are concatenated). Regarding claim 5, Calapodescu, in view of Shih and further in view of Oplustil Gallegos, discloses all of the limitations of claim 4. Calapodescu further discloses wherein the processor system is configured to send output of the one-dimensional convolutional layers to a fully connected layer to produce an output in which, for each of plural tokens, an estimate of an associated token-wise pitch is provided (Calapodescu Fig. 4 reference character 416a shows the fully connected layer, where 410a and 412a are concatenated, then the output is sent to reference character 418, another one-dimensional convolutional layer, where the output then is the predicted pitch). Regarding claim 6, Calapodescu, in view of Shih and further in view of Oplustil Gallegos, discloses all of the limitations of claim 1. However, Calapodescu fails to disclose wherein unvoiced speech is represented with an all-0 vector indicating the unvoiced speech has no pitch. Shih teaches wherein unvoiced speech is represented with an all-0 vector indicating the unvoiced speech has no pitch ("Gap filling 204 may first identify unvoiced regions (gaps) as regions where the pitch (or energy) data is unavailable or shows small values below a set threshold," Shih col. 7 lines 41-43). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Calapodescu’s disclosure of a neural text-to-speech model with prosody control by including Shih’s teaching of indicating where unvoiced or voiceless phonemes are present in the speech. If phonemes are unvoiced or voiceless, i.e. they have no pitch, energy, or any other acoustic information attributed to them, it would be obvious to represent this value based on zero or a vector of zeros. As to claims 7-12, method claims 7-12 and system claims 1-6 are related as system and method of using same, with each claimed element’s function corresponding to the system step, respectively. Accordingly, claims 7-12 are similarly rejected under the same rationale as applied above with respect to the system claims. Regarding claim 15, Calapodescu discloses all of the limitations of claim 13. However, Calapodescu fails to disclose identify plural phonemes, each having a respective phoneme-level pitch value; group the phoneme-level pitch values into N bins, each bin having a same number of phoneme-level pitch values as other bins; and encode each bin with a respective vector such that the phoneme-level pitch values in a respective bin are represented by the respective vector. Shih teaches identify plural phonemes, each having a respective phoneme-level pitch value; and encode each bin with a respective vector such that the phoneme-level pitch values in a respective bin are represented by the respective vector (Shih Fig. 4 reference character 420 associated a phoneme feature vector 422 with the phonemes). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Calapodescu’s disclosure of a neural text-to-speech model with prosody control by including Shih’s teaching of converting the text to phonemes, grouping the phonemes together, and encoding/embedding them into a vector. Converting a textual element into its individual phonemes is a well-known technique within the text-to-speech system art, as it allows for individual acoustic parameters to be discovered, analyzed, and modified. Grouping the phonemes together allows for like phonemes with similar initial acoustic parameters to be analyzed together, thereby simplifying the amount of data that needs to be analyzed. Encoding or embedding the phonemes and groups of phonemes into vector format allows for numerical representation of the data, which further simplifies it and allows for input into neural models. This also allows for easier increase of dimensionality of the data, as neural networks tend to increase the dimensionality of data since multi-dimensionality of data provides additional degrees of freedom to fit complex functions (Shih col. 3 lines 28-33). Oplustil Gallegos teaches group the phoneme-level pitch values into N bins, each bin having a same number of phoneme-level pitch values as other bins (“This vector can then be simplified or “coarse grained” by binning/scaling/grouping the pitch values into N bins or groups according to pitch/frequency in 281. In an example, N is three which provides integer values {0, 1, 2}. These bins may be obtained by calculating the min and max for the entire dataset in 271 and 273 and splitting that range into three equally sized bins. Using these pitch/frequency bins the average phoneme pitch vector is turned into an integer vector containing the values 0 (lowest frequency bin) to 2 (highest frequency bin)”, Oplustil Gallegos para [0271]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Calapodescu’s disclosure of a neural text-to-speech model with prosody control by including Oplustil Gallegos’ teaching of utilizing binning to group pitch values into bins according to pitch. Binning is a common and well-known technique within data processing utilized to simplify data and improve the reliability of the data and its analysis. This would have been an obvious inclusion. Regarding claim 16, Calapodescu discloses all of the limitations of claim 13. However, Calapodescu fails to disclose wherein the instructions are executable to: represent unvoiced speech is represented with an all-0 vector indicating the unvoiced speech has no pitch. Shih teaches wherein the instructions are executable to: represent unvoiced speech is represented with an all-0 vector indicating the unvoiced speech has no pitch ("Gap filling 204 may first identify unvoiced regions (gaps) as regions where the pitch (or energy) data is unavailable or shows small values below a set threshold," Shih col. 7 lines 41-43). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Calapodescu’s disclosure of a neural text-to-speech model with prosody control by including Shih’s teaching of indicating where unvoiced or voiceless phonemes are present in the speech. If phonemes are unvoiced or voiceless, i.e. they have no pitch, energy, or any other acoustic information attributed to them, it would be obvious to represent this value based on zero or a vector of zeros. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADAM MICHAEL WEAVER whose telephone number is (571)272-7062. The examiner can normally be reached Monday-Friday, 8AM-5PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ADAM MICHAEL WEAVER/Examiner, Art Unit 2658 /RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Sep 09, 2024
Application Filed
Mar 11, 2026
Non-Final Rejection mailed — §102, §103
Jun 10, 2026
Response Filed
Sep 02, 2026
Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12664978
FEDERATED KNOWLEDGE DISTILLATION ON AN ENCODER OF A GLOBAL ASR MODEL AND/OR AN ENCODER OF A CLIENT ASR MODEL
3y 6m to grant Granted Jun 23, 2026
Patent 12657219
INFORMATION PROCESSING DEVICE, COMPUTER PROGRAM PRODUCT, AND INFORMATION PROCESSING METHOD
2y 3m to grant Granted Jun 16, 2026
Patent 12651117
METHODS AND SYSTEMS FOR VERIFICATION OF PLANT PROCEDURES' COMPLIANCE TO WRITING MANUALS
4y 0m to grant Granted Jun 09, 2026
Patent 12651266
SYSTEMS AND METHODS FOR RANKING CALL INTENT PROBABILITY
2y 3m to grant Granted Jun 09, 2026
Patent 12639355
IDENTIFYING HALLUCINATIONS IN LARGE LANGUAGE MODEL OUTPUT
2y 9m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
88%
Grant Probability
99%
With Interview (+31.2%)
2y 6m (~5m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 16 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month