Prosecution Insights
Last updated: October 02, 2026
Application No. 17/936,101

SYLLABLE-BASED TEXT CONVERSION FOR PRONUNCIATION HELP

Final Rejection §101§103§112
Filed
Sep 28, 2022
Examiner
CHUNG, DANIEL WONSUK
Art Unit
2659
Tech Center
2600 — Communications
Assignee
International Business Machines Corporation
OA Round
2 (Final)
59%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 59% of resolved cases
59%
Career Allowance Rate
32 granted / 54 resolved
-2.7% vs TC avg
Strong +31% interview lift
Without
With
+31.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 12m
Avg Prosecution
22 currently pending
Career history
84
Total Applications
across all art units

Statute-Specific Performance

§101
24.7%
-15.3% vs TC avg
§103
50.6%
+10.6% vs TC avg
§102
17.6%
-22.4% vs TC avg
§112
6.1%
-33.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 54 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This communication is in response to the Amendments and Arguments filed on 4/9/2026. Claims 21-40 are pending and have been examined. All previous objections / rejections not mentioned in this Office Action have been withdrawn by the examiner. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendments Applicant has cancelled all originally filed claims and has introduced 20 new claims 21-40. Regarding the Applicant’s arguments for the rejections under 35 U.S.C. § 101, examiner broadly interprets the newly introduced claims and maintains the abstract idea rejection. First, the human mind can perform the steps recited in the claims. Moreover, the claim recitations are high level and describe the use of generic computing components to carry out the abstract idea. The claims do not recite specific technological improvement to produce target language syllables from input text in a first language, nor does it describe how the claimed operations improve the functioning of the computer or other technology in a concrete, technical way. Regarding the Applicant’s arguments for the rejections under 35 U.S.C. § 103, applicant has cancelled claims 1-20 and added claims 21-40. The added limitations raise new grounds for rejection. Hence, new references have been applied. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claim 21, 39, and 40 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Specifically, the as filed disclosure does not teach a “training a first machine learning model and a second machine learning model” and “second machine learning model, in response, produces and compares audio tokens representing the first set and the input text”. The specification discloses a second machine learning model to “analyze embeddings representing the one or more spectrograms for the input text and the one or more spectrograms for text of the target language”. Spec. P0020. The as-filed disclosure does not teach the training of the second machine learning model or the second machine learning model producing audio tokens. The specification teaches a specific “audio encoder” that generates audio from text that is trained by processes shown in Fig. 7 and Fig. 8. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claim 39 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claims do not fall within at least one of the four categories of patent eligible subject matter because the limitation of “computer program product”, “computer-readable storage media”, and “program instructions stored on the one or more computer-readable storage media to perform” as drafted, is not directed to a statutory subject matter. The claimed model does not fall within at least one of the four categories of patent eligible subject matter recited in 35 U.S.C. 101 (process, machine, manufacture, or composition of matter). The model can be interpreted as software per se. Claims are not patent eligible if they are not directed to any of the statutory categories and are products that do not have a physical or tangible form. Products without physical or tangible form are information or a computer program per se when claimed as a product without any structural recitations. Therefore, the claim is not patent eligible. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 21, 39 and 40 the limitations of “training a first machine learning model and a second machine learning model in conjunction via: providing input text in a first language and an indicator of a target language to the first machine learning model such that the first machine learning model, in response, produces as output a first set of textual syllable embeddings in the target language that is different from the first language, the first set most closely matching a pronunciation of the input text; providing the first set and the input text to a second machine learning model such that the second machine learning model, in response, produces and compares audio tokens representing the first set and the input text, wherein the audio tokens represent syllable-divided spectrograms corresponding to the first set; and updating the first machine learning model based on a reward determined via the comparing of the audio tokens”, as drafted, are processes that, under broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components. More specifically, the mental process of a human reading text and thinking of text in another language that would pronounce the read text, comparing the phonics of the read text and the text in another language, and adjusting the though process of converting read text into another language text. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the --Mental Processes-- grouping of abstract ideas. Accordingly, the claims recite an abstract idea. This judicial exception is not integrated into a practical application because the recitation of a computer system in claim 40, reads to generalized computer components, based upon the claim interpretation wherein the structure is interpreted using P0038-P0040 in the specification. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are directed to an abstract idea. The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract idea into a practical application, the additional element of using generalized computer components to read text and think of text in another language that would pronounce the read text, compare the phonics of the read text and the text in another language, and adjusting the though process of converting read text into another language text amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are not patent eligible. With respect to claim 22, the claim recites “providing a new input text in the first language to the trained first machine learning model such that, in response, the trained first machine learning model produces as output a second set of textual syllable embeddings in the target language, the second set most closely matching a pronunciation of the new input set; and presenting the second set via a computer”, which reads on a human reading text and thinking of text in another language that would pronounce the read text. No additional limitations are present. With respect to claim 23, the claim recites “wherein the target language comprises multiple languages that are each different from the first language; and wherein the method further comprises: receiving a selection of a second language from amongst the multiple languages of the target language; providing a new input text in the first language and an indicator of the second language to the trained first machine learning model such that, in response, the trained first machine learning model produces as output a second set of textual syllable embeddings in the second language, the second set most closely matching a pronunciation of the new input set; and presenting the second set via a computer”, which reads on a human thinking of text in another selected language that would pronounce the read text. No additional limitations are present. With respect to claim 24, the claim recites “wherein the reward that is used to update the first machine learning model is further determined via providing the textual syllable embeddings in the target language into a machine learning classifier that, in response, produces a classification of a predicted language of the textual syllable embeddings”, which reads on a human thinking of text in another selected language that is classified as another language. No additional limitations are present. With respect to claim 25, the claim recites “wherein the classifier includes an output layer that assigns decimal probabilities to each class in a multi-class problem, each class of the multi-class problem representing a different unique language that is separate from the first language”, which reads on a human identifying predicted language text belonging in another language using probability metric. No additional limitations are present. With respect to claim 26, the claim recites “wherein the target language comprises multiple languages that are each different from the first language; wherein the classifier includes an output layer that assigns decimal probabilities to classes of a multi-class problem; and wherein the classes include a separate class for each of the multiple languages”, which reads on a human thinking of text in multiple languages. No additional limitations are present. With respect to claim 27, the claim recites “wherein the classes further include an additional class for a mixed language”, which reads on a human thinking of text in mixed languages. No additional limitations are present. With respect to claim 28, the claim recites “wherein the second machine learning model comprises a token embedding layer and a transformer, the token embedding layer generating embeddings that are input into the transformer”, which reads on a human making mathematical calculations of phonetic differences using a token embedding layer and transformer. No additional limitations are present. With respect to claim 29, the claim recites “wherein the token embedding layer generates final embeddings that include token embeddings, modal-type embeddings, position embeddings, and language-category embeddings”, which reads on a human performing calculations to generate a final embedding. No additional limitations are present. With respect to claim 30, the claim recites “wherein the token embedding layer receives pairs of: sets of textual syllable tokens and sets of audio syllable tokens”, which reads on a human performing calculations utilizing tokens when comparing phonetic differences. No additional limitations are present. With respect to claim 31, the claim recites “wherein the token embedding layer further receives separator tokens which separate a respective set of the textual syllable tokens and a paired respective set of the audio syllable tokens”, which reads on a human utilizing separator token in calculations. No additional limitations are present. With respect to claim 32, the claim recites “further comprising training the second machine learning model in conjunction with an encoder, wherein the encoder receives a spectrogram as an input and produces a spectrogram vector representing the spectrogram as an output”, which reads on a human utilizing representations of spectrogram that illustrates phonetic similarity or differences. No additional limitations are present. With respect to claim 33, the claim recites “wherein the training of the second machine learning model further uses a codebook that stores vectors representing a group of spectrograms, wherein the training of the second machine learning model include performing nearest-neighbor mapping of the stored vectors in the codebook to find a matching vector that is most similar to the spectrogram vector produced via the encoder”, which reads on a human utilizing codebook for spectrogram representation. No additional limitations are present. With respect to claim 34, the claim recites “further comprising training the encoder by obtaining, from the codebook, an index value that represents the matching vector and converting the index value to a one-hot encoding embedding that is a respective one of the audio tokens”, which reads on a human utilizing codebook for spectrogram representation. No additional limitations are present. With respect to claim 35, the claim recites “wherein the index value is selected from a range from N to N + K, where K is a hyper-parameter and is equal to a total number of unduplicated audio syllables and N is also a hyper-parameter and is equal to a total number of all textual syllable embeddings”, which reads on a human utilizing codebook for spectrogram representation. No additional limitations are present. With respect to claim 36, the claim recites “wherein the encoder is trained in conjunction with a decoder that receives the index value as input and that, in response, produces a reconstructed spectrogram as output, the reconstructed spectrogram being intended to match the spectrogram that is input into the encoder”, which reads on a human utilizing spectrogram representation to calculate a spectrogram. No additional limitations are present. With respect to claim 37, the claim recites “further comprising generating a single-syllable spectrogram with a width smaller than a maximum width by filling in blank areas around the spectrogram with color”, which reads on a human utilizing spectrogram representation to calculate a spectrogram. No additional limitations are present. With respect to claim 38, the claim recites “wherein the syllable-divided spectrograms are generated via: calculating a time to pronounce each of multiple syllables in the first language; identifying points of zero amplitude in an audio waveform generated via pronouncing the input text; and dividing an input text spectrogram for the input text into a spectrogram per syllable of the input text, wherein the dividing is based on the calculated time and on the identified points of zero amplitude”, which reads on a human utilizing spectrogram representation to calculate a spectrogram. No additional limitations are present. These claims further do not remedy the judicial exception being integrated into a practical application and further fail to include additional elements that are sufficient to amount to significantly more than the judicial exception. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 21, 22, 23, 32, 39, and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Prasad et al. (U.S. PG Pub No. 20230116268), hereinafter Prasad, in view of Voss et al. (U.S. PG Pub No. 20220199071), hereinafter Voss. Regarding claim 21 Prasad teaches: (Claim 21) A method comprising: (P0012, The method includes receiving an input text in a first script from a user. Each character of the input text is phonetically mapped with a second script corresponding to the second language. The permutations of mapping of each, input character with, each character of the second script is validated and the input text in the first script is transliterated into an output text in the second script.) (Claim 39) A computer program product comprising: one or more computer-readable storage media; and program instructions stored on the one or more computer-readable storage media to perform operations comprising: (P0090, The transliteration engine comprises modules defining computer program instructions, which when executed by the hardware processor, cause the processor to transliterate input text of first language into the output text of second language.) (Claim 40) A computer system comprising: a processor set; one or more computer-readable storage media; and program instructions stored on the one or more computer-readable storage media to cause the processor set to perform operations comprising: (P0087, The computing device comprises at least one processor and a non-transitory, computer-readable storage medium, for example, a memory unit, for storing computer program instructions defined by modules of the transliteration engine.) training a first machine learning model and a second machine learning model in conjunction via: (P0091, The training module comprises an encoder and a decoder. The encoder trains a pre-trained model with the data files and corresponding transliterated text using transfer learning.) providing input text in a first language and an indicator of a target language to the first machine learning model such that the first machine learning model, in response, produces as output a first set of textual syllable embeddings in the target language that is different from the first language, the first set most closely matching a pronunciation of the input text; (P0084, FIG. 4A-4C exemplarily illustrates a graphical representation displayed on a display unit of an electronic device, showing a transliterated text suggestions on a suggestion bar interface, according to an embodiment herein. When a user invokes an input interface, for example, the keyboard interface, through a user application, the transliteration engine displays a predetermined number of transliterated suggestions for the input text default in the suggestion bar positioned in a row of the keyboard interface.; P0090, The transliteration engine … cause the processor to transliterate input text of first language into the output text of second language.; P0091, The inference module executes the inference stage, where the inference module receives a text file as input, processes the input text data through the pretrained language model, and through the pre-trained customized language model, and generates output text data in a second language.; P0130, The decoding comprises the proposed toolkit that provides varying support for three different decoding schemes. … N-best reranking is then accomplished with the toolkit by configuring the decoder to output the N-best joint G-P sequences and employing RNNLM to rerank the N-best joint sequences.) updating the first machine learning model based on a reward determined via the comparing of the audio tokens. (P0021, The goal of Decoding process is to find the most likely pronunciation given the model.; P0075, A training module comprising an encoder and a decoder for training a pre-trained model and decoding an output text of the trained model to generate text comprising characters of the second language.) Prasad does not specifically teach: providing the first set and the input text to a second machine learning model such that the second machine learning model, in response, produces and compares audio tokens representing the first set and the input text, wherein the audio tokens represent syllable-divided spectrograms corresponding to the first set; and updating the first machine learning model based on a reward determined via the comparing of the audio tokens. Voss, however, teaches: providing the first set and the input text to a second machine learning model such that the second machine learning model, in response, produces and compares audio tokens representing the first set and the input text, wherein the audio tokens represent syllable-divided spectrograms corresponding to the first set; and (P0110, Embedding vectors representing encoded audio data in accordance with certain embodiments of the invention can include grapheme probability vectors (as described above) or other representations of the audio signal, including (but not limited to) the raw signal, derived features such as MFCCs, neural network representations (such as the hidden states or memory states of an LSTM), neural embeddings derived from the network, or any other numerical embeddings.; P0112, Template matching can utilize various common mathematical distance metrics to compute the match signal. Generally, when using embedding vectors, geometric distances (e.g., cosine similarity, Euclidean distance, inner product, etc.) can be appropriate.) updating the first machine learning model based on a reward determined via the comparing of the audio tokens. (P0068, Trained using loss functions.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to compare audio tokens representing spectrograms. It would have been obvious to combine the references because comparing audio tokens is a known technique that yields a predictable result of finding similarity in sections of audio data to a template. (Voss P0017) Regarding claim 22 Prasad in view of Voss teach claim 21. Prasad further teaches: providing a new input text in the first language to the trained first machine learning model such that, in response, the trained first machine learning model produces as output a second set of textual syllable embeddings in the target language, the second set most closely matching a pronunciation of the new input set; and (P0048, The embodiments herein transliterate text in any input language (or first language) to text comprising characters of a base language (or a second language) based on pronunciation.; P0090, The transliteration engine … cause the processor to transliterate input text of first language into the output text of second language.; P0075, A training module comprising an encoder and a decoder for training a pre-trained model and decoding an output text of the trained model to generate text comprising characters of the second language.) presenting the second set via a computer. (P0084, FIG. 4A-4C exemplarily illustrates a graphical representation displayed on a display unit of an electronic device, showing a transliterated text suggestions on a suggestion bar interface.) Regarding claim 23 Prasad in view of Voss teach claim 21. Prasad further teaches: wherein the target language comprises multiple languages that are each different from the first language; and wherein the method further comprises: (P0048, Transliterate text in any input language (or first language) to text comprising characters of a base language (or a second language) based on pronunciation.) receiving a selection of a second language from amongst the multiple languages of the target language; (P0084, When a user invokes an input interface, for example, the keyboard interface, through a user application, the transliteration engine 106 displays a predetermined number of transliterated suggestions for the input text default in the suggestion bar positioned in a row of the keyboard interface. According to an embodiment herein, the transliteration engine displays transliterated suggestions and predictions above the keyboard interface. For example, the user input the word ‘sheershak’ in the typing bar and the transliteration engine generates a plurality of suggestions in the suggestion bar.) providing a new input text in the first language and an indicator of the second language to the trained first machine learning model such that, in response, the trained first machine learning model produces as output a second set of textual syllable embeddings in the second language, the second set most closely matching a pronunciation of the new input set; and (P00048, The embodiments herein transliterate text in any input language (or first language) to text comprising characters of a base language (or a second language) based on pronunciation.; P0090, The transliteration engine … cause the processor to transliterate input text of first language into the output text of second language. P0075, A training module comprising an encoder and a decoder for training a pre-trained model and decoding an output text of the trained model to generate text comprising characters of the second language.) presenting the second set via a computer. (P0084, FIG. 4A-4C exemplarily illustrates a graphical representation displayed on a display unit of an electronic device, showing a transliterated text suggestions on a suggestion bar interface.) Regarding claim 32 Prasad in view of Voss teach claim 21. Prasad does not specifically teach: training the second machine learning model in conjunction with an encoder, wherein the encoder receives a spectrogram as an input and produces a spectrogram vector representing the spectrogram as an output. Voss, however, teaches: training the second machine learning model in conjunction with an encoder, wherein the encoder receives a spectrogram as an input and produces a spectrogram vector representing the spectrogram as an output. (P0147, Audio data can include (but is not limited to) raw audio data, feature vectors derived from audio data, such as mel-frequency cepstral coefficients (MFCC), spectrogram data, neural embeddings (e.g., from a convolutional encoder network), etc.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to produce spectrogram vector that represents audio spectrogram. It would have been obvious to combine the references because producing vectors for spectrogram is a known technique that yields a predictable result utilizing the embedding vector for audio signal comparison. (Voss P0110) Claims 28 is rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Voss and in further view of "Similarity Analysis of Self-Supervised Speech Representations" by Chung et al. Regarding claim 28 Prasad in view of Voss teach claim 21. Prasad in view of Voss does not specifically teach: wherein the second machine learning model comprises a token embedding layer and a transformer, the token embedding layer generating embeddings that are input into the transformer. Chung, however, teaches: wherein the second machine learning model comprises a token embedding layer and a transformer, the token embedding layer generating embeddings that are input into the transformer. (Introduction, We hope to understand the similarity of different self-supervised representations. To carry out this study, we adopt two similarity measures for quantifying the similarity of two given representations.; 2.1 Approaches for measuring representation similarity, For an acoustic feature sequence (in our case, a log Mel spectrogram) x = (x1, x2.…,xT), where xt ∈ ℝ80, from a dataset D, the model M transforms x into a representation M(x) = (m1, m2.…,mT), where mt ∈ ℝ512. Given two representations extracted by two self-supervised models M(1) and M(2), a similarity measure outputs sim(M(1)(x), M(2)(x))∈ℝ that quantifies their similarity.; 3.1. Self-supervised models, Model used includes TRF (transformer).) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to input embedding into a transformer. It would have been obvious to combine the references because utilization of a transformer is a known technique that yields a predictable result of finding similarity in speech representation. Claims 30 is rejected under 35 U.S.C. 103 as being unpatentable over Prasad, in view of Voss, in further view Chung, and further view of Jansen et al. (U.S. PG Pub No. 20260072982), hereinafter Janson. Regarding claim 30 Prasad in view of Voss and further view of Chung teach claim 28. Prasad in view of Voss and further view of Chung does not specifically teach: wherein the token embedding layer receives pairs of: sets of textual syllable tokens and sets of audio syllable tokens. Jansen, however, teaches: wherein the token embedding layer receives pairs of: sets of textual syllable tokens and sets of audio syllable tokens. (P0025, Generating a representation of the user preference in a joint audio-text embedding space by applying a two-tower model comprising an audio embedding network to generate an audio embedding of the initial audio track and a text embedding network to generate a text embedding of the natural language input.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to utilize textual syllable token and audio syllable token. It would have been obvious to combine the references because a proximity of two embeddings in the joint audio-text embedding space is indicative of semantic similarity. Jansen P0003. Claims 33, 34, and 36 are rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Voss and further view of Peng et al. (U.S. PG Pub No. 20250364001), hereinafter Peng. Regarding claim 33 Prasad in view of Voss teach claim 32. Prasad in view of Voss does not specifically teach: wherein the training of the second machine learning model further uses a codebook that stores vectors representing a group of spectrograms, wherein the training of the second machine learning model include performing nearest-neighbor mapping of the stored vectors in the codebook to find a matching vector that is most similar to the spectrogram vector produced via the encoder. Peng, however, teaches: wherein the training of the second machine learning model further uses a codebook that stores vectors representing a group of spectrograms, wherein the training of the second machine learning model include performing nearest-neighbor mapping of the stored vectors in the codebook to find a matching vector that is most similar to the spectrogram vector produced via the encoder. (P0041, The vector quantizer discretizes the learned features in encoding with a set of learnable codebooks according to the target bitrate. … Group quantization is obtained by splitting channels C′ into N groups and coding each group by an independent codebook. Let S denote the number of codewords in each codebook and K=C′/N the dimension of each codeword. … Discrete codes can be determined using a nearest neighbor lookup procedure using a shared embedding space.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to utilize a codebook to map spectrogram vector to codebook vector. It would have been obvious to combine the references because the use of codebook and the process of vector quantization increases the efficiency of processing by reducing the number of bits needed to carry audio information. Peng P0021. Regarding claim 34 Prasad in view of Voss and further view of Peng teach claim 33. Prasad in view of Voss does not specifically teach: training the encoder by obtaining, from the codebook, an index value that represents the matching vector and converting the index value to a one-hot encoding embedding that is a respective one of the audio tokens. Peng, however, teaches: training the encoder by obtaining, from the codebook, an index value that represents the matching vector and converting the index value to a one-hot encoding embedding that is a respective one of the audio tokens. (P0042, An input x can be passed through an encoder to generate an output ze(x), where discrete latent variables z can be determined using a shared embedding space e (having embedding vectors ej) for a nearest neighbor look-up. The encoder output can then be passed through a discretization bottleneck, and then mapped onto a nearest embedding e.; P0069, As quantization is not differentiable, a technique is used to learn the codebook and perform back propagation through the vector quantization process. … The Gumbel-Softmax distribution interpolates between discrete one-hot-encoded categorical distributions and continuous categorical densities.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to utilize an index value and one-hot encoding. It would have been obvious to combine the references because the use of an index value is a known technique that yields predictable result of obtaining codebook index for vector quantization and the use of one-hot encoding is a known technique that yields predictable result of creating an embedding with known indexes. Regarding claim 36 Prasad in view of Voss and further view of Peng teach claim 34. Prasad in view of Voss does not specifically teach: wherein the encoder is trained in conjunction with a decoder that receives the index value as input and that, in response, produces a reconstructed spectrogram as output, the reconstructed spectrogram being intended to match the spectrogram that is input into the encoder. Peng, however, teaches: wherein the encoder is trained in conjunction with a decoder that receives the index value as input and that, in response, produces a reconstructed spectrogram as output, the reconstructed spectrogram being intended to match the spectrogram that is input into the encoder. (P0004, At a decoder portion of a coding framework, the quantized residual-like feature is dequantized and then combined with a prediction from prior reconstructed latent features to provide reconstructed features of a current frame, which can then be processed by a decoder to provide a reconstructed signal.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to utilize a decoder to reconstruct the spectrogram. It would have been obvious to combine the references because the use of a decoder is a known technique that yields a predictable result of reconstructing spectrogram from vector quantized embeddings. Claims 37 is rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Voss, in view of Peng, and further view of Rosenberg et al. (U.S. PG Pub No. 20240029715), hereinafter Rosenberg. Regarding claim 37 Prasad in view of Voss and further view of Peng teach claim 33. Prasad in view of Voss and further view of Peng does not specifically teach: generating a single-syllable spectrogram with a width smaller than a maximum width by filling in blank areas around the spectrogram with color. Rosenberg, however, teaches: generating a single-syllable spectrogram with a width smaller than a maximum width by filling in blank areas around the spectrogram with color. (P0040, The speech encoder receives, as input, each transcribed non-synthetic speech utterance as a sequence of features/vectors (e.g., mel-frequency spectrograms such as the acoustic frames.; P0034, The duration predictor receives the initial textual representation from the embedding extractor and predicts a corresponding text chunk duration (i.e., word, word-piece, phoneme, and/or grapheme duration).; P0037, The number of frames of the alignment output indicates a predicted speech duration of the unspoken textual utterance. Stated differently, the number of frames of the alignment output maps (i.e., aligns) the sequence of text chunks of the unspoken textual utterance to speech frames.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention to generate single syllable spectrogram. It would have been obvious to combine the references because utilizing aligned speech representation can train models without transcribed speech training data. Claims 38 is rejected under 35 U.S.C. 103 as being unpatentable over Prasad in view of Voss, in view of Peng, and view of Rosenberg, and further view of Netzer (U.S. PG Pub No. 20220013120). Regarding claim 38 Prasad in view of Voss teach claim 21. Prasad in view of Voss does not specifically teach: calculating a time to pronounce each of multiple syllables in the first language; identifying points of zero amplitude in an audio waveform generated via pronouncing the input text; and dividing an input text spectrogram for the input text into a spectrogram per syllable of the input text, wherein the dividing is based on the calculated time and on the identified points of zero amplitude. Rosenberg, however, teaches: calculating a time to pronounce each of multiple syllables in the first language; (P0034, The duration predictor receives the initial textual representation from the embedding extractor and predicts a corresponding text chunk duration (i.e., word, word-piece, phoneme, and/or grapheme duration). The text chunk duration indicates a duration the corresponding text chunk would be spoken if a human (or text-to-speech system) spoke the unspoken textual utterance.) dividing an input text spectrogram for the input text into a spectrogram per syllable of the input text, wherein the dividing is based on the calculated time and on the identified points of zero amplitude. (P0040, The speech encoder receives, as input, each transcribed non-synthetic speech utterance as a sequence of features/vectors (e.g., mel-frequency spectrograms such as the acoustic frames of FIG. 1) and generates, as output, for each of a plurality of output steps, an encoded audio representation (es) that corresponds to the transcribed non-synthetic speech utterance at the corresponding output step. In parallel, the alignment model receives the transcription corresponding to the same non-synthetic speech utterance and generates an alignment output according to Equation 1.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to calculate time for pronunciation and divide the spectrogram based on the time. It would have been obvious to combine the references because calculating a duration of time is a known technique to yield a predictable result of mapping sequence of text chunks to speech frames directly. (Rosenberg P0039). Prasad in view of Voss and further view of Rosenberg does not specifically teach: identifying points of zero amplitude in an audio waveform generated via pronouncing the input text; and Netzer, however, teaches: identifying points of zero amplitude in an audio waveform generated via pronouncing the input text; and (P0053, By analyzing the representation of analog speech signal, client of FIG. 1 may be configured to distinguish between segments of the speech signal 300, wherein speech segment (SS) (representing the words ELIAV and DANIELLE respectively) are examples of a speech segment. While segments are an example of silence segment. In some exemplary embodiments, SS303 is attributed to speaking pause, end of speech, silence, or the like due to lack of speech signal or a substantially low speech signal amplitude. … the segment represents speech elements selected from a group comprising of: a syllable; a plurality of syllables; a word; a fraction of a word; a plurality of words; and a combination thereof.) It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to identify points of zero amplitude. It would have been obvious to combine the references because dividing spectrogram by points of zero amplitude can solve the issue of speaking pause, end of speech, silence, or the like due to lack of speech signal or a substantially low speech signal amplitude. (Netzer P0053). Allowable Subject Matter Claim 24-27, 29, 31, and 35 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Regarding claims 24-27 none of the prior art either alone or in combination, teaches or makes obvious the reward being “determined via providing the textual syllable embeddings in the target language into a machine learning classifier that, in response, produces a classification of a predicted language of the textual syllable embeddings”. The closest prior art reference Voss teaches a loss function but does not teach the reward being determined by a machine learning classifier that produces a classification of a predicted language. Regarding claim 29 none of the prior art either alone or in combination, teaches or makes obvious the token embedding layer generating final embedding that includes “token embeddings, modal-type embeddings, position embeddings, and language-category embeddings”. The closest prior art reference Voss teaches the use of acoustic embedding for template matching. Voss, however, does not teach the generation of a final embedding that includes token embeddings, modal-type embeddings, position embeddings, and language-category embeddings. Regarding claim 31 none of the prior art either alone or in combination, teaches or makes obvious the token embedding layer “receiv[ing] separator tokens which separate a respective set of the textual syllable tokens and a paired respective set of the audio syllable tokens”. The closest prior art reference Voss teaches the use of acoustic embedding for template matching. Voss, however, does not teach the receiving of a separator token to specifically separate textual syllable token with audio syllable token. Regarding claim 35 none of the prior art either alone or in combination, teaches or makes obvious the index value selected from a range from “N to N + K, where K is a hyper-parameter and is equal to a total number of unduplicated audio syllables and N is also a hyper-parameter and is equal to a total number of all textual syllable embeddings”. The closest prior art reference Peng teaches the use of S as a notation for the number of codewords in a codebook that reflect audio bitstream and receivers. Peng, however, does not teach the index value selected from a range from “N to N + K, where K is a hyper-parameter and is equal to a total number of unduplicated audio syllables and N is also a hyper-parameter and is equal to a total number of all textual syllable embeddings”. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL WONSUK CHUNG whose telephone number is (571)272-1345. The examiner can normally be reached Monday - Friday (7am-4pm)[PT]. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PIERRE-LOUIS DESIR can be reached at (571)272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /DANIEL W CHUNG/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Sep 28, 2022
Application Filed
Jan 09, 2026
Non-Final Rejection mailed — §101, §103, §112
Mar 31, 2026
Interview Requested
Apr 07, 2026
Applicant Interview (Telephonic)
Apr 08, 2026
Examiner Interview Summary
Apr 09, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744029
Phonemes And Graphemes for Neural Text-to-Speech
2y 3m to grant Granted Sep 22, 2026
Patent 12730969
MULTISTAGE ALIGNMENT FOR GENERATING ARTIFICIAL INTELLIGENCE TRAINING DATA
2y 6m to grant Granted Sep 08, 2026
Patent 12731574
DATA PROCESSING METHOD, AND STORAGE MEDIUM AND ELECTRONIC DEVICE THEREOF
2y 7m to grant Granted Sep 08, 2026
Patent 12711942
CROSS-SPEAKER STYLE TRANSFER SPEECH SYNTHESIS
4y 0m to grant Granted Aug 18, 2026
Patent 12700420
ELECTRONIC APPARATUS, CONTROL METHOD THEREOF AND ELECTRONIC SYSTEM
4y 6m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
59%
Grant Probability
90%
With Interview (+31.2%)
2y 12m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 54 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month