DETAILED ACTION
This office action is in response to Applicant’s submission filed on 1/16/2025. Claims 1-20 are pending in the application of which Claims 1, 9, and 16 are independent and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-7 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception.
The flowchart in MPEP 2106, subsection III, is used to determine whether a claim satisfies the criteria for subject matter eligibility. For analysis purposes, one can follow the flowchart for subject matter eligibility.
PNG
media_image1.png
628
432
media_image1.png
Greyscale
Step 1: The independent Claims is directed to statutory categories:
Step 1: Abstract Idea Groupings – MPEP 2106.04(a)(2)
The enumerated groupings of abstract ideas are defined as:
1) Mathematical concepts – mathematical relationships, mathematical formulas or equations, mathematical calculations (see MPEP § 2106.04(a)(2), subsection I);
2) Certain methods of organizing human activity – fundamental economic principles or practices (including hedging, insurance, mitigating risk); commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations); managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions) (see MPEP § 2106.04(a)(2), subsection II); and
3) Mental processes – concepts performed in the human mind (including an observation, evaluation, judgment, opinion) (see MPEP § 2106.04(a)(2), subsection III).
Claim 1 is a method claim and directed to the process category of patentable subject matter.
Step 2A is a two-prong test.
PNG
media_image2.png
404
780
media_image2.png
Greyscale
Step 2A, Prong One: Does the Claim recite a Judicially Recognized Exception? Abstract Idea? Are these Claims nevertheless considered Abstract as a Mathematical Concept (mathematical relationships, mathematical formulas or equations, mathematical calculations), Mental Process (concepts performed in the human mind (including an observation, evaluation, judgment, opinion), or Certain Methods of Organizing Human Activity (1-fundamental economic principles or practices (including hedging, insurance, mitigating risk), 2-commercial or legal interactions (including agreements in the form of contracts; legal obligations; advertising, marketing or sales activities or behaviors; business relations), 3- managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules or instructions) and fall under the judicial exception to patentable subject matter?)
The broadest reasonable interpretation of steps in the claim limitations is that those steps fall within the mental process groupings of abstract ideas because they cover concepts performed in the human mind, including observation, evaluation, judgment, and opinion. See MPEP 2106.04(a)(2), subsection III.
A method comprising:
processing, using an encoder, one or more audio frames representative of a speech to generate one or more embeddings encoding the speech; and [This is merely amount to taking a speech of a talker and write it down on a piece of paper, and subsequently transforming to another representation such as Cartesian coordinates, a data gathering activity, which can be carry it out on a piece of paper.]
processing, using a plurality of decoders, the one or more embeddings to generate a plurality of transcriptions of the speech, an individual transcription of the plurality of transcriptions; [This is merely amount to a take the transformed data and identify the words that were uttered by the talker and write it down as a text.]
(i) being generated by a respective decoder of the plurality of decoders and; (ii) conforming to a respective text format of a plurality of text formats that differs from other text formats of the plurality of text formats in at least one of capitalization, punctuation, use of non-alphabet characters, or identification of individual utterances of the speech. [This involves writing the text in different format. A human can take a text and rewrite it with Capitalization format, or punctuation format, etc. which can be carried out by a human with the help of pen and paper.]
The rejected Claims recite Mental Processes.
Step 2A, Prong Two: Additional Elements that Integrate the Judicial Exception into a Practical Application? Identifying whether there are any additional elements recited in the claim beyond the judicial exception(s), and evaluating those additional elements to determine whether they integrate the exception into a practical application of the exception. “Integration into a practical application” requires an additional element(s) or a combination of additional elements in the claim to apply, rely on, or use the judicial exception in a manner that imposes a meaningful limit on the judicial exception, such that the claim is more than a drafting effort designed to monopolize the exception. Uses the considerations laid out by the Supreme Court and the Federal Circuit to evaluate whether the judicial exception is integrated into a practical application.
The rejected Claims do not include additional limitations that point to integration of the abstract idea into a practical application. Accordingly, the rejected Claims are directed to the abstract idea that they recite.
Claim 1 is a generic automation of a mental process since a human agent can receive, analyze, generate, evaluate, etc. Other than the mental process under the BRI, there is only the mention of a decoder and encoder, which is considered to be generic computer components that are merely being used as a tool to perform the abstract idea, and as such it is considered a generic component due to lack of specificity. With such a generic extra element, one cannot identify anything that can be relied upon as an improvement. Prong 2 of step 2A, in the 101 analysis, asks whether the abstract idea is integrated into a practical application. The answer is no in this instance because there is no technological solution in the Claim that “integrates” the abstract idea. The Claim only suggests that the abstract idea be applied. It does not describe an application.
These limitations, under their broadest reasonable interpretation, cover performance of the limitation in the mind but for the recitation of generic computer components. That is, other than reciting “encoder”, and “decoder” nothing in the claim element precludes the step from practically being performed in the mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements of using a “encoder”, and “decoder” to perform all of the above-mentioned steps. The use of a “encoder”, and “decoder” is recited at a high-level of generality (i.e., as a generic computer component performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component See MPEP2106.05(f) Mere Instructions to Apply an Exception [R-10.2019].
Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
PNG
media_image3.png
526
736
media_image3.png
Greyscale
Step 2B: Search for Inventive Concept: Additional Element Do not amount to Significantly More: The limitations of " generating one or more embeddings encoding the speech,” or “generating a plurality of transcription” is a well-understood, routine, and conventional machine components that and are being used for their well-understood, routine, and conventional and rather generic functions. Additionally, these limitations are expressed parenthetically and lack nexus to the claim language and as such are a separable and divisible mention to a machine. Merely reciting encoder/decoder without significantly more appears to be equivalent to a generic computer/processor to process a task that a human can process in their mind or with the aid of a paper/pen.
Therefore, the cited additional element of encoder/decoder does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. Accordingly, it is not sufficient to cause the Claim, as a whole, to amount to significantly/substantially more than the underlying abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements of using a “encoder”, and “decoder” to perform all of the above-mentioned steps. The use of a “encoder”, and “decoder” is recited at a high-level of generality (i.e., as a generic computer component device performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component See MPEP2106.05(f) Mere Instructions to Apply an Exception [R-10.2019]. Accordingly, these additional elements do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The dependent claims do not add limitations that would either integrate the recited abstract idea into a practical application or could help the Claim as a whole to amount to significantly more than the Abstract idea identified for the Independent Claim:
Claim 2 recite: “wherein the plurality of decoders comprises one or more of: a connectionist temporal classification (CTC) decoder, or a transducer decoder, a hybrid CTC-transducer decoder, or a token-and-duration transducer decoder.” The claim is reciting additional elements which are generic computer component due to lack of specificity. These additional elements are generic computer components and the hardware is generic computer components that are merely being used as a tool to perform the abstract idea. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim directed toward abstract idea. The claim is not patent eligible.
Claim 3 recite: “wherein the plurality of text formats comprises a normalized text format associated with representing text content with single-case alphabet characters.” Human can provide different form of textual format with the help of a pen and paper. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim directed toward abstract idea. The claim is not patent eligible.
Claim 4 recite: “ wherein the plurality of text formats comprises an inverse normalized text format associated with representing numbers with numeral characters.” Human can carry out the claimed limitation merely by employing a pen and paper. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim directed toward abstract idea. The claim is not patent eligible.
Claim 5 recite: “wherein the inverse normalized text format is further associated with representing at least some of vocabulary words with non-alphanumerical characters.” Similarly, human can carry out the claimed limitation merely by employing a pen and paper. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim directed toward abstract idea. The claim is not patent eligible.
Claim 6 recite: “wherein the plurality of text formats comprises one or more text formats that comprise a combination of at least two other text formats of the plurality of text formats.” Similarly, human can carry out the claimed limitation merely by employing a pen and paper. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim directed toward abstract idea. The claim is not patent eligible.
Claim 7 recite: “wherein the identification of individual utterances of the speech comprises identification of at least one content unit of the speech, the content unit comprising multiple sentences of the speech having at least one of a common meaning, logic, semantics, or context.” Human can listen to a speech and identify the content of the spoken words. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The claim directed toward abstract idea. The claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 8-10, 15 - 17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Thomson et al. (US 20200243094 A1)(herein " Thomson "), and in further view of Liang et al. (“Speech Emotion Recognition Exploiting ASR-based and Phonological Knowledge Representations”)(herein " Liang ").
Regarding claims 1, 9 and 16, Thomson teaches a method comprising [one or more processors to:-claim 16] (Thomson, Par. 0104:” … include memory and at least one processor, which are configured to perform operations as described in this disclosure, among other operations.”)
[processing, using an encoder, one or more audio frames representative of a [training] speech to generate one or more embeddings encoding the [training] speech; and processing, [[using a plurality of decoders,]] the one or more embeddings to generate a plurality of transcriptions of the [training] speech, an individual transcription of the plurality of transcriptions – claims 1, and 9], process, using an encoder, audio data to generate one or more embeddings encoding speech represented in the audio data; process, [[using a plurality of decoders,]] the one or more embeddings to generate a plurality of transcriptions of the speech, wherein an individual transcription of the plurality of transcriptions – claim 16] (Thomson, Par. 0367:” … hypotheses may be input into a neural network, using an encoding method such as one-hot or word embedding, and the neural network ... “, and Par. 1187:” … specified number of speech analysis frames (i.e., a period of time, usually 5-40ms during which speech is considered to be relatively constant, frames may overlap, frame rate is the distance or time between centers of adjacent frames).”, and Par. 1468:”An encoder 8008 may receive the quantized signal and send it to the model trainer 8016. The encoder 8008 may format the quantizer output, such as by packing bits into words ...”, and Par. 0096:” … the multiple transcriptions may be generated by different systems and/or methods.”)
[(ii) conforming/conforms to a respective text format of a plurality of text formats that differs from other text formats of the plurality of text formats in at least one of capitalization, punctuation, use of non-alphabet characters, or identification of individual utterances of the [training] speech. – claims 1, and 9], [(ii) conforms to a respective text format of a plurality of text formats that differs from other text formats of the plurality of text formats; and – claim 16] (Thomson, Par. 0113:” In general, the transcription system 108 may be configured to generate or direct generation of the transcription of audio using one or more automatic speech recognition (ASR) systems. The term “ASR system” … configured to recognize speech in audio and generate a transcription of the audio based on the recognized speech. … In some embodiments, the transcription of the audio generated by the ASR systems may include capitalization, punctuation, and non-speech sounds.”)
[claims 9, and 16] training one or more decoders of the plurality of decoders based at least on at least the plurality of transcriptions of the [training] speech. (Thomson, Par. 0097:” … ASR system to recognize words in speech.”, and Par. 0259:” … , the ASR system 520 may have two language models, one for the decoder 510 and one for the rescorer 512.”, and Par. 0266:” … Training models may also be referred to as training an ASR system.”, and Par. 0318:” … The multiple transcription units may include a first transcription unit 1214a, a second transcription unit 1214b, and a third transcription unit 1214c. The transcription units 1214a, 1214b, and 1214c may be referred to collectively as the transcription units 1214.”, and Par. 0675:” … transcriptions may be used for training of ASR systems, …”).
Thomson, does not teach, however Liang teaches using a plurality of decoders, (i) being/is generated by a respective decoder of the plurality of decoders and (Liang, Page 217:”… The acoustic features, ASR-based and phonological knowledge representations are concatenated together as input features, with Conformer as a classifier. … In our work, because of the correlated objectives of phonological knowledge representation learning, we design a multi-task-based model which is a shared encoder-multi decoder (SEMD) model. The shared encoder can capture high level speech representation in phonological space.”, and Section 3.3:” … We propose the shared encoder to learn the one phonological space shared by these subtasks and the multi-decoder to solve the problem that there are different PFs labels in each subtask.”)
PNG
media_image4.png
334
556
media_image4.png
Greyscale
Liang is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson further in view of Liang to use a plurality of decoders, (i) being generated by a respective decoder of the plurality of decoders. Motivation to do so would allow the model to share a heavy acoustic-feature encoder while using distinct decoders for specialized outputs.
Regarding claim 8, Thomson, as modified above, teaches the method of claim 1.
Thomson, as modified above, does not teach, however, Liang further teaches wherein the plurality of decoders are trained using an aggregated loss value computed based at least on at least a plurality of loss values, an individual loss value of the plurality of loss values being associated with an accuracy of a training text output generated, by a corresponding decoder of the plurality of decoders, for a training speech input. (Liang, Page 218:” … The multi-decoder consists of several Transformer layers and takes encoder output and different PFs as input and output labels. Then we fine-tune the pre-trained PFs model on the IEMOCAP to extract the encoder-out feature as phonological representations. At the stage of pre-training and fine-tuning, the training loss is the average CE loss of these six subtasks,”) Note: summation of all 6 encoder loss stages reads on aggregated loss value.
PNG
media_image5.png
34
526
media_image5.png
Greyscale
Regarding claims 10, and 17, Thomson, as modified above, teaches the method, and the system of claims 9, and 16 respectively.
Thomson, as modified above, does not teach, however, Liang further teaches wherein the training the one or more decoders comprises: computing a plurality of loss values, an individual loss value of the plurality of loss values characterizing accuracy of a corresponding transcription of the plurality of transcriptions of the training speech; (Liang, Page 218:” … the training loss is the average CE loss of these six subtasks,”).
computing, based at least on at least the plurality of loss values, an aggregated loss value; and (Liang, Page 218:” … At the stage of pre-training and fine-tuning, the training loss is the average [aggregated] CE loss of these six subtasks,”) Note: summation of all 6 encoder loss stages reads on aggregated loss value.
PNG
media_image5.png
34
526
media_image5.png
Greyscale
modifying, using the aggregated loss value, parameters of the one or more decoders. (Liang, Page 218:” The multi-decoder consists of several Transformer layers and takes encoder output and different PFs as input and output labels. … the training loss is the average [aggregated] CE loss of these six subtasks,”) Note: During training, loss value serves as a direct mathematical reflection of the network's error. Parameters (weights and biases) are iteratively modified to reduce the loss, where in this case instead of individual loss, the aggregated loss is being modified.
Regarding claim 15, Thomson, as modified above, teaches the method of claim 9.
Thomson, as modified above, does not teach, however, Liang further teaches training, using the plurality of transcriptions of the training speech, the encoder. (Liang teaches training an encoder-decoder structure in section 3.2, and 3.3)
Regarding claim 20, Thomson, as modified above, teaches the system of claim 16.
Thomson, as modified above, further teaches wherein the system is comprised in at least one of: an in-vehicle infotainment system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing one or more medical operations; a system for performing one or more factory operations; a system for performing one or more analytics operations; (Thomson teaches Analytic in Pars. 0155, 0784 and table 4.)a system implementing one or more inference microservices; a system for performing light transport simulations; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; (Thomson teaches deep learning in Pars. 0616, 1256, and table 9.) a system implemented using an edge device; a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content; a system implemented using a robot; a system for performing one or more conversational AI operations; (Thomson teaches conversational AI in Par. 1553, and table 5.) a system implementing one or more large language models (LLMs);( Thomson teaches large language model in Pars. 0262, 0295.) a system implementing one or more vision language models (VLMs); a system implementing one or more multi-modal language models (MMLMs); a system implementing one or more language models; (Thomson teaches language model in Pars. 0073, 0258, 0259, 0262, etc.) a system for performing one or more generative AI operations; (Thomson teaches generative AI in Pars. 0616, 1337, and table 5.) a system for generating synthetic data; (Thomson teaches synthetic data generation in Par. 1337.) a system incorporating one or more virtual machines (VMs); (Thomson teaches virtual machine in Pars. 0204, 1699.) a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (Thomson teaches cloud computing in Pars. 0204, 1698, 1702).
Claim 2, is rejected under 35 U.S.C. 103 as being unpatentable over Thomson, and Liang et al. and in further view of Hainan et al. (DE102024110264A1)(herein " Hainan ").
Regarding claim 2, Thomson, as modified above, teaches the method of claim 1.
Thomson, as modified above, does not teach, however, Liang further teaches wherein the plurality of decoders comprises one or more of: a connectionist temporal classification (CTC) decoder, or a transducer decoder, (Liang, P. 218, Section 3.2:” At the stage of fine-tuning, the training loss is the combined connectionist temporal classification (CTC) and cross-entropy (CE) loss, …”).
Thomson, as modified above, does not teach, however, Hainan teaches a hybrid CTC-transducer decoder, or a token-and-duration transducer decoder. (Hainan, Fig. 3 represents an issuance probability grid of a token-and-duration model according to at least one embodiment;”).
Hainan is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Hainan to comprise a hybrid CTC-transducer decoder, or a token-and-duration transducer decoder. Motivation to do so would accelerate inference while improving transcription accuracy.
Claim 3, is rejected under 35 U.S.C. 103 as being unpatentable over Thomson, and Liang et al. and in further view of Wang et al. (US 20190087417 A1)(herein " Wang ").
Regarding claim 3, Thomson, as modified above, teaches the method of claim 1.
Thomson, as modified above, however Wang teaches wherein the plurality of text formats comprises a normalized text format associated with representing text content with single-case alphabet characters. (Wang, Par. 0005:” … In some instances, electronic text messages can be tokenized and/or normalized, for example, by removing control characters, adjusting character widths, replace XML characters, etc.”, and Par. 0056:”The tokenized text from the tokenization module 208 can be provided to the lower-casing module 210 to produce a message having lowercase letters and preferably no uppercase letters. Note that, in some implementations, lower-casing can be applied to languages that do not have lowercase letters (e.g., unicameral scripts including Japanese, Korean, etc.) because chat messages in those languages may contain words from languages that do have lowercase letters (e.g., bicameral scripts including Latin-based, Cyrillic-based languages, etc.).”)
Wang is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Wang to comprise a normalized text format associated with representing text content with single-case alphabet characters. Motivation to do so would allow to process text and typography much more fluidly without the added burden of parsing capitalization rules.
Claims 4, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Thomson, and Liang and in further view of Sunkara et al. (US 12182498 B1)(herein " Sunkara ").
Regarding claims 4, and 14, Thomson, as modified above, teaches the method of claims 1, and 9 respectively.
Thomson, as modified above, however Sunkara teaches wherein the plurality of text formats comprises an inverse normalized text format associated with representing numbers with numeral characters. (Sunkara, Col. 7, ll. 62-65:” The normalized text provided by neural network inverse text normalization model 440 may be provided along with a corresponding confidence score or value associated with the predicted normalization of the text provided by the model 440. Inverse text normalization model selection 450 may apply threshold or other criteria based on the confidence score to …”).
Sunkara is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Sunkara to comprise an inverse normalized text format associated with representing numbers with numeral characters. Motivation to do so would improve the performance of redaction techniques in machine transcription systems (Sunkara, Col. 2, ll. 24-25).
Claim 5, is rejected under 35 U.S.C. 103 as being unpatentable over Thomson, Liang, and Sunkara and in further view of Sahbhaumik et al. (US 20200211543 A1)(herein " Sahbhaumik ").
Regarding claim 5, Thomson, as modified above, teaches the method of claim 4.
Thomson, as modified above, does not teach, however, Sahbhaumik teaches wherein the inverse normalized text format is further associated with representing at least some of vocabulary words with non-alphanumerical characters. (Sahbhaumik, Par. 0022:” … For example, if the speech-to-text conversion engine 211 detects that a user has spoken or otherwise input a special character (e.g., by speaking the words “exclamation point,” “dollar sign,” etc.), then the parsing and context engine 212 may either convert those words back into the appropriate special character, or may strip out those characters altogether from the text string.”, and Par. 0038:” … For example, if the speech-to-text conversion engine 211 detects that a user has spoken or otherwise input a special character (e.g., by speaking the words “exclamation point,” “dollar sign,” etc.), then in steps 403 or 404 a parsing and context engine 212 may either convert those words back into the appropriate special character (e.g., “!” or $“, etc.), or may strip out those characters altogether from the text string.”)
Sahbhaumik is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Sahbhaumik to wherein the inverse normalized text format is further associated with representing at least some of vocabulary words with non-alphanumerical characters. Motivation to do so would provide a strong password when non-alphanumeric characters are used.
Claim 6, is rejected under 35 U.S.C. 103 as being unpatentable over Thomson, and Liang and in further view of Galley et al. (US 20160352656 A1)(herein " Galley ").
Regarding claim 6, Thomson, as modified above, teaches the method of claim 1.
Thomson, as modified above, however Galley teaches wherein the plurality of text formats comprises one or more text formats that comprise a combination of at least two other text formats of the plurality of text formats. (Galley, Par. 0061:” … A combination format response is a combination of two or more different formats. In other words, a combination format response includes a combination of two or more of a text format response, ...”).
Galley is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Galley to wherein the plurality of text formats comprises one or more text formats that comprise a combination of at least two other text formats of the plurality of text formats. Motivation to do so would increase flexibility, accessibility, and compatibility which allows content to be tailored for different systems, devices, and human needs without losing data or meaning.
Claim 7, is rejected under 35 U.S.C. 103 as being unpatentable over Thomson, and Liang and in further view of Kobayashi et al. (US 20100049500A1)(herein " Kobayashi ").
Regarding claim 7, Thomson, as modified above, teaches the method of claim 1.
Thomson, as modified above, however Kobayashi teaches wherein the identification of individual utterances of the speech comprises identification of at least one content unit of the speech, the content unit comprising multiple sentences of the speech having at least one of a common meaning, logic, semantics, or context. (Kobayashi, Par. 0122:” The text segmentation unit 850 segments the incoming text according to a specific segmentation rule and inputs the segmented text items sequentially to the morphological analysis unit 104 and speech synthesis unit 102. The segmentation rule may be, for example, to segment the incoming text in sentences or in linguistic units larger than sentences (e.g., topics). When the incoming text is segmented in topic units, the text is segmented on the basis of the presence or absence of a linefeed or of a representation of topic change.”)
Kobayashi is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Kobayashi to comprise comprises identification of at least one content unit of the speech, the content unit comprising multiple sentences of the speech having at least one of a common meaning, logic, semantics, or context. Motivation to do so would provide communication efficiency and informativeness and reveal how accurate and relevant information is successfully being conveyed.
Claims 11, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Thomson, and Liang and in further view of Cormack et al. (US20160371369 A1)(herein " Cormack ").
Regarding claims 11, and 18, Thomson, as modified above, teaches the method, and the system of claims 10, and 17, respectively.
Thomson, as modified above, does not teach, however Cormack teaches wherein at least two loss values of the plurality of loss values are weighted with unequal weights in the aggregated loss value. (Cormack, Par. 0070:” In certain embodiments, various loss measures (e.g., lossr and losse) are aggregated to form a combined loss measure. When aggregating the various loss measures, each individual loss function may be weighted. In certain embodiments, the loss measures are weighted equally. In certain embodiments, the loss measures are unequally weighted.”)
Cormack is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Cormack to wherein at least two loss values of the plurality of loss values are weighted with unequal weights in the aggregated loss value. Motivation to do so would allow the network to prioritize specific decoding streams based on their relative importance, reliability, or specific task goals.
Claims 12, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Thomson, and Liang and in further view of Salama et al. (US 20240427997 A1)(herein " Salama "), and Ananthakrishnan et al. (“Automatic diacritization of Arabic Transcripts for Automatic Speech recognition”)(herein “Ananthakrishnan”).
Regarding claims 12, and 19, Thomson, as modified above, teaches the method, and the system of claims 10, and 17, respectively.
Thomson, as modified above, does not teach, however Salama teaches wherein the plurality of loss values are computed based at least on a plurality of ground truth transcriptions generated by processing the training speech using a [[cascade automatic speech recognition (ASR) system]] that comprises: (Salama, Par. 0003:” … each speech recognition hypothesis representing a corresponding candidate transcription of the training query and generated by a speech recognizer from audio data characterizing the training query; and a corresponding ground-truth transcription of the training query. … processing, using a transcription decoder, the corresponding NSP embedding to generate a corresponding predicted transcription of the training query; and determining a corresponding first loss based on the corresponding predicted transcription of the training query and the corresponding ground-truth transcription of the training query. The operations further include training, based on the first losses determined for the set of training queries, the NSP model to learn how to predict user intents associated with the operations specified by the training queries.”)
Salama is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Salama to wherein the plurality of loss values are computed based at least on a plurality of ground truth transcriptions generated by processing the training speech. Motivation to do so would guide the model away from misalignment, leading to faster, more robust convergence.
Thomson, as modified above, does not teach, however Ananthakrishnan teaches cascade automatic speech recognition (ASR) system (Ananthakrishnan, Section 1, Introduction :” Transcription generated and are normalized it subsequently”). Note: transcript generation along with the normalized transcript read on a cascade ASR system.
an ASR model configured to process the training speech to generate a first ground truth transcription of the plurality of ground truth transcriptions; and a text processing model configured to process the first ground truth transcription to generate a second ground truth transcription of the plurality of ground truth transcriptions. (Ananthakrishnan, Section 1, Introduction :” Transcription generated and are normalized by removal of one or more short vowels, or one or more diacritics”). Note: the normalized transcript read on a second ground truth transcription.
Ananthakrishnan is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Ananthakrishnan to use cascade automatic speech recognition (ASR) system, process the training speech to generate a first ground truth transcription of the plurality of ground truth transcriptions; and a text processing model configured to process the first ground truth transcription to generate a second ground truth transcription of the plurality of ground truth transcriptions. Motivation to do so would remove non-semantic differences such as letter casing, punctuation, and spelling variants, resulting in a cleaner and more accurate evaluation of recognition model performance.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Thomson, Liang, Salama, and Ananthakrishnan and in further view of Wang, and Sahbhaumik.
Regarding claim 13, Thomson, as modified above, teaches the method of claim 12.
Thomson, as modified above, does not teach, however Salama teaches wherein the first ground truth transcription conforms to a normalized text format associated with representing text content with single-case alphabet characters, and (Wang, Par. 0005:” … In some instances, electronic text messages can be tokenized and/or normalized, for example, by removing control characters, adjusting character widths, replace XML characters, etc.”, and Par. 0056:”The tokenized text from the tokenization module 208 can be provided to the lower-casing module 210 to produce a message having lowercase letters and preferably no uppercase letters. Note that, in some implementations, lower-casing can be applied to languages that do not have lowercase letters (e.g., unicameral scripts including Japanese, Korean, etc.) because chat messages in those languages may contain words from languages that do have lowercase letters (e.g., bicameral scripts including Latin-based, Cyrillic-based languages, etc.).”)
Wang is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Wang to wherein the first ground truth transcription conforms to a normalized text format associated with representing text content with single-case alphabet characters. Motivation to do so would allow to process text and typography much more fluidly without the added burden of parsing capitalization rules.
Thomson, as modified above, does not teach, however Sahbhaumik teaches wherein the second ground truth transcription conforms to text format representing at least one of: a capitalized text format, a punctuated text format, a text format with non-alphabet characters, or a text format that identifies utterances of the first ground truth transcription. (Sahbhaumik, Par. 0022:” … For example, if the speech-to-text conversion engine 211 detects that a user has spoken or otherwise input a special character (e.g., by speaking the words “exclamation point,” “dollar sign,” etc.), then the parsing and context engine 212 may either convert those words back into the appropriate special character, or may strip out those characters altogether from the text string.”, and Par. 0038:” … For example, if the speech-to-text conversion engine 211 detects that a user has spoken or otherwise input a special character (e.g., by speaking the words “exclamation point,” “dollar sign,” etc.), then in steps 403 or 404 a parsing and context engine 212 may either convert those words back into the appropriate special character (e.g., “!” or $“, etc.), or may strip out those characters altogether from the text string.”)
Sahbhaumik is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Thomson, as modified above, further in view of Sahbhaumik to conform to text format representing at least one of: a capitalized text format, a punctuated text format, a text format with non-alphabet characters, or a text format that identifies utterances of the first ground truth transcription.. Motivation to do so would provide a strong password when non-alphanumeric characters are used.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Kim et al. (US 20210050016 A1) teaches in Par. 0008:” receiving, from the device, an encoder output value derived from an encoder of an end-to-end automatic speech recognition (ASR) model in the device; identifying a domain corresponding to the received encoder output value; selecting a decoder corresponding to the identified domain from among a plurality of decoders of an end-to-end ASR model included in the server; obtaining a text string from the received encoder output value using the selected decoder; and providing the obtained text string to the device, wherein the encoder output value is derived by the device by encoding the speech signal input to the device.”
Examiner's Note: Examiner has cited particular columns and line numbers and/or paragraph numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner.
In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DARIOUSH AGAHI whose telephone number is (408)918-7689. The examiner can normally be reached Monday - Thursday and alternate Fridays, 7:30-4:30 PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
DARIOUSH AGAHI, P.E.
Primary Examiner
/DARIOUSH AGAHI/Primary Examiner, Art Unit 2656