Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1,2,-9,12,13,16-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Nachmani et al (“LMs with a Voice: Spoken Language Modeling beyond Speech Tokens”, ,ay 24, 2023, arXiv:2305.15255v2) in view of Manchandra et al (2024032595).
Top of Form
Bottom of Form
As per claim 1, Nachmani et al (“LMs with a Voice:…”) teaches a computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:
receiving a conversational training dataset comprising a plurality of conversational training samples, each conversational training sample in the conversational training dataset associated with a corresponding conversation (as, performing natural language processing (pp1, introduction, and applied in inference during spoken speech prompts in a dialog setting – pp4, top 5 lines) and comprising:
corresponding audio data characterizing a corresponding current utterance spoken by a user during a current turn in the corresponding conversation (as, analyzing the input speech for speech features – pp4, section 3.1.1., input speech processing, equation 1); a corresponding context for the corresponding current utterance, the corresponding context comprising a transcript of a previous turn in the corresponding conversation that precedes the current turn (pp4, section 3.1.1., equation 2, and preceding text – transcripts are generated from the speech utterance processing, and equation 2 shows a mapping of the feature index into the text tokens);
a corresponding ground-truth transcription of the corresponding current utterance (see pp 5, section 3.2.1, using a ground truth distribution of the transcript continuation, and equation 9);
and a corresponding
and for each particular conversational training sample in the conversational training dataset, training a speech model on the particular conversational training sample to teach the speech model to learn how to predict the corresponding logical relationship from the corresponding audio data and the corresponding context (as, performing the training for both speech continuation, as well as speech recognition and transcript continuation – pp , section 3.2, referring into further detail, Figure 1 on pp2). Nachmani et al (“LMs with a Voice:…”) does not explicitly teach ground-truth checking/filtering on the intermediate reasoning/CoT; Manchandra et al (20240320595) teaches producing ‘intermediate reasoning steps’ (also known as Chain-of-Thought) that are tested against ground-truth and fined tuned – para 0126. Therefore, it would have been obvious to one of ordinary skill in the art of machine learning based conversation training to improve upon the Chain-of-Thought/Intermediate Reasoning step of Nachmani et al (“LMs with a Voice:…”) with ground-truth improvement of the C-O-T results, as taught by Manchandra et al (20240320595), because it would advantageously improve the ability to reach the right answer, with the ground-truth fine tuning (see Manchandra et al (20240320595), end of para 0126).
As per claim 2, the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) teaches the computer-implemented method of claim 1, wherein the corresponding logical relationship represented by the corresponding ground truth (see Manchandra et al (20240320595), para 0126) CoT annotation indicates that at least one term is contained in both the corresponding ground-truth transcription of the corresponding current utterance and the transcript of the previous turn (as, in the training example, the CoT results are reflected with post and pre-net – see figure 1).
As per claims 5,6, the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) teaches the computer-implemented method of claim 1, wherein training the speech model on the particular conversational training sample comprises:
processing, by the speech model, the corresponding audio data and the corresponding context to generate a predicted logical relationship between the corresponding current utterance and the previous turn in the corresponding conversation; determining a first cross-entropy loss term for the particular conversational training sample based on the predicted logical relationship and the corresponding logical relationship represented by the corresponding ground truth CoT annotation; determining a second cross-entropy loss term for the particular conversational training sample based on the predicted transcription for the corresponding current utterance and the corresponding ground-truth transcription of the corresponding current utterance (as, Nachmani et al (“LMs with a Voice:…”) see section 3.2.3., with 2 cross entropy losses – one cross entropy loss is for the speech recognition section, the second entropy loss is for the transcription continuation, referring back to section 3.2.2, and 3.2.1); see mapping above in claim 1 for the CoT annotation tied to the transcription continuation; and Manchandra et al (20240320595) teaching ground-truth intermediate reasoning – para 0126);
and training the speech model on the first cross-entropy loss term and the second cross-entropy loss term determined for the particular conversational training sample (as training the speech model (see end of section 3.2.3, where the language models are updated and used to perform the speech recognition).
As per claim 7 the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) teaches the computer-implemented method of claim 6, wherein the speech model is trained to generate the predicted logical relationship between the corresponding current utterance and the previous turn in the corresponding conversation as an intermediate processing step prior to generating the predicted transcription for the corresponding current utterance (pp 4, section 3.1.3. – see chain-of-thought text based language models – that are paralleled to the predicted text that serves as “an intermediate reasoning”; the intermediate text serves as the annotation for the transcripts).
As per claim 8, the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) teaches the computer-implemented method of claim 1, wherein the speech model comprises a speech-text language model comprising: an audio encoder (pp3, section 3.1, first two lines -- as spoken speech encoded); and a large language model decoder (and decoder language model – pp3, section 3.1, first two lines and pp 4, section 3.1.3 “decoder language model”).
As per claim 9, the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) teaches the computer-implemented method of claim 8, wherein the large language model decoder, during inference, generates a CoT annotation for using CoT reasoning during speech recognition (see pp4, section 3.1.3, wherein the large language model predicts text transcription, text continuation, and speech embedding – the text continuation is in the form of annotation).
Claims 12,13,16-20 are system claims that perform the steps found in method claims 1, 2, 5-9 above and as such, claims 12,13,16-20 are similar in scope and content to method claims 1,2,5-9; therefore, claims 12,13,16-20 are rejected under similar rationale as presented against claims 1,2,5-9 above. Furthermore, Nachmani et al (“LMs with a Voice:…”) teaches systems performing the disclosed steps (see Introduction).
Claim(s) 3, 4, 14, 15 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Nachmani et al (“LMs with a Voice:…”) in view of Manchandra et al (20240320595), in further view of CN113033664.
As per claim 3, the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) teaches the computer-implemented method of claim 1, wherein the corresponding logical relationship represented by the corresponding ground truth CoT annotation indicates that at least one term contained in the transcript of the previous turn Nachmani et al (“LMs with a Voice:…”) in view of Manchandra et al (20240320595) does not explicitly detail topical relevance in the comparison of the transcripts and the ground truth distribution. CN113033664 teaches dialog-turn tracking management (pp 13, para 3-4), that tracks content and determines sets of positive and negative examples, based on the fit of the current topic of the dialog turn -- page 13 last 2 paragraphs and pp 14, first 2 lines and first full paragraph. Therefore, it would have been obvious to one of ordinary skill in the art of dialog turn management to further specify the predictive continuation text of Nachmani et al (“LMs with a Voice:…”) in view of Manchandra et al (20240320595) with context/content tracking, as taught by CN113033664 because it would advantageously make the context logic mor fluent and improve the accuracy of the response (see CN113033664, pp14, second full paragraph).
As per claim 4, the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) in view of CN113033664 teaches the computer-implemented method of claim 1, wherein the corresponding logical relationship represented by the corresponding ground truth CoT annotation indicates that the transcript of the previous turn does not contain any terms that are topically relevant to the corresponding ground-truth transcription of the corresponding current utterance (as generating a set of negative examples during the dialog turn -- CN113033664, pp 13, last paragraph; Nachmani et al (“LMs with a Voice:…”) see pp 5, section 3.2.1, using a ground truth distribution of the transcript continuation, and equation 9) ).
Claims 14,15 are system claims that perform the steps in claims 3,4 above and as such, claims 14,15 are similar in scope and content to claims 3,4; therefore, claims 14,15 are rejected under similar rationale as presented against claims 3,4 above.
Claim(s) 10, 11, 21, 22 are rejected under 35 U.S.C. 103 as being unpatentable over the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) in further view of Hanson et al (20220199079).
As per claim 10, the combination Nachmani et al (“LMs with a Voice:…”) of Manchandra et al (20240320595) teaches the computer-implemented method of claim 1, wherein the corresponding ground truth CoT annotation for at least one of the plurality of conversational training samples (see above in claim 1,2), however, does not explicitly teach that the transcript is manually written by a human labeler; Hanson et al (20220199079) teaches analysis of dialog turns with transcriptions (abstract) with human commentary/corrections (para 0208, as humans ranking the results of the AI machine; and comparison of the automated results and human evaluations – par 0210). Therefore, it would have been obvious to one of ordinary skill in the art of speech transcription to modify the process of the combination Nachmani et al (“LMs with a Voice:…”) in view of Manchandra et al (20240320595) with an additional step of human evaluation of the results, as taught by Hanson et al (20220199079) because it would advantageously make the results more human like (Hanson et al (20220199079), end of para 0210).
As per claim 11, the combination of Nachmani et al (“LMs with a Voice:…”) in view of Hanson et al (20220199079) teaches the computer-implemented method of claim 1, wherein the corresponding CoT annotation for at least one of the plurality of conversational training samples (see Nachmani et al (“LMs with a Voice:…”) as applied to claim 1,2 above) is generated by a knowledge graph (see Hanson et al (20220199079), using a knowledge graph in determining dialog turns – para 0069, which would make the result more human-like, as noted above, in Hanson et al (20220199079)).
Claims 21, 22 are system claims that perform the steps in claims 10,11 above and as such, claims 21,22 are similar in scope and content to claims 10,11; therefore, claims 21,22 are rejected under similar rationale as presented against claims 10,11 above.
Response to Arguments
Applicant's arguments filed 6/15/2026 have been fully considered but they are not persuasive. Applicants replacement abstract is accepted, and the objection to the abstract is overcome, and the objection has been removed. As per applicants arguments against the prior art rejections, applicants arguments are toward the amended claim language of ‘ground truth chain of thought’. Examiner notes the introduction of the Manchandra et al (20240320595)
reference teaching ground-truth testing/improvement on the intermediate reasoning/CoT steps; see rejection above.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Please see related art listed on the PTO-892 form.
Furthermore, the following references were found containing features in applicants claims/specification:
Sun et al (20240249080) teaches ground truth labeling/testing on the Chain-of-Thought step (see para 0101)
Chen et al (20210280170) teaches cross entropy loss functions, forward and backward, in ASR systems (para 0046).
Qui et al (20220310080) teaches cross entropy loss (para 0051) in spoken dialog systems (para 0031)
Wang et al (20220383887) teaches cross entropy loss (para 0043, 0045) in editable live captioning systems (para 0002).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael Opsasnick, telephone number (571)272-7623, who is available Monday-Friday, 9am-5pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Mr. Richemond Dorvil, can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/Michael N Opsasnick/Primary Examiner, Art Unit 2658 09/15/2026