DETAILED ACTION
Introduction
1. This office action is in response to Applicant’s submission filed on 12/10/2024. Claims 1-15 are pending in the application and have been examined.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
3. The drawings filed on 12/10/2024 have been accepted and considered by the Examiner.
Claim Objections
4. Claims 1, 12, and 13 are objected to because of the following informalities: Said claims recite S100, S200A, S200B, S300A, S300B, S400, S500, and S600 which appears to referencing Fig. elements. Claim 3 also recites “, -like process.” Please, remove the recited references from said claims. Appropriate correction is required.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
5. Claim(s) 1-15 is/are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Thomson et al., (U.S. Patent Application Publication: 2020/0175987), cited in IDS filed 04/14/2026, and hereinafter referred to as THOMSON.
With respect to Claim 1, THOMSON discloses:
1. A method for generating accurate speech-to-text, STT, transcriptions in communication sessions, wherein the method comprises the steps of: - S100 establishing, by a first user, a communication session with at least one second user (See e.g., “…method may include obtaining first audio data originating at a first device during a communication session between the first device and a second device…”, “…communication session audio may include speech from one or more speakers participating in the communication session from other locations or using other communication devices such as on a conference communication session or an agent-assisted communication session…” See e.g., THOMSON, Abstract, paras. 299, 1136, Figs. 1, 4); - S200A capturing, by at least one Machine Learning, ML, model entity, audio data from a communication session (See e.g., “…estimated accuracy of transcriptions of audio generated by a first transcription unit or ASR system may be based on transcriptions of the audio generated by a second transcription unit or ASR system…methods may be used to estimate the accuracy of transcriptions. Embodiments describing how the accuracy tester 430 may generate the estimated accuracy…” and “…a language model other than the language model used by the decoder 510, such as a rescoring language model. A rescoring language model may, for example, be a neural net-based or an n-gram based language model. In some embodiments, the application information may include intelligence gained from user preferences or behaviors, syntax checks, rules…”; “…some embodiments, the ASR system 520 may have two language models, one for the decoder 510 and one for the rescorer 512… the model for the decoder 510 may include an n-gram based language model. The model for the rescorer 512 may include an RNNLM (recurrent neural network language model…” See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 1136, Figs. 1, 4, 5, 13); - S200B capturing, by at least one Natural Language Processing, NLP, entity, audio data from a communication session (See e.g., capabilities for capturing any combination of different models according to Fig. 5, comprising inter alia natural language processing, See e.g., paras. See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 1136, Figs. 1, 4, 5, 13); - S300A producing, by the at least one ML model entity time-stamped transcription chunks of the audio data (See e.g., “…errors of a transcription by the ASR system may then be provided to the error type model for estimating or predicting the reliability of the transcription for purposes of alignment and/or voting. A similar error type model may be determined for a pair of ASR systems, using the method described above for an ASR system and a reference transcription. In these and other embodiments, the error type model may be built for a given ASR system using a language modeling method based on, for example, n-grams, or using other machine learning methods such as neural networks…”; and “…transcription unit 1114 may also be configured to train at least one ASR system, for example, by training or updating models, using samples of the person's voice…may be speaker-dependent or speaker-independent…models that may be trained may include acoustic models, language models, lexicons, and runtime parameters or settings, among other models, including…transcription unit 1114 may include an ASR system 1120, a diarizer 1102, a voiceprints database 1104, an ASR model trainer 1122, and a speaker profile database 1106… diarizer 1102 may be configured to identify a device that generates audio for which a transcription is to be generated by the transcription unit 1114…” See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 306-309, 395, 1136, Figs. 1, 4, 5, 13); - S300B producing, by the at least one NLP entity, time-stamped transcription chunks of the audio data (See e.g., “…errors of a transcription by the ASR system may then be provided to the error type model for estimating or predicting the reliability of the transcription for purposes of alignment and/or voting. A similar error type model may be determined for a pair of ASR systems, using the method described above for an ASR system and a reference transcription. In these and other embodiments, the error type model may be built for a given ASR system using a language modeling method based on, for example, n-grams, or using other machine learning methods such as neural networks…”; “…transcription unit 1114 may also be configured to train at least one ASR system, for example, by training or updating models, using samples of the person's voice…may be speaker-dependent or speaker-independent…models that may be trained may include acoustic models, language models, lexicons, and runtime parameters or settings, among other models, including…transcription unit 1114 may include an ASR system 1120, a diarizer 1102, a voiceprints database 1104, an ASR model trainer 1122, and a speaker profile database 1106… diarizer 1102 may be configured to identify a device that generates audio for which a transcription is to be generated by the transcription unit 1114….” See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 306-309, 395, 1136, Figs. 1, 4, 5, 13); - S400 synchronizing and merging, by one or more chunk synchronization entity, the transcription chunks according to their time-stamps (See e.g., “the hypotheses may be represented as a string of tokens. The string of tokens may include one or more of sentences, phrases, or words. A token may include a word, subword, character, or symbol… fuser 1324…the fuser 1324 may be configured to merge the transcriptions generated by the ASR systems 1320 to create a fused transcription…the fused transcription may include an accuracy that is improved with respect to the accuracy of the individual transcriptions combined to generate the fused transcription…the fuser 1324 may generate multiple transcriptions…” and “…a synchronizer 1902, a first transcription unit 1914 a, and a second transcription unit 1914 b, collectively the transcription units 1914. The first transcription unit 1914 a may be a revoiced transcription unit. The second transcription unit 1914 b may be a non-revoiced transcription unit. Each of the transcription units 1914 may be configured to generate transcriptions from audio and provide the transcriptions to the synchronizer 1902…” See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 336, 337, 369, 511-515, 825, 1136, Figs. 1, 4, 5, 13); - S500 processing, by one or more transformer model processing entity, the synchronized and merged transcription chunks; - S600 ending the method (See e.g., “…outputs of the feature extractors 4004 may be communicated to the joint processor 4010. The joint processor 4010 may include components of an ASR system as described above with reference to FIG. 5, including to a feature transformer, probability calculator, rescorer, capitalizer, punctuator, and scorer, among others…” and “the hypotheses may be represented as a string of tokens. The string of tokens may include one or more of sentences, phrases, or words. A token may include a word, subword, character, or symbol… fuser 1324…the fuser 1324 may be configured to merge the transcriptions generated by the ASR systems 1320 to create a fused transcription…the fused transcription may include an accuracy that is improved with respect to the accuracy of the individual transcriptions combined to generate the fused transcription…the fuser 1324 may generate multiple transcriptions…” and “…a synchronizer 1902, a first transcription unit 1914 a, and a second transcription unit 1914 b, collectively the transcription units 1914. The first transcription unit 1914 a may be a revoiced transcription unit. The second transcription unit 1914 b may be a non-revoiced transcription unit. Each of the transcription units 1914 may be configured to generate transcriptions from audio and provide the transcriptions to the synchronizer 1902…” See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 336, 337, 369, 511-515, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 2, THOMSON discloses:
2. The method according to claim 1, wherein the at least one ML model entity and the at least one NLP entity run on different endpoints (See e.g., “… the CA profile may include one or more ASR modules that may be trained with respect to the speaker profile of the CA 118. The speaker profile may include models or links to models such as acoustic models and feature transformation models such as neural networks or MLLR or fMLLR transforms…” and “… ASR endpoints and/or text. For example, the audio delay 1 4740 a may receive endpoints and text from the transcription unit 4730 and audio delay 2 4740 b and audio delay 3 4740 c may receive text from text editor 1 4742 a and text editor 2 4742 b, respectively. When an audio delay 4740 receives text, the audio delay 4740 may use an ASR system to generate endpoints…” See e.g., THOMSON, Abstract, paras. 129, 232, 259, 260, 273, 299, 336, 337, 369, 511-515, 893, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 3, THOMSON discloses:
3. The method according to claim 1, wherein the at least one ML model entity is an automatic speech recognition, ASR, model, and the at least one NLP entity is a WebSpeech Application Programming Interface, API, -like process (See e.g., “…the first ASR system 420 a may be speaker-independent or speaker-dependent. The second ASR system 420 b may use the audio from the communication session to generate transcriptions. The second transcription unit 414 b may be configured in any manner described in this disclosure. For example, the second transcription unit 414 b may include an ASR system that is speaker-independent. In some embodiments, the ASR system may be an ASR service that the second transcription unit 414 b communicates with through an application programming interface (API) of the ASR service…” See e.g., THOMSON, Abstract, paras. 129, 232, 241, 259, 260, 273, 299, 336, 337, 369, 511-515, 893, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 4, THOMSON discloses:
4. The method according to claim 1, wherein the at least one ML model entity has context-awareness and allows domain-specific errors (See e.g., “…the list or context rules may be replaced by a natural language processor, a set of rules, or a model trained on data where profane and innocent terms have been labeled. In these and other embodiments, a function may be constructed that generates an output denoting whether the term is likely to be offensive…”; and how “…the processor may detect domain specific or application-specific terms or use knowledge of the domain to correct errors, format terms in a transcription, or configure a language model 511 for speech recognition…” See e.g., THOMSON, Abstract, paras. 129, 232, 241, 259, 260, 273, 274, 299, 336, 337, 369, 511-515, 893, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 5, THOMSON discloses:
5. The method according to claim 1, wherein the at least one NLP entity has an updatable domain-specific vocabulary/grammar (See e.g., how “…the processor may detect domain specific or application-specific terms or use knowledge of the domain to correct errors, format terms in a transcription, or configure a language model 511 for speech recognition…” See e.g., THOMSON, Abstract, paras. 129, 232, 241, 259, 260, 273, 274, 299, 336, 337, 369, 511-515, 893, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 6, THOMSON discloses:
6. The method according to claim 1, wherein multiple ML models are incorporated simultaneously (See e.g., “…method described above for an ASR system and a reference transcription. In these and other embodiments, the error type model may be built for a given ASR system using a language modeling method based on, for example, n-grams, or using other machine learning methods such as neural networks…” See e.g., THOMSON, Abstract, paras. 129, 232, 241, 259, 260, 273, 274, 299, 336, 337, 369, 511-515, 893, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 7, THOMSON discloses:
7. The method according to claim 1, wherein the method further comprises the step of updating, by one or more grammar/vocabulary update entity, the grammar/vocabulary, in case new grammar/vocabulary is detected (See e.g., “… an error type model, consider an example reference transcription (e.g., what was actually spoken) “Hermits have no peer pressure” and a hypothesis transcription (e.g., what the ASR system output) “Hermits no year is pressure…”; “…the align text process 1406 may be configured to align the tokens, such that subwords may be aligned as well as words. For example, the phrase “I don't want anything” may be transcribed by three ASR systems…” See e.g., THOMSON, Abstract, paras. 129, 232, 241, 259, 260, 273, 274, 299, 336, 337, 369, 393-398, 511-515, 893, 1136, Figs. 1, 4, 5, 13)
With respect to Claim 8, THOMSON discloses:
8. The method according to claim 7, wherein the step of updating the grammar/vocabulary is performed at various levels selected from user level, organization level, conversation/ channel level, and/or any combination of said levels, and/or wherein updating is performed in real-time (See e.g., “…errors of a transcription by the ASR system may then be provided to the error type model for estimating or predicting the reliability of the transcription for purposes of alignment and/or voting. A similar error type model may be determined for a pair of ASR systems, using the method described above for an ASR system and a reference transcription. In these and other embodiments, the error type model may be built for a given ASR system using a language modeling method based on, for example, n-grams, or using other machine learning methods such as neural networks…”; and “…transcription unit 1114 may also be configured to train at least one ASR system, for example, by training or updating models, using samples of the person's voice…may be speaker-dependent or speaker-independent…models that may be trained may include acoustic models, language models, lexicons, and runtime parameters or settings, among other models, including…transcription unit 1114 may include an ASR system 1120, a diarizer 1102, a voiceprints database 1104, an ASR model trainer 1122, and a speaker profile database 1106… diarizer 1102 may be configured to identify a device that generates audio for which a transcription is to be generated by the transcription unit 1114…” See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 306-309, 395, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 9, THOMSON discloses:
9. The method according to claim 7, wherein updates are added through interfaces and/or API calls (See e.g., “…the first ASR system 420 a may be speaker-independent or speaker-dependent. The second ASR system 420 b may use the audio from the communication session to generate transcriptions. The second transcription unit 414 b may be configured in any manner described in this disclosure. For example, the second transcription unit 414 b may include an ASR system that is speaker-independent. In some embodiments, the ASR system may be an ASR service that the second transcription unit 414 b communicates with through an application programming interface (API) of the ASR service…” See e.g., THOMSON, Abstract, paras. 129, 232, 241, 259, 260, 273, 299, 336, 337, 369, 511-515, 893, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 10, THOMSON discloses:
10. The method according to claim 1, wherein transcription chunks are collected from multiple endpoints (See e.g., “… the CA profile may include one or more ASR modules that may be trained with respect to the speaker profile of the CA 118. The speaker profile may include models or links to models such as acoustic models and feature transformation models such as neural networks or MLLR or fMLLR transforms…” and “… ASR endpoints and/or text. For example, the audio delay 1 4740 a may receive endpoints and text from the transcription unit 4730 and audio delay 2 4740 b and audio delay 3 4740 c may receive text from text editor 1 4742 a and text editor 2 4742 b, respectively. When an audio delay 4740 receives text, the audio delay 4740 may use an ASR system to generate endpoints…” See e.g., THOMSON, Abstract, paras. 129, 232, 259, 260, 273, 299, 336, 337, 369, 511-515, 893, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 11, THOMSON discloses:
11. The method according to claim 10, wherein the transcription chunks are synchronized and merged into time-frame windows of 1 to 60 seconds, 1 to 30 seconds, 1 to 15 seconds, 1 to 10 seconds, and/or 30 to 60 seconds, 45 to 60 seconds, 50 to 60 seconds; and/or all transcriptions chunks of a complete transcription are synchronized and merged (See e.g., “…outputs of the feature extractors 4004 may be communicated to the joint processor 4010. The joint processor 4010 may include components of an ASR system as described above with reference to FIG. 5, including to a feature transformer, probability calculator, rescorer, capitalizer, punctuator, and scorer, among others…” and “the hypotheses may be represented as a string of tokens. The string of tokens may include one or more of sentences, phrases, or words. A token may include a word, subword, character, or symbol… fuser 1324…the fuser 1324 may be configured to merge the transcriptions generated by the ASR systems 1320 to create a fused transcription…the fused transcription may include an accuracy that is improved with respect to the accuracy of the individual transcriptions combined to generate the fused transcription…the fuser 1324 may generate multiple transcriptions…” and “…a synchronizer 1902, a first transcription unit 1914 a, and a second transcription unit 1914 b, collectively the transcription units 1914. The first transcription unit 1914 a may be a revoiced transcription unit. The second transcription unit 1914 b may be a non-revoiced transcription unit. Each of the transcription units 1914 may be configured to generate transcriptions from audio and provide the transcriptions to the synchronizer 1902…” See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 336, 337, 369, 511-515, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 12, THOMSON discloses:
12. The method according to claim 1, wherein the method further comprises the step of employing overlapping time-frame windows before step S400, wherein said employing step comprises: - configuring, by the one or more overlapping windows entity, the level of overlapping; and - capturing, by the at least one ML model entity and the at least one NLP entity, different context (See e.g., “…for example, the selector 1806 may direct the first switch 1804 a to direct audio to the first transcription unit 1814 a and also direct the second switch 1804 b to direct audio to the second transcription unit 1814 b, in overlapping time periods. In these and other embodiments, both the first transcription unit 1814 a and the second transcription unit 1814 b receive the same audio at approximately the same or at the same time… both the first transcription unit 1814 a and the second transcription unit 1814 b may generate transcriptions and/or other data…”; and “… See e.g., a transcription unit that includes the CA client 3622 may be configured to transcribe the segment at overlapping time periods as the segment is transcribed using the ASR systems 3620. By streaming segments to CA clients 3622 based on predicted accuracy, round-trip latency to and from the CA clients 3622 may be reduced…segments may continue to stream to the CA clients 3622 until the predicted accuracy rises above the threshold…” THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 336, 337, 369, 498, 511-515, 744, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 13, THOMSON discloses:
13. The method according to claim 1, wherein the transcription chunks are evaluated in conjunction in step S500, leveraging grammar/vocabulary from the transcription chunks produced by the at least one NLP entity, and context from the transcription chunks, produced by the at least one ML model entity (See e.g., “… an error type model, consider an example reference transcription (e.g., what was actually spoken) “Hermits have no peer pressure” and a hypothesis transcription (e.g., what the ASR system output) “Hermits no year is pressure…”; “…the align text process 1406 may be configured to align the tokens, such that subwords may be aligned as well as words. For example, the phrase “I don't want anything” may be transcribed by three ASR systems…”; and see e.g., “…for example, the selector 1806 may direct the first switch 1804 a to direct audio to the first transcription unit 1814 a and also direct the second switch 1804 b to direct audio to the second transcription unit 1814 b, in overlapping time periods…both the first transcription unit 1814 a and the second transcription unit 1814 b receive the same audio at approximately the same or at the same time… both the first transcription unit 1814 a and the second transcription unit 1814 b may generate transcriptions and/or other data…”; and how “… see e.g., a transcription unit that includes the CA client 3622 may be configured to transcribe the segment at overlapping time periods as the segment is transcribed using the ASR systems 3620. By streaming segments to CA clients 3622 based on predicted accuracy, round-trip latency to and from the CA clients 3622 may be reduced…segments may continue to stream to the CA clients 3622 until the predicted accuracy rises above the threshold…” See e.g., THOMSON, Abstract, paras. 129, 232, 241, 259, 260, 273, 274, 299, 336, 337, 369, 393-398, 511-515, 744, 893, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 14, THOMSON discloses:
14. A system for generating accurate speech-to-text transcription in communication sessions, wherein the system is configured to perform the method according to claim 1 (See e.g., “…method may include obtaining first audio data originating at a first device during a communication session between the first device and a second device…”, “…communication session audio may include speech from one or more speakers participating in the communication session from other locations or using other communication devices such as on a conference communication session or an agent-assisted communication session…” THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 336, 337, 369, 498, 511-515, 744, 1136, Figs. 1, 4, 5, 13).
With respect to Claim 15, THOMSON discloses:
15. The system according to claim 14 wherein the system comprises: - at least one ML model entity; - at least one NLP entity; - one or more grammar/vocabulary update entity; - one or more chunk synchronization entity; - one or more overlapping windows entity; - one or more transformer model processing entity; and/or - a media server (See e.g., “…outputs of the feature extractors 4004 may be communicated to the joint processor 4010. The joint processor 4010 may include components of an ASR system as described above with reference to FIG. 5, including to a feature transformer, probability calculator, rescorer, capitalizer, punctuator, and scorer, among others…” and “the hypotheses may be represented as a string of tokens. The string of tokens may include one or more of sentences, phrases, or words. A token may include a word, subword, character, or symbol… fuser 1324…the fuser 1324 may be configured to merge the transcriptions generated by the ASR systems 1320 to create a fused transcription…the fused transcription may include an accuracy that is improved with respect to the accuracy of the individual transcriptions combined to generate the fused transcription…the fuser 1324 may generate multiple transcriptions…” and “…a synchronizer 1902, a first transcription unit 1914 a, and a second transcription unit 1914 b, collectively the transcription units 1914. The first transcription unit 1914 a may be a revoiced transcription unit. The second transcription unit 1914 b may be a non-revoiced transcription unit. Each of the transcription units 1914 may be configured to generate transcriptions from audio and provide the transcriptions to the synchronizer 1902…” See e.g., THOMSON, Abstract, paras. 232, 259, 260, 273, 299, 336, 337, 369, 511-515, 825, 1136, Figs. 1, 4, 5, 13).
Conclusion
6. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Szymański et al., (Szymański et al., 2023. Why Aren’t We NER Yet? Artifacts of ASR Errors in Named Entity Recognition in Spontaneous Speech Transcripts. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1746–1761, Toronto, Canada), discloses an architecture comprising, how, see e.g., “…transcripts of spontaneous human speech present a significant obstacle for traditional NER models. The lack of grammatical structure of spoken utterances and word errors introduced by the ASR make downstream NLP tasks challenging. In this paper, we examine in detail the complex relationship between ASR and NER errors which limit the ability of NER models to recover entity mentions from spontaneous speech transcripts…” (See e.g., Szymański et al., Abstract).
Please, see PTO-892 for more details.
7. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Edgar Guerra-Erazo whose telephone number is (571) 270-3708. The examiner can normally be reached on M-F 7:30a.m.-5:00p.m. EST. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Bhavesh Mehta can be reached on (571) 272-7453. The fax phone number for the organization where this application or proceeding is assigned is (571) 273-8300.
8. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at
http://www.uspto.gov/interviewpractice.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/EDGAR X GUERRA-ERAZO/Primary Examiner, Art Unit 2656