DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 5, 8, 10, 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Chollampatt Muhammed Ashraf et al. (US 2023/0351123 A1, “Ashraf”) in view of Chevrier et al. (US 2019/0228774 A1, “Chevrier”).
As to claims 1, 8, Ashraf discloses one or more non-transitory computer-readable media that include stored thereon computer-executable instructions that when executed by at least a processor of a computing system cause the computing system to:
preemptively establish a set of live connections to an automatic speech recognition service that are available for use, wherein the set of live connections includes fewer live connections than a total K of participants in a virtual meeting (pool of available transcription and translation processes 314, 316, para. 0065-0070);
in response to a participant of the virtual meeting becoming active, dedicate one live connection from the set of live connections to real-time transcription of an individual audio stream from the participant (for each client device that requests translation, the virtual conference provider 310 will allocate the appropriate transcription and translation processes 314, 316, para. 0069);
in real-time, label transcription results received back through the one live connection with a username of the participant (virtual conference provider may provide real-time translation with participant identification, para. 0071); and
in real-time, inject the labeled transcription results back into the virtual meeting for display in a user interface of the virtual meeting (translated text is provided as closed captions or in a separate area of the requesting client’s GUI, para. 0067).
Ashraf differs from claims 1, 8 in that it does not disclose the above underlined limitation.
Chevrier teaches establishing a network connection with a transcription system in response to a voice becoming active (para. 0040). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ashraf with the above teaching of Chevrier in order to establish a connection with a transcription system only when needed, that is, when there is a voice input to be transcribed.
As to claims 5, 11, Ashraf in view of Chevrier teaches: close live connections that are in excess of a baseline count C of live connections and which have been available for use longer than a threshold amount of time T (Chevrier: a network connection with the transcription system may be terminated in response to the voice being inactive for longer than a time period threshold, para. 0039, 0055).
As to claim 10, Ashraf in view of Chevrier teaches: in response to the participant of the virtual meeting becoming inactive, releasing the one live connection from dedication to the participant back into the set of live connections that are available for use (Ashraf: de-allocated connections are returned to the pool, para. 0070).
As to claim 14, Ashraf in view of Chevrier teaches: wherein the real-time transcription includes translation from a first human language to a second human language, wherein speech in the individual audio stream is in the first human language, and the transcription results are in the second human language (Ashraf: translation from a first to a second language, Abstract).
Claim(s) 2, 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ashraf in view of Chevrier, as applied to claims 1, 8 above, and further in view of Shin (US 2025/0023930 A1).
Ahsraf in view of Chevrier differs from claims 2, 13 in that it does not teach:
monitor a mute/unmute status of the participant to determine when the participant becomes active;
in response to the mute/unmute status changing from mute to unmute, allocate the one live connection for sole use by the participant and connect the individual audio through the one live connection to an individual session of the automatic speech recognition service; and
in response to the mute/unmute status changing from unmute to mute, disconnect the individual audio stream from the one live connection, and deallocate the one live connection back to the set of live connections that are available for use.
Shin teaches providing captured audio data to a transcription engine in response to detecting an unmute button (para. 0042). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ashraf in view of Chevrier with the above teaching of Shin in order to determine an active speaker.
Claim(s) 3, 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ashraf in view of Chevrier, as applied to claims 1, 8 above, and further in view of Togami et al. (US 2024/0096346 A1, “Togami”).
Ashraf in view of Chevrier differs from claims 3, 9 in that it does not teach:
associate a session ID of the one live connection with a user ID of the participant; and
send the individual audio stream of the participant to the automatic speech recognition service through the one live connection to cause the automatic speech recognition service to transcribe speech from the individual audio stream into the transcription results in real-time, wherein audio streams of other participants are not sent through the one live connection.
Togami teaches providing a plurality of single-talker audio streams to a transcription system which generates a plurality of corresponding single-talker transcriptions, each single-talker transcription being a transcription of speech only from a respective talker, and no other talkers (para. 0014). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ashraf in view of Chevrier with the above teaching of Togami in order to clearly identify each talker.
Claim(s) 6, 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ashraf in view of Chevrier, as applied to claims 1, 8 above, and further in view of D’Amore et al. (US 2019/0005047 A1, “D’Amore”)
Ashraf in view of Chevrier differs from claims 6, 12 in that it does not teach: expand the set of live connections to the automatic speech recognition service by preemptively establishing additional live connections when the live connections that are available for use falls to a threshold number.
D’Amore teaches increasing a number of connections available in a connection pool prior to the number of available connections dropping below a particular minimum (para. 0071). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ashraf in view of Chevrier with the above teaching of D’Amore in order to increase system efficiency and reduce connection delays, as taught by D’Amore (para. 0023, 0072).
Claim(s) 7, 15, 17-18, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ashraf in view of Chevrier, as applied to claim 1 above, and further in view of Pho et al. (US 2025/0118297 A1, “Pho”).
Ashraf in view of Chevrier differs from claims 7, 15 in that it does not teach: wherein the live connections to the automatic speech recognition service are WebSocket connections.
Pho teaches the well known use of WebSocket connection as an immediate linkage to a low-latency speech-to-text service for transcribing telephonic audio data (para. 0069, 0073). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ashraf in view of Chevrier with the above teaching of Pho in order to use a well known WebSocket connection for establishing a bidirectional connection.
As to claim 17, Ashraf in view of Chevrier and Pho teaches: in response to the participant of the virtual meeting becoming inactive, releasing the one WebSocket connection from dedication to the participant back into the set of WebSocket connections that are available for use (Ashraf: de-allocated connections are returned to the pool, para. 0070).
As to claim 18, Ashraf in view of Chevrier and Pho teaches: close WebSocket connections that are in excess of a baseline count C of WebSocket connections and which have been available for use longer than a threshold amount of time T (Chevrier: a network connection with the transcription system may be terminated in response to the voice being inactive for longer than a time period threshold, para. 0039, 0055).
As to claim 20, Ashraf in view of Chevrier and Pho teaches: cause the computing system to join the virtual meeting as an additional participant to obtain the individual audio stream input by the participant (Ashraf: client device may join an existing meeting and participate using audio devices, para. 0026, 0029).
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ashraf in view of Chevrier and Pho, as applied to claim 15 above, and further in view of Togami.
Ashraf in view of Chevrier and Pho differs from claim 16 in that it does not teach:
associate a session ID of the one WebSocket connection with a user ID of the participant; and
send the individual audio stream of the participant to the automatic speech recognition service through the one WebSocket connection to cause the automatic speech recognition service to transcribe speech from the audio stream into the transcription results in real-time, wherein audio streams of other participants are not sent through the one WebSocket connection.
Togami teaches providing a plurality of single-talker audio streams to a transcription system which generates a plurality of corresponding single-talker transcriptions, each single-talker transcription being a transcription of speech only from a respective talker, and no other talkers (para. 0014). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ashraf in view of Chevrier and Pho with the above teaching of Togami in order to clearly identify each talker.
Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ashraf in view of Chevrier and Pho, as applied to claims 1, 8 above, and further in view of D’Amore.
Ashraf in view of Chevrier and Pho differs from claim 19 in that it does not teach: expand the set of WebSocket connections to the automatic speech recognition service by preemptively establishing additional WebSocket connections when the WebSocket connections that are available for use falls to a threshold number.
D’Amore teaches increasing a number of connections available in a connection pool prior to the number of available connections dropping below a particular minimum (para. 0071). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ashraf in view of Chevrier and Pho with the above teaching of D’Amore in order to increase system efficiency and reduce connection delays, as taught by D’Amore (para. 0023, 0072).
Allowable Subject Matter
Claim 4 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Anderson et al. (US 2002/0065915 A1) teach increasing the number of available communication connections in a pool if the number falls below a desired amount.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Stella L Woo whose telephone number is (571)272-7512. The examiner can normally be reached Monday - Friday, 8 a.m. to 5 p.m.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ahmad Matar can be reached at 571-272-7488. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Stella L. Woo/ Primary Examiner, Art Unit 2693