DETAILED ACTION
Response to Amendment
The Applicant’s Amendment filed 06/19/2026 has been entered. Claims 21-40 are pending in the Application.
Response to Arguments
Applicant's Terminal Disclaimer filed 06/19/2026 has overcome the Double Patenting rejection. Accordingly, the Double Patenting rejection has been withdrawn.
Applicant's arguments filed 06/19/2026 with respect to the prior art rejection have been fully considered but they are not persuasive.
Regarding claims 21, 28 and 34, the Applicant argues that the cited art Chavez (US 20220060525) fails to teach the newly amended limitation “causing the first computing device to determine a first action of the first set of actions to perform” and “causing the first computing device to determine a second action of the second set of actions to perform”. The Applicant submits that Chavez makes determinations purely based on server-side processing thus does not teach or suggests that the endpoints perform any determinations of actions for the endpoints to perform. The Examiner respectfully disagrees. Contrary to the Applicant’s submission, Chavez explicitly discloses that the muting/unmuting decisions can be made either on server-side or at the endpoint (see para 0096, Muting may be performed automatically by a processor of a server, such as the server 110 providing the conference content, or by a signal to the particular endpoint 104 to execute a mute circuit that, when received by the associated participants 102, performs the muting action). Chavez further explicitly discloses a signal causing the endpoint to determine an action and executing the action at the endpoint (see para 0090, the server 110 sends a muting notification/action signal 208 to the endpoint 104B and, in response, the endpoint 104B activates a notification circuit or logic to prompt the participant 102B to manually activate a muting feature of the endpoint 104B). Therefore, Chavez discloses the argued limitation as claimed.
Based on the reasoning above, the rejection should be maintained. Please see below for the detailed rejections.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 21-40 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chavez et al US 20220060525.
Regarding claim 21, Chavez teaches a system comprising:
at least one processor (see para 0009, a system is provided to achieve an intelligent muting/unmuting of endpoints, which may be performed by a microprocessor(s)); and
memory storing instructions that, when executed by the at least one processor, causes the system to perform a set of operations (see para 0059, a computer-readable storage medium… that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device), the set of operations comprising:
receiving, by the system, an interaction intent metric for a user of a first computing device, wherein the user is a communication participant of a communication session (see para 0010-0013, receiving intent metric such as: the movement of the participant's lips, other facial features involved in speech, direction of his gaze (e.g., towards vs. away from endpoint, elsewhere, etc.), and/or facial expressions);
when it is determined that the interaction intent metric exceeds a positive intent threshold and input for the first computing device to the communication session is muted (see para 0016-0017, upon determining an active participant is speaking on mute and based on the confidence score that the participant is active speaking in the conference):
identifying a first set of actions associated with a positive intent to engage in the communication session (see para 0019-0021, set of actions including visual indicator, audible announcement or automatically unmute); causing the first computing device to determine a first action of the first set of actions to perform and performing a first action of the first set of actions with respect to the first computing device of the participant (see para 0096, Muting may be performed automatically by a processor of a server, such as the server 110 providing the conference content, or by a signal to the particular endpoint 104 to execute a mute circuit that, when received by the associated participants 102, performs the muting action); and
when it is determined that the interaction intent metric is below a negative intent threshold and input for the first computing device to the communication session is not muted (see para 0022, When a determination is made that audio provided, while the endpoint is not on mute, is not for inclusion in the conference):
identifying a second set of actions associated with a negative intent to engage in the communication session (see para 0024-0026, set of actions including visual indicator, audible announcement or automatically mute); causing the first computing device to determine a second action of the second set of actions to perform and performing a second action of the second set of actions with respect to the first computing device of the participant (see para 0096, muting may be performed automatically by a processor of a server, such as the server 110 providing the conference content, or by a signal to the particular endpoint 104 to execute a mute circuit that, when received by the associated participants 102, performs the muting action).
Regarding claim 22, Chavez further teaches displaying at least a part of the identified first set of actions or the identified second set of actions for selection by the user (see para 0090, the server 110 sends a muting notification/action signal 208 to the endpoint 104B and, in response, the endpoint 104B activates a notification circuit or logic to prompt the participant 102B to manually activate a muting feature).
Regarding claim 23, Chavez further teaches providing, to another computing device associated with the communication session, an indication of at least a part of the identified first set of actions or the identified second set of actions for presentation to a user of the another computing device (see para 0090, the server 110 sends a muting notification/action signal 208 to the endpoint 104B).
Regarding claim 24, Chavez further teaches the first set of actions includes providing a recommendation to unmute input for the communication session (see para 0019-0021, trigger notification or automatically unmute); and the second set of actions includes automatically muting input for the communication session (see para 0024-0026, trigger notification or automatically mute).
Regarding claim 25, Chavez further teaches the input includes at least one of: audio input from an audio device associated with the computing device; and video input from a video device associated with the computing device (see para 0029, the incoming real time stream (e.g., video and audio) from a participant's endpoint).
Regarding claim 26, Chavez further teaches performing the first action or the second action in response to an indication received from a computing device of a participant of the communication session (see para 0090, the server 110 sends a muting notification/action signal 208 to the endpoint 104B and, in response, the endpoint 104B activates a notification circuit or logic to prompt the participant 102B to manually activate a muting feature of the endpoint 104B).
Regarding claim 27, Chavez further teaches the first action is automatically performed in response to determining that the interaction intent metric exceeds the positive intent threshold and input for the first computing device to the communication session is muted (see para 0021, Automatically unmute the participant's audio); or the second action is automatically performed in response to determining that the interaction intent metric is below the negative intent threshold and input for the first computing device to the communication session is not muted (see para 0026, Automatically mute the participant's endpoint).
Regarding claim 28, Chavez teaches a method for managing a communication session, the method comprising:
obtaining, by a server device of the communication session, an interaction intent metric for a first computing device of a communication session (see figure 3B, server 110, see para 0010-0013, obtaining intent metric such as: the movement of the participant's lips, other facial features involved in speech, direction of his gaze (e.g., towards vs. away from endpoint, elsewhere, etc.), and/or facial expressions);
when it is determined that the interaction intent metric exceeds an intent threshold and therefore does not match an input state for the first computing device (see para 0016-0017, upon determining an active participant is speaking on mute and based on the confidence score that the participant is active speaking in the conference):
identifying a set of actions associated with the interaction intent metric as it relates to the intent threshold (see para 0019-0021, trigger notification or automatically unmute); and automatically transmitting, by the server device to the first computing device, an indication of an action of the set of, causing the first computing device to determine an action of the set of actions to perform and perform the action of the set of actions (see para 0090, the server 110 sends a muting notification/action signal 208 to the endpoint 104B and, in response, the endpoint 104B activates a notification circuit or logic to prompt the participant 102B to manually activate a muting feature of the endpoint 104B, also see para 0096, muting may be performed automatically by a processor of a server, such as the server 110 providing the conference content, or by a signal to the particular endpoint 104 to execute a mute circuit that, when received by the associated participants 102, performs the muting action).
Regarding claim 29, Chavez further teaches providing, by the server and to the first computing device, a recommendation to unmute input for the communication session when the interaction intent metric exceeds a positive intent threshold (see para 0019-0021, trigger notification or automatically unmute); or automatically muting the first computing device when the interaction intent metric exceeds a negative intent threshold (see para 0024-0026, trigger notification or automatically mute).
Regarding claim 30, Chavez further teaches providing, to another computing device associated with the communication session, an indication associated with the mismatched input state of the first computing device (see para 0114, Video conferencing system determines whether the muting on the endpoint 104B is in error. In some embodiments, analysis of the video and/or audio portion contributed from the endpoint 104C may result in the determination that the participant 102B is attempting to speak to the video conference 800. For example, based on analysis of the video portion contributed from the endpoint 102B, video conferencing system may determine that the gaze of the participant 102B is directed at the endpoint 104B, and that the mouth/lips/other facial features of the participant 102B are moving. Additionally or alternatively, NLP may be used to determine a question requiring a spoken response was directed to the participant 102B. The video conferencing system sends an alert 804B (e.g., a tone, message, pop-up visual indicator, etc.) to the participant 102B/the endpoint 104B to unmute the erroneously muted participant 102B/the endpoint 104B).
Regarding claim 31, Chavez further teaches the interaction intent metric is generated using a machine learning model trained to process a set of factors associated with the communication session (see para 0032, the analyzing the participants' contributed audio and/or video using NLP/Artificial Intelligence (AI), which may also include machine learning, deep learning, or other machine intelligence and voice recognition techniques to make a determination that the user is not speaking in the video conference); and the machine learning model was trained based at least in part on training data associated with at least one of the communication participant or a population of users (see para 0028, the data gathered as described above, may then be used to train one or more Machine Learning (ML) models).
Regarding claim 32, Chavez further teaches the set of factors includes:
a semantic factor associated with audio input (see para 0012, determine the context of the sentence);
a linguistic factor associated with audio input (see para 0013, analyzed for audio characteristics such as intensity/loudness, pitch, tone, etc);
a video input factor (see para 0010, analyzing the video portion of the media stream);
a meeting context associated an ongoing communication session (see para 0013, Other data, such as participant rooster, conference agenda, etc.); and
a historical user characteristic factor associated with a communication participant of the ongoing communication session (see para 0086, the AI Driven Facial Movement Recognition and Analysis module might employ one or more AI Vision libraries which will be trained with numerous samples of the human facial structure and facial characteristics).
Regarding claim 33, Chavez further teaches the interaction intent metric is received from the first computing device (see para 0010, analyzing the video portion of the media stream received from an endpoint to determine whether the participant in the video portion is actively speaking or not speaking).
Regarding claim 34, Chavez teaches a method for managing a communication session, the method comprising:
obtaining an interaction intent metric for a user of a first computing device, wherein the user is a communication participant of a communication session (see para 0010-0013, receiving intent metric such as: the movement of the participant's lips, other facial features involved in speech, direction of his gaze (e.g., towards vs. away from endpoint, elsewhere, etc.), and/or facial expressions);
when it is determined that the interaction intent metric exceeds an intent threshold and therefore does not match an input state for the first computing device (see para 0016-0017, upon determining an active participant is speaking on mute and based on the confidence score that the participant is active speaking in the conference):
generating a set of actions based on the interaction intent metric as it relates to the intent threshold (see para 0024-0026, set of actions including visual indicator, audible announcement or automatically mute); causing the first computing device to determine an action of the set of actions to perform; and performing an action of the set of actions with respect to the first computing device of the participant (see para 0096, muting may be performed automatically by a processor of a server, such as the server 110 providing the conference content, or by a signal to the particular endpoint 104 to execute a mute circuit that, when received by the associated participants 102, performs the muting action).
Regarding claim 35, Chavez further teaches displaying at least a part of the set of actions for selection by the user (see para 0090, the server 110 sends a muting notification/action signal 208 to the endpoint 104B and, in response, the endpoint 104B activates a notification circuit or logic to prompt the participant 102B to manually activate a muting feature).
Regarding claim 36, Chavez further teaches providing, to another computing device associated with the communication session, an indication of at least a part of the set of actions for presentation to a user of the another computing device (see para 0090, the server 110 sends a muting notification/action signal 208 to the endpoint 104B).
Regarding claim 37, Chavez further teaches providing a recommendation to unmute input for the communication session (see para 0019-0021, trigger notification or automatically unmute) or automatically muting input for the communication session (see para 0024-0026, trigger notification or automatically mute).
Regarding claim 38, Chavez further teaches the input includes at least one of: audio input from an audio device associated with the computing device; and video input from a video device associated with the computing device (see para 0029, the incoming real time stream (e.g., video and audio) from a participant's endpoint).
Regarding claim 39, Chavez further teaches performing the action in response to an indication received from a computing device of a participant of the communication session (see para 0090, the server 110 sends a muting notification/action signal 208 to the endpoint 104B and, in response, the endpoint 104B activates a notification circuit or logic to prompt the participant 102B to manually activate a muting feature of the endpoint 104B).
Regarding claim 40, Chavez further teaches the action is automatically performed in response to determining that the interaction intent metric exceeds the intent threshold and therefore does not match the input state for the first computing device (see para 0021, Automatically unmute the participant's audio).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Sun et al US 20150156598 discloses microphone muting/unmuting decision in a localized endpoint setting.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHONG H DANG whose telephone number is (571)272-0470. The examiner can normally be reached Monday-Friday 9:30AM - 6:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henry Tsai can be reached at (571)272-4176. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHONG H DANG/Primary Examiner, Art Unit 2184