DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Regarding claim 1, the phrase " in response to receiving user responses to unmute and play the recording " renders the claim indefinite because it is unclear whether the limitation(s) following the phrase are part of the claimed invention. See MPEP § 2173.05(d).
Claim 1 recites the limitation " presenting a notification to the user conveying the user is on mute while presenting to the one or more other participants " and “the user responses to unmute and play recording..”. There is insufficient antecedent basis for this limitation in the claim.
Contingent Limitations
Section MPEP 2111.04(II) sets forth, “The broadest reasonable interpretation of a method (or process) claim having contingent limitations requires only those steps that must be performed and does not include steps that are not required to be performed because the condition(s) precedent are not met.” The following are contingent limitations which are not required to be found in the prior art under broadest reasonable interpretation. This does not necessarily include contingent limitations which are required to be found in the prior art.
Claim 1 recites the following contingent limitations “in response to receiving user responses to unmute and play the recording, unmuting an audio feed from the user to the one or more other participants and playing the recording for the one or more other participant”. Under the BRI of a method claim, the steps of “unmuting an audio feed from the user to the one or more other participants and playing the recording” are contingent limitations. The contingent limitations is for the compact prosecution but is not required by the prior art to read on the claim.
Claim 4 recites the following contingent limitations “in response to receiving a user response within the user responses to display the speech-to-text translation of the recording”. Under the BRI of a method claim, the steps of “in response to receiving a user response within the user responses to display the speech-to-text translation of the recording” are contingent limitations. The contingent limitations is for the compact prosecution but is not required by the prior art to read on the claim.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Deole et al. (US Patent Publication Application No. 2021/0359872) in view of Finlow-Bates et al. (US Patent Application Publication No. 2016/0055859).
Regarding claim 1, Deole discloses a processor-implemented method, the method comprising [see para. 0108; Components attached to network include memory, data storage, input/output device(s), and/or other components that may be accessible to processor]:
in response to determining a user is muted while speaking to one or more other participants of a web conference [see para. 0006; Electronic conferences or meetings, with at least two participants or groups of participants communicating via communication endpoints over a network (herein, “conference”) are common in business and other settings. Unfortunately, it is also common to have a speaker talking but without realizing they're on mute, resulting in confusion and wasted time and continuity of the conference.], capturing a recording of audio received by a microphone associated with the user [see para. 0008; a system is provided to recognize the fact that the speaker is speaking on mute and intelligently take action and/or a system that recognizes the fact that sound (e.g., an extraneous conversation), not relevant to the conference, is being picked up and included in the conference and similarly automatically taking action before any manual intervention is required to reduce the extraneous sound within the conference; which corresponds to a recording made of the audio portion received 304 while on mute and replayed into the conference];
presenting a notification to the user conveying the user is on mute while presenting to the one or more other participants, wherein the notification includes two or more prompts which require a user response [see para. 0087-0088; Replaying all or a portion of speech re-prompts user 102A to provide a response. If user did provided a response, such as while on mute, a recording made of response speech received while on mute and replayed into the conference. For example, user begin providing speech, by saying a word or two (e.g., “For the . . . ”), while endpoint is on mute. After endpoint is unmuted, server buffer the words provided after endpoint is unmuted and the recorded speech followed by the buffered speech played back into the conference as conference content until speech is live. If the portion of speech provided during muting is more than a few words (e.g., more than ten seconds), then user prompted to either initiate the playback the portion of speech provided while on mute or repeating speech again; which corresponds to notify the participant 102A that they are on mute and prompt the participant 102A to manually unmute the endpoint];
receiving the user response to each prompt [see para. 0088-0089; unmuting notification/action automatically unmute endpoint 104A to provide speech as a portion of the conference content, unmuting notification/action further include signaling endpoint wherein the signal causes a notification to be presented by only endpoint, that they are off mute (e.g., tone, message, pop-up message, etc.). As a further option, all endpoints notified of the on-mute/off-mute state of endpoints and, when changed, each endpoint is updated accordingly, such as with a message (e.g., “Alice is on mute” or “Alice is off mute.”) or graphical icon having a meaning associated with the muting state. Optionally, speech buffered and replayed as conference content, so that any speech provided before the unmuting notification/action results in the unmuting of endpoint is provided as uninterrupted speech but with a delay determined by the beginning of speech and the occurrence of the unmuting action]; however, Deole fails to explicitly teach in response to receiving user responses to unmute and play the recording, unmuting an audio feed from the user to the one or more other participants and playing the recording for the one or more other participants.
Finlow-Bates discloses in response to receiving user responses to unmute and play the recording, unmuting an audio feed from the user to the one or more other participants and playing the recording for the one or more other participants [see para. 0028-0031; the communication device processor configured with software to automatically detect when to turn off the mute function and initiate playback, such as by recognizing when the caller is speaking, communication device configured with a smart un-mute user interface input (e.g., a virtual button or icon) to enable the caller to select when to playback a portion of the recorded audio to the other participants. Such a smart un-mute user interface also enable the user to designate the duration or portion of the buffered audio to be played back; which corresponds to a smart mute for recovering speech spoken while a communication device is muted during a voice call].
It would have been obvious to one of an ordinary skill in the art, having the teachings of Deole and FInlow-Bates before the affective filing date of the claimed invention to modify, the notification and recorvery process of Deole to include multiple user prompts requiring responses as taught by Finlow-Bates.
One would have been motivated to make this modification to provide the user with control over the recovered speech is injected into the conference, thereby improving user experience and reducing risk of unwanted playback of private or unintended speech .
Regarding claims 2 and 9, Finlow-Bates discloses wherein the two or more prompts comprise requesting whether the user wishes to unmute and whether the user wishes for playback of the recording [see para. 0028-0031; the communication device processor configured with software to automatically detect when to turn off the mute function and initiate playback, such as by recognizing when the caller is speaking, communication device configured with a smart un-mute user interface input (e.g., a virtual button or icon) to enable the caller to select when to playback a portion of the recorded audio to the other participants. Such a smart un-mute user interface also enable the user to designate the duration or portion of the buffered audio to be played back; which corresponds to a smart mute for recovering speech spoken while a communication device is muted during a voice call].
Regarding claims 3 and 10, Finlow-Bates discloses wherein the two or more prompts further comprise whether to display a speech-to-text translation of the recording to the one or more other participants [see para. 0058; The predetermined conditions may include an association with at least one of a key word, a language, a context of the active voice call, a recognized voice, a sensor input of the communication device, and/or other considerations used to analyze the buffered input audio. A speech-to-text function performed on the buffered audio data by the processor may identify words within in the buffered input audio for analysis. For example, a speech-to-text function may help determine that the active voice call involved only a single language, while the initially muted input audio included words from a different language].
Regarding claims 4 and 11, Finlow-Bates discloses further comprising: performing a translation of the recording to text using a speech-to-text technology; and in response to receiving a user response within the user responses to display the speech-to-text translation of the recording, displaying the translation on a device display screen associated with each of the one or more other participants and the user [see para. 0058, 0110; Identification of the use of particular key words (listed in the key word module 734) or particular voices (identified through voice recognition module) a predetermined condition trigger the smart mute feature in addition to turning off the mute function. The key word module operate in conjunction with the speech-to-text module. The speech-to-text module decode and interpret the conversations/voice samples from all participants in real time. Identified high frequency words may be added to the key word module list]. It would have been well known and used in web-conferencing environments.
Regarding claims 5 and 12, Deole discloses wherein capturing the recording is further performed only in response to also determining a user gaze is focused on a graphical user interface of a web conferencing software [see para. 0098 and figure 6; Speech provided by humans, such as a particular participant providing speech for inclusion in a conference content, versus speech provided to other, non-conference content, different in terms of speech attributes. For example, one speaking to a group of remote conference participants have a particular manner of speaking that differs when speaking to a colleague or other party face-to-face. These manners quantified as various speech attributes and, utilized to determine whether speech provided by the participant is or is not intended for inclusion into the conference content].
Regarding claims 6 and 13, Deole discloses wherein capturing the recording is further performed only in response to also determining speech [see para. 0097; The determination that the muting is in error performed by test is variously embodied, a preceding portion of the conference content, such as provided by a different endpoint addressed the participant associated with the particular endpoint, such as by name, role, location, etc. An attribute of the speech provided in the audio from the particular endpoint matches an attribute of speech, within a previously determined threshold, of prior speech from the participant when known to be providing speech intended to be included in the conference content].
Finlow-Bates discloses emitted by the user, is intended for the one or more other participants [see para. 0058; he analysis of the buffered input audio in block used by the processor in determining whether the mute function is supposed to be on, based on predetermined conditions identified from the buffered audio itself, and possibly other inputs. The predetermined conditions include an association with at least one of a key word, a language, a context of the active voice call, a recognized voice, a sensor input of the communication device, and/or other considerations used to analyze the buffered input audio. A speech-to-text function performed on the buffered audio data by the processor identify words within in the buffered input audio for analysis. For example, a speech-to-text function help determine that the active voice call involved only a single language, while the initially muted input audio included words from a different language].
One would have been motivated to make this modification to provide the user with control over the recovered speech is injected into the conference, thereby improving user experience and reducing risk of unwanted playback of private or unintended speech.
Regarding claims 7 and 14, Deole discloses wherein playing the recording to the one or more other participants further comprises concurrently playing the recording to the user [see para. 0098; speech provided by humans, such as a particular participant providing speech for inclusion in a conference content, versus speech provided to other, non-conference content, different in terms of speech attributes. For example, one speaking to a group of remote conference participants may have a particular manner of speaking that differs when speaking to a colleague or other party face-to-face. These manners quantified as various speech attributes and, utilized to determine whether speech provided by the participant is or is not intended for inclusion into the conference content].
Regarding claims 8 and 15, Deole discloses a computer system, the computer system comprising: one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising [see para. 0060; computer-readable storage medium, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory]:
in response to determining a user is muted while speaking to one or more other participants of a web conference [see para. 0006; Electronic conferences or meetings, with at least two participants or groups of participants communicating via communication endpoints over a network (herein, “conference”) are common in business and other settings. Unfortunately, it is also common to have a speaker talking but without realizing they're on mute, resulting in confusion and wasted time and continuity of the conference.], capturing a recording of audio received by a microphone associated with the user [see para. 0008; a system is provided to recognize the fact that the speaker is speaking on mute and intelligently take action and/or a system that recognizes the fact that sound (e.g., an extraneous conversation), not relevant to the conference, is being picked up and included in the conference and similarly automatically taking action before any manual intervention is required to reduce the extraneous sound within the conference; which corresponds to a recording made of the audio portion received 304 while on mute and replayed into the conference];
presenting a notification to the user conveying the user is on mute while presenting to the one or more other participants, wherein the notification includes two or more prompts which require a user response [see para. 0087-0088; Replaying all or a portion of speech re-prompts user 102A to provide a response. If user did provided a response, such as while on mute, a recording made of response speech received while on mute and replayed into the conference. For example, user begin providing speech, by saying a word or two (e.g., “For the . . . ”), while endpoint is on mute. After endpoint is unmuted, server buffer the words provided after endpoint is unmuted and the recorded speech followed by the buffered speech played back into the conference as conference content until speech is live. If the portion of speech provided during muting is more than a few words (e.g., more than ten seconds), then user prompted to either initiate the playback the portion of speech provided while on mute or repeating speech again; which corresponds to notify the participant 102A that they are on mute and prompt the participant 102A to manually unmute the endpoint];
receiving the user response to each prompt [see para. 0088-0089; unmuting notification/action automatically unmute endpoint 104A to provide speech as a portion of the conference content, unmuting notification/action further include signaling endpoint wherein the signal causes a notification to be presented by only endpoint, that they are off mute (e.g., tone, message, pop-up message, etc.). As a further option, all endpoints notified of the on-mute/off-mute state of endpoints and, when changed, each endpoint is updated accordingly, such as with a message (e.g., “Alice is on mute” or “Alice is off mute.”) or graphical icon having a meaning associated with the muting state. Optionally, speech buffered and replayed as conference content, so that any speech provided before the unmuting notification/action results in the unmuting of endpoint is provided as uninterrupted speech but with a delay determined by the beginning of speech and the occurrence of the unmuting action]; however, Deole fails to explicitly teach in response to receiving user responses to unmute and play the recording, unmuting an audio feed from the user to the one or more other participants and playing the recording for the one or more other participants.
Finlow-Bates discloses in response to receiving user responses to unmute and play the recording, unmuting an audio feed from the user to the one or more other participants and playing the recording for the one or more other participants [see para. 0028-0031; the communication device processor configured with software to automatically detect when to turn off the mute function and initiate playback, such as by recognizing when the caller is speaking, communication device configured with a smart un-mute user interface input (e.g., a virtual button or icon) to enable the caller to select when to playback a portion of the recorded audio to the other participants. Such a smart un-mute user interface also enable the user to designate the duration or portion of the buffered audio to be played back; which corresponds to a smart mute for recovering speech spoken while a communication device is muted during a voice call].
It would have been obvious to one of an ordinary skill in the art, having the teachings of Deole and FInlow-Bates before the affective filing date of the claimed invention to modify, the notification and recovery process of Deole to include multiple user prompts requiring responses as taught by Finlow-Bates.
One would have been motivated to make this modification to provide the user with control over the recovered speech is injected into the conference, thereby improving user experience and reducing risk of unwanted playback of private or unintended speech.
Regarding claims 16-20, directly or indirectly dependent on claim 15, essentially correspond to those of claims 2-7 respectively. Accordingly, the same reasoning as in claims 2-7 applies to claims 16-20.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (See PTO-892).
Chavez (US 11,082,465) discloses systems, methods, and software to provide intelligent detection and automatic correction of erroneous audio settings in a video conference. Electronic conferences can often be the source of frustration and wasted resources as participants may be forced to contend with extraneous sounds, such as background/ambient noises, or conversations not intended for the conference, provided by an endpoint that should be muted. Similarly, participants may speak with the intention of providing their speech to the conference while their associated endpoint is muted. As a result, the conference may be awkward and lack a productive flow while endpoints are erroneously muted or non-muted. By intelligently processing at least the video portion of a video conference, endpoints/participants may be prompted to mute/unmute or automatically muted/unmuted].
A reference to specific paragraphs, columns, pages, or figures in a cited prior art reference is not limited to preferred embodiments or any specific examples. It is well settled that a prior art reference, in its entirety, must be considered for all that it expressly teaches and fairly suggests to one having ordinary skill in the art. Stated differently, a prior art disclosure reading on a limitation of Applicant's claim cannot be ignored on the ground that other embodiments disclosed were instead cited. Therefore, the Examiner's citation to a specific portion of a single prior art reference is not intended to exclusively dictate, but rather, to demonstrate an exemplary disclosure commensurate with the specific limitations being addressed. In re Heck, 699 F.2d 1331, 1332-33,216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006,1009, 158 USPQ 275, 277 (CCPA 1968)). In re: Upsher-Smith Labs. v. Pamlab, LLC, 412 F.3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir. 2005); In re Fritch, 972 F.2d 1260, 1264, 23 USPQ2d 1780, 1782 (Fed. Cir. 1992); Merck & Co. v. Biocraft Labs., Inc., 874 F.2d 804, 807, 10 USPQ2d 1843, 1846 (Fed. Cir. 1989); In re Fracalossi, 681 F.2d 792,794 n.1,215 USPQ 569, 570 n.1 (CCPA 1982); In re Lamberti, 545 F.2d 747, 750, 192 USPQ 278, 280 (CCPA 1976); In re Bozek, 416 F.2d 1385, 1390, 163 USPQ 545, 549 (CCPA 1969).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAO H NGUYEN whose telephone number is (571)272-4053. The examiner can normally be reached on Mon-Fri 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kieu Vu can be reached on 571-272-4057. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CAO H NGUYEN/Primary Examiner, Art Unit 2171