Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 6, 7,9, 13, 15, 12, 18 are rejected under 35 U.S.C. 103 as being unpatentable over Maxwell (US 20240233745) in view of Price (US 20250239025).
Regarding claim 1, Maxwell teaches, a method for visual cueing of audio context in an assistive audio call for a deaf or a hard of hearing (HOH) participant (abstract: Video relay services, Paragraph 24: The artificial intelligence engine is also configured to receive the audio stream including voice data from the hearing-capable user, analyze the voice data (e.g., using voice recognition software) to translate the voice data into a language supported by the system that is understood by the hearing-impaired user. In some embodiments, the artificial intelligence engine then communications the translated data (e.g., text and/or gestures) to the hearing-impaired user during the call) comprising:
establishing an assistive call with an HOH participant and a counterpart participant (Fig. 2, el. 102, 104 hearing impaired and hearing capable);
receiving an audio stream from the counterpart participant (Fig. el. 216 and Paragraph 35: the hearing-capable user speaks into the microphone of the far-end communication device 104, which transmits the voice data with an audio stream to the video relay service 106) and identifying speech audio within the audio stream (Paragraph 35);
submitting the speech audio to a speech transformation engine in order to transform the speech audio to a visual form of the speech for consumption by the HOH participant (Paragraph 35: At operation 218, the AI translation engine 110 of the video relay service 106 receives and analyzes the audio stream, and recognizes the spoken language (e.g., English, Spanish, etc.) according to various voice recognition systems. This translation may occur using various voice recognition services that translate voice data into text data as known in the art or other speech-to-text systems that use phonetic sound libraries 252 and grammar rules to recognize words and phrases using contextual information or that are configured to read text outload. As a result, the spoken language is translated into a textual based language understood by the hearing-capable user (e.g., text of English, Spanish, etc.). The AI translation engine 110 may transmit the translation as text data to the far-end communication device 104 at operation 220. The translated text is then displayed on the electronic display of the video communication device 102 “text: is visual form”);
displaying the visual form in a user interface to the assistive audio call (Paragraph 68: the hearing-impaired user's own video stream may be displayed on the video communication device during the call. The user interface 1300 may also include a second video area 1320 for displaying the avatar received from the video relay service corresponding to the translation of the far-end user's audio into an avatar performing sign language. In some embodiments, the user interface 1300 may include a text area 1330 for displaying the translated text received from the video relay service corresponding to the translation of the hearing-impaired user's near-end video).
Maxwell teaches displaying multiple information (Fig. 13, el. 1320, 1340, 1330, 1310).
Maxwell does not explicitly teach processing a portion of the
audio stream separate from the transformation of the speech audio to the visual form in order to identify an audio context of the audio stream; matching the audio context to a visual cue; and, supplementing the visual form in the user interface with the visual cue of the audio context.
Price teaches (abstract: detect an ambient sound from audio data captured by a plurality of microphones on a display device. A display device may determine, based on the audio data, a location of a sound source of the ambient sound. A display device may generate contextual information about the ambient sound based on an audio segment of the audio data that includes the ambient sound. A display device may display the contextual information based on the location of the sound source) Price also teaches processing a portion of the audio stream separate from the transformation of the speech audio to the visual form in order to identify an audio context of the audio stream; matching the audio context to a visual cue; and, supplementing the visual form in the user interface with the visual cue of the audio context (Fig. 3, el. 330 and Paragraph 28: The display device provides a technical solution for enabling a display device with perception-enhanced directional capabilities for ambient sounds, including determining a directionality of the ambient sound (e.g., there is a honking car to your right), labeling the source of information (e.g., Matt is speaking), overlaying background audio sources into the user's experience (e.g., there is a child crying upstairs), and/or labeling attribute data (e.g., a 22-year woman is speaking on your right, saying [insert translation]). In some examples, the display device provides a technical effect for hearing-impaired users such as converting auditory sounds into visual cues for the user).
Therefore, it would have been obvious to one with ordinary skill in the art before the filing date of the claimed invention to modify Maxwell with Price in order to improve the system and enhance user’s experience.
Regarding claim 3, Maxwell in view of Price teaches, wherein the audio context is a determination of a type of background noise (Paragraph 13 and Fig. 3, el. 330-1, 330-2).
Regarding claim 6, Maxwell in view of Price teaches, wherein the visual form is captioned text speech recognized from the speech audio in the audio stream (Maxwell Paragraph 28-29).
Regarding claim 7, see claim 1 rejection.
Regarding claim 9, see claim 3 rejections.
Regarding claim 12, see claim 6 rejections.
Regarding claim 13, see claim 1 rejection.
Regarding claim 15, see claim 3 rejections.
Regarding claim 18, see claim 6 rejections.
Claims 2, 8, 14 are rejected under 35 U.S.C. 103 as being unpatentable over Maxwell (US 20240233745) in view of Price (US 20250239025) in view of Eronen (US 20110190008).
Regarding claim 2, Maxwell in view of Price teaches, the audio context.
Maxwell in view of Price does not teach wherein the audio context is a determination of one of a masculine voice and a feminine voice.
Eronen teaches the audio context is a determination of one of a masculine voice and a feminine voice (Paragraph 78).
Therefore, it would have been obvious to one with ordinary skill in the art before the filing date of the claimed invention to modify Maxwell in view of Price with Eronen in order to improve the system and enhance user’s experience.
Regarding claim 8, see claim 2 rejections.
Regarding claim 14, see claim 2 rejections.
Claims 4, 10, 16 are rejected under 35 U.S.C. 103 as being unpatentable over Maxwell (US 20240233745) in view of Price (US 20250239025) in view of Karunamuni (US 20180335939),
Regarding claim 4, Maxwell in view of Price teaches,
Maxwell in view of Price does not teach wherein the audio context is a volume level of the speech indicative of tone.
Karunamuni teaches wherein the audio context is a volume level of the speech indicative of tone. (Paragraph 309 and FIGS. 5E24-5E27)
Therefore, it would have been obvious to one with ordinary skill in the art before the filing date of the claimed invention to modify Maxwell in view of Price with Karunamuni in order to improve the system and enhance user’s experience.
Regarding claim 10, see claim 4 rejections.
Regarding claim 16, see claim 4 rejections.
Claims 5, 11, 17 are rejected under 35 U.S.C. 103 as being unpatentable over Maxwell (US 20240233745) in view of Price (US 20250239025) in view of BREEDVELT-SCHOUTEN (US 20210097142).
Regarding claim 5, Maxwell in view of Price teaches, the audio context indicating sentiment (Price: Fig. 2, el. 230: happy people sound)
Maxwell in view of Price does not teach, wherein the audio context is a produced by a sentiment analysis engine.
BREEDVELT-SCHOUTEN teaches wherein the audio context is a produced by a sentiment analysis engine (Paragraph 90: The display 600 shows not only the fact that the conversation velocity of posts 610 exceeds the conversation velocity threshold 640 for the second period of time 650, but also the change in emojis 630 from happy (as indicated by smiley faces) to angry (as indicated by thunderbolts), indicating that the participants are angry).
Therefore, it would have been obvious to one with ordinary skill in the art before the filing date of the claimed invention to modify Maxwell in view of Price with BREEDVELT-SCHOUTEN in order to improve the system and enhance user’s experience.
Regarding claim 11, see claim 5 rejections.
Regarding claim 17, see claim 5 rejections.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARIA EL-ZOOBI whose telephone number is (571)270-3434. The examiner can normally be reached Monday-Friday 7-4.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn Edward can be reached at (571)270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARIA EL-ZOOBI/Primary Examiner, Art Unit 2692