DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed on 6/30/26 have been fully considered but they are not persuasive. Applicant argues on pp11 of Applicant Remarks that Abate thus fails to teach or suggest "identifying a highest priority level among the determined priority level and each received priority level corresponding to the at least one audio input,".
In response, examiner disagrees. Abate teaches identifying a highest priority level among the determined priority level and each received priority level corresponding to the at least one audio input (e.g. e.g. identifying a sound level above the threshold value a highest priority level among the determined priority level as suggested in Fig 4a-b and each received priority level corresponding to the at least one audio input [0062]-[0066]); The user's sound level is defined as the volume level of the user's speech [0062], which can increase or decrease.
Furthermore, Applicant argues on pp12, with regard to claim 6 that Abate fails to teach "determining a confidence level whether the audio input detected at the microphone includes a voice signal".
In response, examiner disagrees. Abate teaches determining a volume level corresponding to the audio input detected at the microphone (e.g. The user's sound level is defined as the volume level of the user's speech [0062]); determining a confidence level whether the audio input detected at the microphone includes a voice signal (e.g. the audio input which is above the threshold value of Fig 4a-c is determining a confidence level whether the audio input detected at the microphone includes a voice signal); and determining, based on the volume level and the confidence level, the priority level of the audio input detected at the microphone (e.g. determining whether the user is active or inactive based on the volume level and the confidence level, as suggested in [0046], [0060]).
Examiner explained during the interview how the threshold value of Fig 4a-c was interpreted as the confidence level wherein the system has a high confidence level with a user’s speech above the threshold value and a low confidence level with the user’s speech below the threshold value as taught in Fig 4a-c.
Because the applied prior art, as a whole, still reads on the claim limitations as previously presented, examiner maintains his rejection.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-10, 12-20 are rejected under 35 U.S.C. 103 as being unpatentable over Abate (US 2012/0182381) in view of Chiang (US 2018/0184045).
With respect to claim 1 (similarly claims 19-20), Abate teaches an electronic device (e.g. a first user terminal 104 Fig 1-2 [0038] and [0044]), comprising:
a microphone (e.g. microphone [0039]);
one or more processors (e.g. a local processor in the user terminal 104 [0039]);
a memory (e.g. memory 208 Fig 2 [0047]); and
one or more programs (e.g. program 110 [0039], see also claim 35), wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors (e.g. the program is stored in the memory and executed by the local processor of [0039], claim 35), the one or more programs including instructions for:
receiving at least one audio input (e.g. receive at least one audio input as suggested in [0042], Fig 2 [0045]), wherein each audio input of the at least one audio input is associated with a respective priority level (e.g. each audio input is associated with a respective sound level [0046]);
determining a priority level of an audio input detected at the microphone (e.g. determining a priority level i.e. active or inactive [0046] of an audio input detected at the microphone of [0039]);
identifying a highest priority level among the determined priority level and each received priority level corresponding to the at least one audio input (e.g. e.g. identifying a sound level above the threshold value a highest priority level among the determined priority level as suggested in Fig 4a-b and each received priority level corresponding to the at least one audio input [0062]-[0066]);
However, Abate fails to teach obtaining a textual representation of a respective audio input corresponding to the identified highest priority level; and
displaying, on the display, the obtained textual representation.
Chiang teaches obtaining a textual representation of a respective audio input corresponding to an identified highest priority level (e.g. obtaining a textual representation of a respective audio input, see Figs 1-3, Fig 5 S506 [0060] corresponding to an identified highest level i.e. only the speaker who is currently speaking as suggested in [0020]; and
displaying, on the display, the obtained textual representation (e.g. displaying, on text interface, the obtained textual representation, Fig 5 S510 [0061]).
Abate and Chiang are analogous art because they all pertain to video calling and improvement thereof. Therefore, it would have been obvious to people having ordinary skill in the art before the effective filing date of the claimed invention to modify Abate with the teachings of Chiang to include: obtaining a textual representation of a respective audio input corresponding to the identified highest priority level; and
displaying, on the display, the obtained textual representation, as disclosed in Fig 5 [0060]-[0061]. The benefit of the modification would be to not only display the speaker in the video call, but to convert his/her conversation into text and display the textual representation of the call, thus satisfying the user.
With respect to claim 2, Abate teaches the device of claim 1, comprising: while the electronic device is in a two-way communication session with a second electronic device (e.g. while terminal 104 is in a two-way communication session with a second terminal, as suggested in Fig 1 [0042], [0045] where users A-D are in a conference call):
detecting a first power level of the audio input (e.g. detecting a first power level/volume of the audio input, as suggested in Fig 4a-c [0062]) detected at the microphone (e.g. detected at the microphone of [0039]);
detecting a second power level of an audio input (e.g. detecting a second power level/volume of an audio input, as suggested in Fig 4a-c [0062]) from the second electronic device (e.g. from the second device i.e. any terminal device 116,118,120 Fig 1);
in accordance with a determination that both the first power level and the second power level are below a threshold power level (e.g. in accordance with a determination that both the first power level and the second power level are below a threshold power level i.e. the threshold value of Fig 4a-c), disabling a voice activity detector (e.g. demote the user to the inactive display (the roster) 304, Fig 3 [0052], thus disabling a voice detector 204 Fig 2 when the user is inactive as suggested in Fig 3); and
in accordance with a determination that both the first power level and the second power level are above the threshold power level (e.g. in accordance with a determination that both the first power level and the second power level are above the threshold power level i.e. the threshold value of Fig 4a-c), enabling a voice activity detector (e.g. keeping the user in display 302 of Fig 3 [0050]-[0051], thus enabling a voice detector 204 Fig 2 when the user is active to continue to monitor his speech activity).
With respect to claim 3, Abate in view of Chiang teaches the device of claim 1, comprising: while the electronic device is in a two-way communication session with a second electronic device (e.g. while terminal 104 is in a two-way communication session with a second terminal, as suggested in Fig 1 [0042], [0045] where users A-D are in a conference call), and while a voice activity detector is enabled (e.g. while voice detector 204 Fig 2 is enabled, as suggested in [0046]), in accordance with a determination that a voice signal is detected in either the audio input detected at the microphone or an audio input from the second electronic device (e.g. in accordance with a determination that a voice signal is detected in either the audio input detected at headset 112/microphone or an audio input from the second terminal device/ any terminal device 116,118,120 Fig 1 as suggested in [0045]-[0046]): enabling a speech recognition processor at the electronic device (e.g. enabling a comparator 206 which determines whether the users in the video call are active or inactive participants in the call using a sound level threshold value input to comparator 206 [0046], see also [0056] of Chiang where a voice recognition software is enabled); and
obtaining the textual representation using the speech recognition processor (Chiang e.g. obtaining the textual representation using the voice recognition software of [0056], as disclosed in [0060]).
With respect to claim 4, Abate teaches the device of claim 1, comprising: while the electronic device is in a communication session with a plurality of additional electronic devices (e.g. while terminal device 104 Fig 1 is in communication session with terminal devices 116,118,120 Fig 1 as suggested in [0045]-[0046]), enabling a voice activity detector (e.g. enabling voice detector 204 Fig 2, as suggested in [0045]-[0046]).
With respect to claim 5, Abate teaches the device of claim 4, comprising: while the electronic device is in a communication session with a plurality of additional electronic devices (e.g. while terminal device 104 Fig 1 is in communication session with terminal devices 116,118,120 Fig 1 as suggested in [0045]-[0046]): detecting a number of devices participating in the communication session (e.g. detecting a number of devices participating in the conference call, as suggested in Fig 3, 6-7); and
in accordance with a determination that the number of devices participating in the communication session is less than three devices (e.g. when one of the users in Fig 3 has been quiet for more than 10 seconds, see Fig 4c [0072], i.e. a determination that the number of devices participating in the communication session is less than three devices), selectively enabling the voice activity detector based on audio power level (e.g. selectively enabling the voice detector based on the sound level, as suggested in [0072]).
With respect to claim 6, Abate teaches the device of claim 1, comprising: determining a volume level corresponding to the audio input detected at the microphone (e.g. The user's sound level is defined as the volume level of the user's speech [0062]); determining a confidence level whether the audio input detected at the microphone includes a voice signal (e.g. the audio input which is above the threshold value of Fig 4a-c is determining a confidence level whether the audio input detected at the microphone includes a voice signal); and determining, based on the volume level and the confidence level, the priority level of the audio input detected at the microphone (e.g. determining whether the user is active or inactive based on the volume level and the confidence level, as suggested in [0046], [0060]).
With respect to claim 7, Abate teaches the device of claim 1, wherein each audio input of the at least one audio input is associated with a respective additional electronic device (e.g. the speech data is associated with additional electronic device, see Fig 1 [0045]), comprising: after determining the priority level of the audio input detected at the microphone, transmitting, to each respective additional electronic device, the determined priority level (e.g. the speech data received and transmitted between terminal devices as suggested in [0045] include the sound level which is transmitted along with the speech data to each respective additional electronic device 16,118,120).
With respect to claim 8, Abate teaches the device of claim 1, wherein the respective priority level associated with each audio input of the at least one audio input is received with each audio input (e.g. receiving the speech data, including the sound level [0046], from terminals 16,118,120 [0045] suggest the respective priority level associated with each audio input of the at least one audio input is received with each audio input).
With respect to claim 9, Abate in view of Chiang teaches the device of claim 1, wherein the textual representation is obtained subject to predetermined criteria being satisfied (Chiang e.g. the textual representation is obtained in S506 of Fig 5 subject to predetermined criteria being satisfied), wherein the predetermined criteria at least one of:
a criterion that a supported transcription language matches a keyboard settings language, a criterion that a supported transcription language matches a user-to-digital assistant interaction language, and a criterion that a supported transcription language matches a default language setting of the electronic device (Chiang e.g. The voice recognition engine 562 can convert the audio file 560 and convert it to a text file 564 suitable for display on the UEs 108, 552, 554. [0060] wherein a text file 564 suitable for display on the UEs 108, 552, 554 suggest a criterion that a supported transcription language matches a user-to-digital assistant interaction language, and a criterion that a supported transcription language matches a default language setting of the electronic device).
With respect to claim 10, Abate in view of Chiang teaches the device of claim 1, wherein displaying, on the display, the obtained textual representation (Chiang e.g. displaying the textual representation as disclosed in Fig 1-3, Fig 5 S510) comprises: displaying, on the display, a video representation of a communication between the electronic device and at least one additional electronic device (e.g. displaying in the video window 102 Fig 1 a video representation of a communication between the electronic device and at least one additional electronic device as suggested in [0018]-[0019]); identifying a location of an image sensor of the electronic device (e.g. identifying a location of an image sensor of the electronic device, see Figs 1-3 the oval-shaped dashed lines on top of device 100/200/300); and
displaying, adjacent to the image sensor, the obtained textual representation (Chiang e.g. and displaying, adjacent to the image sensor, the obtained textual representation, see Fig 1-3).
With respect to claim 12, Abate in view of Chiang teaches the device of claim 1, comprising: receiving a first touch input on the displayed textual representation (Chiang e.g. highlighting a portion on the screen of the UE 108) [0028]-[0029] i.e. receiving a first touch input on the displayed textual representation), wherein the displayed textual representation includes a first number of lines of text (Chiang e.g. the displayed textual representation includes a first number of lines of text, see Figs 1-3); and
in response to receiving the touch input, displaying an expanded textual representation by increasing a displayed size of the displayed textual representation, wherein the expanded displayed textual representation includes a second number of lines of text greater than the first number of lines of text (Chiang e.g. highlighting a portion on the screen of the UE 108) [0028]-[0029] is displaying an expanded textual representation by increasing a displayed size of the displayed textual representation, wherein the expanded displayed textual representation includes a second number of lines of text greater than the first number of lines of text because the screen of UE 108 is a touch-sensitive display 422 Fig 4 [0055] which is equipped to handle these changes/editions/manipulations, as suggested in [0063]).
With respect to claim 13, Abate in view of Chiang teaches the device of claim 12, comprising: receiving a second touch input on the expanded displayed textual representation, wherein the second touch input includes a predetermined motion pattern in a first direction (Abate, as modified by Chiang, receives a second input i.e. the scroll of Fig 6c [0059]); and in response to receiving the second touch input, displaying additional text of the textual representation by scrolling the expanded displayed textual representation in a vertical direction (Abate, as modified by the highlighted, expanded, displayed textual representation of Chiang, uses the scrollable area 307 Fig 6c [0059] to scroll the expanded displayed textual representation in a vertical direction or any other direction).
With respect to claim 14, Abate in view of Chiang teaches the device of claim 1, comprising: storing, at the electronic device, a predetermined duration of each received audio input and the audio input detected at the microphone (Chiang e.g. storing, in the call log 106 of Fig 1-3 a predetermined duration of each received audio input and the audio input detected at the microphone); and in accordance with a determination that at least two voice signals detected among the received audio input and the audio input detected at the microphone (Chiang e.g. in accordance with a determination that at least two voice signals detected among the received audio input and the audio input detected at the microphone as suggested in Fig 1-3), performing speech recognition using a stored predetermined duration of a received audio input (Chiang e.g. performing speech recognition [0037] using a stored predetermined duration of a received audio input, see call log 106 Fig 1-3).
With respect to claim 15, Abate in view of Chiang teaches the device of claim 1, comprising: while obtaining the textual representation of the respective audio input (Chiang e.g. while obtaining the textual representation of the respective audio input, as suggested in Figs 1-3, Fig 5 S506), detecting a respective priority level of a second audio input (Chiang e.g. detecting a respective priority level of a second audio input for caller 2, Figs 1-3, as each caller has a turn to speak and the caller who is currently speaking has the priority as suggested in [0020]), different from the respective audio input (Chiang e.g. different from the respective audio input from caller 1 Fig 1-3), as having a highest priority level among the determined priority level of the audio input detected at the microphone and each received priority level corresponding to the at least one audio input (Chiang e.g. caller 2 has a highest priority level as being the one who is currently speaking [0020] among caller 1 and 2); and in response to detecting the respective priority level of the second audio input having the highest priority level (Chiang e.g. in response to detecting that the caller 2 is currently speaking i.e. detecting the respective priority level of the second audio input having the highest priority level), continuing to obtain the textual representation of the respective audio input (Chiang e.g. continue to obtain the textual representation of the respective audio input i.e. caller 2 speech converted to text and displayed on display 102 of Fig 1-3).
With respect to claim 16, Abate in view of Chiang teaches the device of claim 15, comprising: in accordance with a determination that the textual representation of the respective audio input has been obtained (Chiang e.g. in accordance with a determination that the textual representation of the respective audio input has been obtained as suggested in Fig 1-3), retrieving, from a buffer component (Chiang e.g. retrieving from call log 106 Fig 1-3), a second textual representation of a respective stored duration of audio corresponding to the second audio input (Chiang e.g. a second textual representation of a respective stored duration of audio corresponding to the second audio input, see Fig 1-3).
With respect to claim 17, Abate in view of Chiang teaches the device of claim 1, wherein the electronic device is one of a mobile phone, a tablet device, a laptop computer, a desktop computer, or a wearable device (Abate e.g. The user terminal 104 may be, for example, a personal computer ("PC"), personal digital assistant ("PDA"), a mobile phone, a gaming device or other embedded device able to connect to the network 106 [0038]).
With respect to claim 18, Abate in view of Chiang teaches the device of claim 1, comprising: prior to displaying the obtained textual representation (Chiang e.g. prior to displaying the obtained textual representation at S510 of Fig 5), displaying, on the display, a video representation corresponding to a communication session between the electronic device and at least one additional electronic device (Chiang e.g. displaying, on display 102 Fig 1-3, a video representation corresponding to a communication session between the electronic device and at least one additional electronic device, see S502 of Fig 5 [0057]).
Claim(s) 11 is rejected under 35 U.S.C. 103 as being unpatentable over Abate (US 2012/0182381) in view of Chiang (US 2018/0184045) and further in view of Deluca (US 2012/0077480).
With respect to claim 11, Abate in view of Chiang teaches the device of claim 10.
However, Abate fails to teach comprising: detecting a change in orientation of the electronic device; in response to the detected change in orientation of the electronic device: rotating the display of the video representation of the communication between the electronic device and at the least one additional electronic device, wherein the rotation maintains the orientation of the displayed video representation; and maintaining the displayed position of the obtained textual representation adjacent to the image sensor.
Deluca teaches detecting a change in orientation of the electronic device (e.g. orientation component 366 Fig 1 detects a change in orientation of the electronic device [0027]); in response to the detected change in orientation of the electronic device (e.g. in response to the detected change in orientation of the electronic device as suggested in [0027]):
rotating the display of the information representation of the communication between the electronic device and at the least one additional electronic device (e.g. rotate the display of information representation of the communication between the electronic device and at the least one additional electronic device, see Fig 5 [0029]), wherein the rotation maintains the orientation of the displayed information representation (e.g. the rotation maintains the orientation of the displayed information representation, see Fig 5); and maintaining the displayed position of the obtained textual representation adjacent to the image sensor (e.g. maintaining the displayed position of the obtained textual representation adjacent to the image sensor as suggested in Fig 5 [0029]).
Abate and Deluca are analogous art because they all pertain to displaying video/information on the display interface of an electronic device. Therefore, it would have been obvious to people having ordinary skill in the art before the effective filing date of the claimed invention to modify Abate with the orientation component 366 of Fig 1 to include: comprising: detecting a change in orientation of the electronic device; in response to the detected change in orientation of the electronic device: rotating the display of the video representation of the communication between the electronic device and at the least one additional electronic device, wherein the rotation maintains the orientation of the displayed video representation; and maintaining the displayed position of the obtained textual representation adjacent to the image sensor, as taught by Deluca in Fig 1 [0027]-[0029]. The benefit of the modification would be to display the video/information to the user even the electronic device is rotated to 90 degrees.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IBRAHIM SIDDO whose telephone number is (571)272-4508. The examiner can normally be reached 9:00-5:30PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Akwasi Sarpong can be reached at 5712703438. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IBRAHIM SIDDO/Primary Examiner, Art Unit 2681