DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Allowable Subject Matter
Claims 3-12, 16-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 13-15 and 19-20 are rejected under 35 U.S.C. 102(a)(1) as anticipated by or, in the alternative, under 35 U.S.C. 103 as obvious over Yamashita (and if necessary) in view of Yu OR Zhang and further in view of Katae et al (US 2011/0060590), Ramachandra et al (US 2021/0326372), Cooper (US 2008/0260350) OR Costanzo (US 2016/0330408).
Claims 1, 13, 14 and 19, Yamashita teaches an apparatus, an electronic device, a storage media and a method for audio processing in a multi-view mode (see Fig. 4 for multi-view mode), the method comprising:
based on an occurrence of a preset audio output adjusting event, triggering corresponding views to perform audio output by using matching audio output devices based on a policy (an audio adjustment unit configured to adjust focus values each indicating a degree of highlighting of audio data of each content displayed in a plurality of display areas, [0005]), that a first application uses a real audio output device and a non-first application uses a virtual audio output device, (Application example 4, [0085-0087]: degrees of importance may be given to the keywords. It is thus possible for the focus ratio setting unit 122 to set the focus value of the audio of the secondary program in response to the degree of importance of the keyword. For example, when a program matching the keyword having a high degree of importance is detected, the focus ratio setting unit 122 may set the focus value of the program to be greater (e.g., equal to or greater than the focus value of the audio of the primary program), thereby strongly notifying the viewer of the broadcasting of the corresponding program. On the other hand, when a program matching the keyword having a low degree of importance is detected, the focus ratio setting unit 122 may set the focus value of the program to be smaller (e.g., smaller than the focus value of the audio of the primary program));
Yamashita does not teach “wherein the virtual audio output device is a pseudo-rendering instance having an audio and video synchronization function for realizing audio synchronization through analog audio output”.
Based on the current Specs, [0051], “The virtual sound output device is a pseudo-rendering instance having an audio and video synchronization function. That is to say, the virtual sound output device is not the real sound output device, but only the pseudo-rendering instance for realizing audio synchronization through analog audio output”. A well-known technique in the art as examiner wishes to provide several references addressing this claimed feature. Please note examiner maps the claimed pseudo-rendering to the creating of the synthetic audio/video feed, material, signal and/or media file as examiner presents:
Katae: [0047] The synthetic speech text-input device 1 can be used, for example, as a device with which a user enters a text that is converted into a synthetic speech and is added (inserted) in synchronization with video data in a video editing system. The present embodiment is explained with reference to a case, as an example, where the synthetic speech text-input device 1 is used for inputting a text for a synthetic speech to be added to a designated section in video data.
Ramachandra: [0063] The synthetic video files 108 and the synthetic audio files 112 may be merged, for example, using audio video synchronization, to generate final synthesized videos, where the final synthesized videos may be designated as the synthetic media files 116. A plurality (e.g., 60) of such synthetic media files 116 may be generated, and stored in a database.
[0091] After creation of the synthetic audio files 112, the synthetic audio files 112 may be merged with the synthetic video files 108 as disclosed herein to generate the synthetic media files 116. During creation of the synthetic video files 108, the actor may be directed to speak as per a relevant conversation topic for the digital persona 104. Thus, during merging of the synthetic video files 108 and the synthetic audio files 112, lip synchronization may be automatically accounted for.
Cooper, via FIG. 2 and [0053], shows an example of an "unobtrusive synchronizer" device configured according to one embodiment of the invention. Essentially, this embodiment functions by providing frequent synthetic but non-obtrusive audio video synchronization signals, typically every few seconds. As previously discussed, these non-obtrusive signals are designed to be intense enough to be reliably detected by automated equipment designed for this purpose, but unobtrusive enough as to not detract from the viewer's enjoyment of the program. According to the invention, these events may be unobtrusive enough to be either dismissed by the viewer as background audio and visual noise; or may be completely undetectable by human viewers; or, alternatively, may be unobtrusive enough so as to be capable of being effectively subtracted from the final signal by automated audio and visual signal processing equipment. And via par. [0066] and It is also possible for a user to adjust the rate or timing of generation of events (13) and (12) via automated or manual user adjustment (9). For example, in programs, like sports programs, where the potential for large or sudden changes in audio or video signal processing is high (due for example to the difficulty of compressing scenes with a lot of detail and motion), the speed (rate of generation of synthetic unobtrusive audio and video synchronization events) may be manually or automatically increased to facilitate quick downstream analysis of audio to video timing, OR
Costanzo teaches, “generating said synthetic view transitions containing novel audio-video at the determined time intervals and for the determined durations in synchronization with time alignment of the audio-video capture feeds or streams, wherein the synthetic view transitions represent at least one of a plurality of possible trajectories in said tri-dimensional space of the venue, [0068]”.
wherein the audio output adjusting event is an event causing the audio output devices used by the views to not match the policy, wherein when a mixing mode is off, the first application is an application currently in an audio focus, and wherein, when the mixing mode is on, the first application is an application currently participating in mixing ( Yamashita teaches, per [0098], “Traffic information and music or television programs may also be presented to the car navigation system at the same time. Here, while listening to the music from the car navigation system, an interruption such as traffic jam information or road construction information may be carried out using the audio output system of the present embodiment, and an audio mixed with the interruption information may be output without stopping the music reproduction. In this case, since the interruption information is considered as information having a high degree of importance, the focus ratio setting unit 122 may decrease the focus value of the music and increase the focus value of the composed audio of the interruption information, thereby allowing the interruption information to be readily listened to.
Further clarification “policy” which Yamashita does not use the term “policy”. Yet he teaches, [0086] Here, degrees of importance may be given to the keywords. It is thus possible for the focus ratio setting unit 122 to set the focus value of the audio of the secondary program in response to the degree of importance of the keyword. For example, when a program matching the keyword having a high degree of importance is detected, the focus ratio setting unit 122 may set the focus value of the program to be greater (e.g., equal to or greater than the focus value of the audio of the primary program), thereby strongly notifying the viewer of the broadcasting of the corresponding program. On the other hand, when a program matching the keyword having a low degree of importance is detected, the focus ratio setting unit 122 may set the focus value of the program to be smaller (e.g., smaller than the focus value of the audio of the primary program). Here examiner maps “policy” to low degree of importance, [0086] or high degree of importance, [0098] or highest degree of importance, [0087] of keyword, [0085]). Yamashita suggests or at least by obviousness presents “policy>
To support this obviousness, examiner provides Yu who teaches, “ the mixing policy may specifically include a type of audio data that needs to be mixed and an application from which the audio data that needs to be mixed comes, and/or a type of audio data that does not need to be mixed and an application from which the audio data that does not need to be mixed comes, [0107]. Therefore, in this embodiment of this application, a mixing policy may be preset in the master device (for example, the mobile phone), and the mixing policy may be used to indicate whether to allow to mix audio data of different types, [0716]”.
Zhang teaches, per [0352] Figs. 13B is a schematic diagram of a data flow direction in an Android operating system according to an embodiment of this application. For example, in this embodiment of this application, an application running in the foreground/background (which is referred to as a foreground/background application for short below) may invoke the media player (MediaPlayer) at the framework layer. Then, an audio stream played by the foreground/background application is output, and the output audio stream is sent to AudioFlinger. AudioFlinger sends the audio stream to the audio mixing module. The audio mixing module processes the audio stream and provides a corresponding audio mixing policy. Finally, the audio mixing module at the framework layer invokes AudioHAL at the hardware abstraction layer HAL, and AudioHAL sends the output audio stream to a device at the hardware layer for playing. Where FIG. 13B is specifically subdivided into the audio source classifier and the policy library.
Therefore it would have been obvious to the ordinary artisan before the effective filing date to incorporate the teaching of Yu or Zhang into the teaching of Yamashita for the purpose of explicitly defining the policy whether it is an audio playing policy, an audio mixing policy to carry out the method of the invention without confusion to improve an audio control, a greater flexibility switch control and thus to improve the user experience and also to incorporate the teaching of Katae, Ramachandra, Cooper and/or Costanzo into the teaching of Yamashita for the purpose of providing a greater improvement of synchronization technology, i.e., the accuracy in the prediction of the acceptable text amount; providing unobstructive event in the audio/video signal; better optimization of the bandwidth usage and of the required processing resources on both the server and the client side and is scalable to any number of interactive users; and/or better implementation of the conversational digital persona with ethical human centered computing and multimodal analysis of data may be performed by sensing users and environment, and creating the digital persona.
Claims 2 and 20. The method of claim 1, wherein the audio output adjusting event comprises at least one of view modes being switched, the audio focus being changed and the mixing mode being off in the multi-view mode, the mixing mode being turned on or off in the multi-view mode, or a set of applications participating in mixing being changed. (Yamashita: In the remote 200 of the present embodiment, for example, a display mode switching button for switching the display area of the display unit 110 between a one-screen mode and a two-screen mode, a focus value ratio change slider for changing the ratio of the focus value with respect to the audio of each program in the two screen modes, and so forth are disposed. The user may change the ratio of the focus value of each program using the focus value ratio change slider, so that the audio of the interesting program may be more clearly output and the audio of the other program with a recognizable sound quality may also be output naturally without disruption. That is, it is possible to change the audio to the sound quality according to degrees of interest of the viewers. It is thus possible to satisfy the interest of each viewer and for the viewers to share the same time and place, [0026]).
Claim 15. The electronic device of claim 14, wherein the audio output adjusting event comprises at least one of view modes being switched, the audio focus being changed and the mixing mode being off in the multi-view mode, the mixing mode being turned on or off in the multi-view mode, or a set of applications participating in mixing being changed. (See claims 2 and 20 above and in addition see Fig. 4 for sliding scale from 0 to 100 where examiner maps 0 to OFF).
Response to Arguments
Applicant’s arguments with respect to claim(s) 7/10/26 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Applicant argues and submits that Yamashita, alone or in combination with Yu and/or Zhang, fails to disclose "based on an occurrence of a preset audio output adjusting event, triggering corresponding views to perform audio output by using matching audio output devices based on a policy that a first application uses a real audio output device and a non-first application uses a virtual audio output device, wherein the virtual audio output device is a pseudo-rendering instance having an audio and video synchronization function for realizing audio synchronization through analog audio output, wherein the audio output adjusting event is an event causing the audio output devices used by the views to not match the policy, wherein, when a mixing mode is off, the first application is in an audio focus, and wherein, when the mixing mode is on, the first application is currently participating in mixing," as presently recited.
Examiner respectfully disagrees as examiner has produced additional reference to address the applicant’s argument.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHUNG-HOANG J. NGUYEN whose telephone number is (571)270-1949. The examiner can normally be reached Reg. Sched. 6:00-3:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 571-272-7503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHUNG-HOANG J NGUYEN/Primary Examiner, Art Unit 2691