DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 101
3. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
4. Claims 8-14 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Claims 8-14 each is directed to a computer program comprising a computer readable storage medium. Computer program is software per se which is subject-matter ineligible, and the medium as disclosed in Applicant’s specification is not limited to non-transitory embodiments, the claims as recited could possibly be directed to a transitory medium; therefore claims 8-14 are subject-matter ineligible.
Claim Rejections - 35 USC § 103
5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1,148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating
obviousness or nonobviousness.
6. Claims 1, 3, 5-6, 8, 10, 12-13, 15-16, and 18-19 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Zavesky et al. (US Publication 2023/0057722) in view of Spreitzer et al. (US Publication 2022/0293133), and further in view of Xu et al. (WIPO Publication WO 2023116208 06-2023).
Regarding claim 1, Zavesky discloses a method comprising:
capturing sounds from a plurality of objects within a volumetric video (Zavesky, para. 0062, fig. 2, in operations 202-204, the volumetric content capture system 102 captured and stores, in the storage 124, the volumetric video 110 in association with the audio 122. In operations 206-208, volumetric content capture system provides the volumetric video 110 and the audio 122 to the volumetric content enhancement system 126; para. 0066, the volumetric content analysis module 132 can utilize semantic classification concepts and technologies to classify scenes as dialog, suspense, action, or some other classification. The annotations 142 can identify the scenes with tags/labels according to the classifications. Analysis of a main view (e.g., default or original view) of the volumetric video 110 can be used to determine actors, objects, and various other contexts within and/or between scenes. Static inferences from the environment, time, landmarks, and the like depicted in the volumetric video 110 can also be used for additional annotations 142 and/or to refine existing annotations 142. The annotations 142 can provide a general sentiment depicted in the volumetric video 110, the audio 122, or both. The volumetric content analysis module 132 can associate the annotations 142 hierarchically with an entire scene, individual actors within a scene, individual objects within a scene, or the respective regions(s) 138 or voxels 136);
analyzing the captured sounds to generate context of the captured sounds; (Zavesky, para. 0066, the volumetric content analysis module 132 can utilize semantic classification concepts and technologies to classify scenes as dialog, suspense, action, or some other classification. The annotations 142 can identify the scenes with tags/labels according to the classifications. Analysis of a main view (e.g., default or original view) of the volumetric video 110 can be used to determine actors, objects, and various other contexts within and/or between scenes. Static inferences from the environment, time, landmarks, and the like depicted in the volumetric video 110 can also be used for additional annotations 142 and/or to refine existing annotations 142. The annotations 142 can provide a general sentiment depicted in the volumetric video 110, the audio 122, or both. The volumetric content analysis module 132 can associate the annotations 142 hierarchically with an entire scene, individual actors within a scene, individual objects within a scene, or the respective regions(s) 138 or voxels 136).
Zavesky does not explicitly disclose:
determining whether to remove any existing sounds or to add any new sounds resulting in sound variations;
dynamically modifying the volumetric video to synchronize the sound variations with the plurality of objects by employing a generative adversarial network (GAN) model; and
generating a new volumetric video exhibiting synchronization between the plurality of objects and the sound variations.
Spreitzer discloses:
determining whether to remove any existing sounds or to add any new sounds resulting in sound variations; dynamically modifying the volumetric video to synchronize the sound variations with the plurality of objects; and generating a new volumetric video exhibiting synchronization between the plurality of objects and the sound variations (Spreitzer, para’s 0016-0021, 0043-0050, and 0067-0073, scanning a video to find objects in the scene, such as a ball or a person’s head; tracking how those objects move over time. If the motion matches certain patterns, such as a collision, a sudden change in direction, or repeated bouncing, the system selects an audio or visual effect to match that motion. For example, it may add a “boom” sound at the moment of impact or slow down the video around a dramatic collision. It can also add music whose beat matches the rhythm of repeated motion, such as bouncing or juggling. The playback speed of the video or audio may be adjusted so the beat and motion line up better. Generating a new version of the video based on the edited result).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Spreitzer’s features into Zavesky’s invention for enhancing user’s playback experience by enriching sound effect that provides more harmony with objects depicted in the video.
Zavesky-Spreitzer does not explicitly disclose but Xu discloses does not explicitly disclose dynamically modifying the volumetric video to synchronize the sound variations with the plurality of objects by employing a generative adversarial network (GAN) model (Xu, para. 0007, generating a digital object is provided, the method comprising: acquiring target audio and a target digital object; obtaining a first driving video of the reference digital object driven by the target audio based on the target audio and a pre-trained speech-driven model of the reference digital object; wherein the speech-driven model is trained based on the audio and video of the reference digital object; and transferring the pose of the reference digital object in each video frame of the first driving video to the target digital object to obtain a second driving video of the target digital object driven by the target audio; para. 0046, a large amount of audio of the digital object A and video of the digital object A that matches the audio can be obtained (it can also be said that the actions and expressions in the video are synchronized with the audio, and in this disclosure it can also be called video synchronized with the audio). Then, the audio is used as a training sample and the video is used as a sample label to train the neural network model and obtain the speech-driven model of the digital object A; para. 0062, the audio of a reference digital object can be used as a training sample and input into a neural network model such as a generative adversarial network (GAN), and the video synchronized with the audio can be used as a sample label to train a speech-driven model of the reference digital object).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Xu’s features into Zavesky-Spreitzer’s invention for enhancing user’s playback experience by effectively synchronizing the enriched sound effect with video objects of a volumetric video content.
Regarding claim 3, Zavesky-Spreitzer-Xu discloses the method of claim 1, wherein the captured sounds are correlated with the plurality of objects and correlated with actions taken by the plurality of objects (Zavesky, para. 0066, the volumetric content analysis module 132 can utilize semantic classification concepts and technologies to classify scenes as dialog, suspense, action, or some other classification. The annotations 142 can identify the scenes with tags/labels according to the classifications. Analysis of a main view (e.g., default or original view) of the volumetric video 110 can be used to determine actors, objects, and various other contexts within and/or between scenes. Static inferences from the environment, time, landmarks, and the like depicted in the volumetric video 110 can also be used for additional annotations 142 and/or to refine existing annotations 142. The annotations 142 can provide a general sentiment depicted in the volumetric video 110, the audio 122, or both. The volumetric content analysis module 132 can associate the annotations 142 hierarchically with an entire scene, individual actors within a scene, individual objects within a scene, or the respective regions(s) 138 or voxels 136; Spreitzer, para’s 0016-0021, 0043-0050, and 0067-0073, scanning a video to find objects in the scene, such as a ball or a person’s head; tracking how those objects move over time. If the motion matches certain patterns, such as a collision, a sudden change in direction, or repeated bouncing, the system selects an audio or visual effect to match that motion. For example, it may add a “boom” sound at the moment of impact or slow down the video around a dramatic collision. It can also add music whose beat matches the rhythm of repeated motion, such as bouncing or juggling. The playback speed of the video or audio may be adjusted so the beat and motion line up better. Generating a new version of the video based on the edited result; Xu, para. 0046, a large amount of audio of the digital object A and video of the digital object A that matches the audio can be obtained (it can also be said that the actions and expressions in the video are synchronized with the audio, and in this disclosure it can also be called video synchronized with the audio). Then, the audio is used as a training sample and the video is used as a sample label to train the neural network model and obtain the speech-driven model of the digital object A).
The motivation to combine the references and obviousness arguments are the same as claim 1.
Regarding claim 5, Zavesky-Spreitzer-Xu discloses the method of claim 1, wherein additional context is added to the volumetric video based on movements of the plurality of objects within the volumetric video (Spreitzer, para’s 0016-0021, 0043-0050, and 0067-0073, scanning a video to find objects in the scene, such as a ball or a person’s head; tracking how those objects move over time. If the motion matches certain patterns, such as a collision, a sudden change in direction, or repeated bouncing, the system selects an audio or visual effect to match that motion. For example, it may add a “boom” sound at the moment of impact or slow down the video around a dramatic collision. It can also add music whose beat matches the rhythm of repeated motion, such as bouncing or juggling. The playback speed of the video or audio may be adjusted so the beat and motion line up better. Generating a new version of the video based on the edited result).
The motivation to combine the references and obviousness arguments are the same as claim 1.
Regarding claim 6, Zavesky-Spreitzer-Xu discloses the method of claim 1, wherein a distance of the plurality of objects generating the sounds is determined when dynamically modifying the volumetric video to synchronize the sound variations with the plurality of objects (Zavesky, para. 0066, the annotations 142 can provide a general sentiment depicted in the volumetric video 110, the audio 122, or both. The volumetric content analysis module 132 can associate the annotations 142 hierarchically with an entire scene, individual actors within a scene, individual objects within a scene, or the respective regions(s) 138 or voxels 136. The annotations 142 can identify position, time stamp, and orientation; Xu, para. 0066, distance between feature points).
The motivation to combine the references and obviousness arguments are the same as claim 1.
Claims 8, 10, 12-13, 15-16, and 18-19 are rejected the same reasons set forth in claim 1, 3 and 5-6. Zavesky-Spreitzer-Xu further discloses processor(s), memory module(s), and computer readable medium (see Zavesky, para. 0057).
7. Claims 2 and 9 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Zavesky-Spreitzer-Xu, as applied to claims 1 and 8 above, in view of Chopra et al. (US Publication 2023/0112369).
Regarding claim 2, Zavesky-Spreitzer-Xu discloses the method of claim 1,.
Zavesky-Spreitzer-Xu does not explicitly disclose but Chopra discloses wherein a knowledge corpus is created based on the captured sounds (Chopra, para. 0030, the database 102a can store a digital audio recording of the voice call and/or a transcript of the voice call (e.g. as generated by a speech-to-text module that converts the digital audio recording into unstructured text).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Chopra’s features into Zavesky-Spreitzer-Xu’s invention for enhancing user’s video editing experience by providing a knowledge database of predetermined audio recording and corresponding text transcripts.
Claim 9 is rejected for the same reasons set forth in claim 2.
8. Claims 4, 11, and 17 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Zavesky-Spreitzer-Xu, as applied to claims 1, 8, and 15 above, in view of Colbaugh et al. (US Publication 2013/0289401).
Regarding claim 4, Zavesky-Spreitzer-Xu discloses the method of claim 1,
Zavesky-Spreitzer-Xu does not explicitly disclose but Colbaugh discloses wherein the context of the captured sounds is based on effects on the captured sounds including a Doppler effect or sound echoes (Colbaugh para. 0077, the electrical signals generated by the ultrasonic transducer probe 112 based on the returned sound waves (echoes) are provided to processor 116 which examines the signals in order to identify in the signals the characteristic modulation (e.g., 30-40 Hz or some other frequency range or ranges) associated with OSA, if present, to determine whether the patient has OSA. Processor 116 is connected to display 120, which provides a visual indication of the result of the processing).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Colbaugh’s features into Zavesky-Spreitzer-Xu’s invention for enhancing user’s video editing experience by providing special effect to sound content of correlated video objects.
Claims 11 and 17 are rejected for the same reasons set forth in claim 4.
9. Claims 7, 14, and 20 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Zavesky-Spreitzer-Xu, as applied to claims 1, 8, and 15 above, in view of Lardaro et al. (US Publication 2020/0178852).
Regarding claim 4, Zavesky-Spreitzer-Xu discloses the method of claim 1.
Zavesky-Spreitzer-Xu does not explicitly disclose but Lardaro discloses wherein digital simulations of the sounds are performed to identify any sound that is related to an activity produced by one or more of the plurality of objects, sound effects, and environmental parameters (Lardaro, para. 0136, the hearing enhancement simulation performs digital audio processing in real time of some samples audio files (i.e. a recorded speech played against different noisy environments, such as, but not limited to: cocktail party, restaurant, church, at home with the TV on etc.) or of live ambient sounds).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Lardaro’s features into Zavesky-Spreitzer-Xu’s invention for enhancing user’s video editing experience by providing digital simulation of audio recordings.
Claims 14 and 20 are rejected for the same reasons set forth in claim 7.
10. The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. These include:
Khalid et al., US Publication 2019/0182471
Ponochevnyi, US Publication 2023/0260548
Conclusion
11. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LOI H TRAN whose telephone number is (571)270-5645. The examiner can normally be reached 8:00AM-5:00PM PST FIRST FRIDAY OF BIWEEK OFF.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, THAI TRAN can be reached at 571-272-7382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LOI H TRAN/Primary Examiner, Art Unit 2484