DETAILED ACTION
Information Disclosure Statement
1. The information disclosure statements submitted are being considered by the examiner.
Double Patenting
2. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
3. Claims 1-17, 19 and 20 are rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over claims 1-6, 8-10, 12-15 and 17 of U.S. Patent No. 12,217,768 in view of Rakshit, U.S. Patent Application Publication No. 2018/0210697 (hereinafter Rakshit).
Regarding claim 1 of the instant application, claim 1 of U.S. Patent No. 12,217,768 discloses a computer-implemented method, comprising:
receiving, by a computing device, an audio waveform associated with a plurality of video frames of video content;
determining, by a neural network and from the audio waveform, a neural representation comprising one or more audio frames, wherein each audio frame of the one or more audio frames comprises a respective plurality of coefficients, and wherein the respective plurality of coefficients represent one or more audio features in an encoded mixture of the audio waveform;
predicting, by the neural network and based on the neural representation, one or more estimated audio sources associated with the plurality of video frames;
associating, by the neural network, a time-invariant embedding of an estimated audio source of the one or more estimated audio sources with a spatio-temporal location of a video embedding;
determining, by the neural network, whether the estimated audio source corresponds to an on-screen object, or is an off-screen sound.
Still on the issue of claim 1 of the instant application, claim 8 of U.S. Patent No. 12,217,768 does not teach providing, by the computing device, a selectable user control to modify the estimated audio source. All the same, Rakshit discloses providing, by the computing device, a selectable user control to modify the estimated audio source (from abstract, see Based on a selection of a viewing perspective from which to view the scene, the method determines an audio mix for the audio portions given the selected viewing perspective). Therefore, it would have been obvious to one of ordinary skill in the art to modify claim 8 of U.S. Patent No. 12,217,768 with providing, by the computing device, a selectable user control to modify the estimated audio source as taught by Rakshit. This modification would have improved the viewer experience by providing more realistic sound as suggested by Rakshit.
Regarding claim 2, the combination of claim 8 of U.S. Patent No. 12,217,768 and Rakshit discloses receiving, by the computing device, a user selection of the selectable user control to modify the estimated audio source; and modifying the estimated audio source from the audio waveform (from abstract of Rakshit, see Based on a selection of a viewing perspective from which to view the scene, the method determines an audio mix for the audio portions given the selected viewing perspective).
Regarding claim 3, the combination of claim 8 of U.S. Patent No. 12,217,768 and Rakshit discloses wherein the modifying of the estimated audio source from the audio waveform comprises one or more of enhancing an audio content of the estimated audio source or deleting the estimated audio source from the audio waveform (from abstract of Rakshit, see Based on a selection of a viewing perspective from which to view the scene, the method determines an audio mix for the audio portions given the selected viewing perspective).
Regarding claim 4, claim 1 of U.S. Patent No. 12,217,768 discloses wherein the determining of the neural representation is performed by a time-domain convolutional masking network of the neural network.
Regarding claim 5, claim 2 of U.S. Patent No. 12,217,768 discloses predicting, by the neural network and for a given audio frame of the one or more audio frames, a respective mask, and wherein the predicting of the one or more audio sources comprises applying an inverse transform to a result of multiplying the respective mask with the plurality of coefficients corresponding to the given audio frame.
Regarding claim 6, claim 3 of U.S. Patent No. 12,217,768 discloses extracting, by a video embedding network, one or more visual features from each of the plurality of video frames; and generating, for each of the plurality of video frames, the video embedding based on the one or more visual features.
Regarding claim 7, claim 4 of U.S. Patent No. 12,217,768 discloses generating, by the video embedding network, a global video embedding comprising a global representation of the one or more visual features, and a plurality of spatio-temporal locations of the video content.
Regarding claim 8, claim 5 of U.S. Patent No. 12,217,768 discloses generating, by the video embedding network, a local video embedding comprising, for each video frame of the plurality of video frames, a temporal representation of the one or more visual features in the video frame.
Regarding claim 9, claim 6 of U.S. Patent No. 12,217,768 discloses the computer-implemented method of claim 8, further comprising: generating, by the video embedding network, a global video embedding based on a plurality of local video embeddings corresponding to the plurality of video frames.
Regarding claim 10, claim 8 of U.S. Patent No. 12,217,768 discloses generating, by the neural network, one or more audio embeddings corresponding to the one or more estimated audio sources; generating, by the neural network and for each audio embedding corresponding to the one or more estimated audio sources and based on the video embedding, a spatio-temporal audio-visual embedding based on an attention operation that aligns the one or more predicted audio sources with spatio-temporal positions of on-screen objects in the plurality of video frames, and wherein the determining of whether the estimated audio source corresponds to an on-screen object, or is an off-screen sound is based on the spatio-temporal audio-visual embedding.
Regarding claim 11, claim 9 of U.S. Patent No. 12,217,768 discloses the neural network comprises a classifier, wherein a first attention operation is applied to generate the one or more audio embeddings, a second attention operation is applied to generate the video embedding, and wherein the determining of whether the estimated audio source corresponds to an on-screen object, or is an off-screen sound comprises applying the classifier based on the one or more audio embeddings and the video embedding.
Regarding claim 12, claim 10 of U.S. Patent No. 12,217,768 discloses the neural network comprises a classifier, and the attention operation is applied to the one or more audio embeddings and the video embedding, to produce a representation, wherein the determining of whether the estimated audio source corresponds to an on-screen object, or is an off-screen sound comprises applying the classifier based on the representation.
Regarding claim 13, claim 12 of U.S. Patent No. 12,217,768 discloses the neural network comprises: an audio separation network to generate one or more estimated audio sources; and an audio embedding network to generate one or more audio embeddings based on the one or more estimated audio sources, wherein the one or more audio embeddings comprise a representation of audio features.
Regarding claim 14, claim 13 of U.S. Patent No. 12,217,768 discloses training the neural network to receive a particular audio waveform associated with a particular plurality of video frames and predict one or more particular audio sources in the particular audio waveform.
Regarding claim 15, claim 14 of U.S. Patent No. 12,217,768 discloses wherein the training further comprises: training the neural network to receive the particular audio waveform and predict a version of the particular audio waveform comprising particular audio sources that correspond to particular objects in the particular plurality of video frames.
Regarding claim 16, claim 15 of U.S. Patent No. 12,217,768 discloses wherein the training of the neural network comprises training a classifier based on active combinations cross entropy.
Regarding claim 17, claim 17 of U.S. Patent No. 12,217,768 discloses the training of the neural network comprises unsupervised mixture invariant training.
As per claims 19 and 20 of the instant application, although these claims are not identical to claim 8 of U.S. Patent No. 12,217,768, they are not patentably distinct from each other because claims 19 and 20 are respectively directed towards a computing device and an article of manufacture while 8 is directed towards a computer-implemented method. Indeed, the computing device and the article of manufacture of the instant application can be used to practice the computer-implemented method of the patent. And so, the pending claims are an obvious variant of the patented claims.
4. Claim 18 is rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over the combination of claim 13 of U.S. Patent No. 12,217,768 and Rakshit in further view of Nikolenko et al, U.S. Patent Application Publication No. 2020/0320346 (hereinafter Nikolenko).
On the issue of claim 18, the combination of claim 13 of U.S. Patent No. 12,217,768 and Rakshit does not clearly teach the training of the neural network is based on a training dataset comprising in-the-wild videos. All the same, Nikolenko discloses the training of the neural network is based on a training dataset comprising in-the-wild videos (from paragraph 0017, see In accordance with the various aspects and embodiments of the invention, “seed data” and “captured data” are used in relation to data that is real data. The real data may come from any source, including video, real dynamic images and real static images. In accordance with one aspect of the invention, real data includes real objects in any setting or environment, including native or natural environment or unnatural environment). Therefore, it would have been obvious to one of ordinary skill in the art to further modify the combination of claim 13 of U.S. Patent No. 12,217,768 and Rakshit wherein the training of the neural network is based on a training dataset comprising in-the-wild videos as taught by Nikolenko. This modification would have improved flexibility by allowing for the use of different types of datasets as suggested by Nikolenko.
Conclusion
5. Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLISA ANWAH whose telephone number is 571-272-7533. The examiner can normally be reached Monday to Friday from 8.30 AM to 6 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn Edwards can be reached on 571-270-7136. The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications.
Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2600.
Olisa Anwah
Patent Examiner
August 19, 2026
/OLISA ANWAH/Primary Examiner, Art Unit 2692