DETAILED ACTION
Claims 1, 6-12, 17-18, 22-23 are pending.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 7/6/2026 has been entered.
Terminal Disclaimer
The application/patent being disclaimed has been improperly identified since the number used to identify the Co-pending application disclaimed is incorrect.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp.
Claims 1-23 rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of Co-pending Application 19/244,346 in view of Siagian et al. (US 11871068).
Although the claims at issue are not identical, they are not patentably distinct from each other because the instant application recites similar features and limitations as the patented application.
The instant App 18/991,074 and Co-pending Application 19/244,346 both recite features regarding extraction of video and audio frames and performing a synchronization detection technique..
The instant App 18/991,074 additionally recites the specific feature of: “facial image lists”.
Siagian teaches the specific feature of: “facial image lists” (i.e. neural network of facial images) (col. 14, lines 35-58).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the present invention to have provided specific facial image lists as taught by Siagian to the Co-pending Application 19/244,346 to detect speech (col. 2, lines 17-55).
Present Application 18/991,074
Co-pending Application 19/244,346
1. An audio and video synchronization detection method, comprising:
extracting, based on a preset interval, audio frames from a video segment of a target length and
generating an audio data list comprising the audio frames;
extracting, based on the preset interval, image frames from the video segment of the target length, and
generating an image data list comprising the image frames;
determining, from each image frame, one or more face images each having a face area greater than a preset threshold;
obtaining one or more face identifiers by tracking the one or more face images in each image frame based on face features;
marking, with each of the one or more face identifiers, image frames that contain a face image of the face identifier;
obtaining, for each of the one or more face identifiers, a respective face list by grouping image frames that contain the face identifier, wherein each face list corresponds to a respective person;
traversing, for each face list, face images in the face list, extracting mouth features of the face images,
generating a mouth feature list for the face list based on the mouth features, and determining the
mouth feature list containing mouth features corresponding to opening and closing changes;
obtaining, for each face list, a labial-sound similarity based on (i) the mouth feature list containing mouth features corresponding to opening and closing changes, and (ii) an audio feature sequence of the audio data list; and
performing, on the video segment, audio and video synchronization detection based on the labial-sound similarity for each face list to determine a synchronization result of the video segment.
1. An audio and video synchronization detection method, comprising:
extracting first image frames and first audio frames from a video;
obtaining a respective target type of each first image frame by identifying types of the first image frames;
determining a respective target audio and video synchronization detection algorithm according to the respective target type; and
performing audio and video synchronization detection on the first image frames and the first audio frames based on the target audio and video synchronization detection algorithms.
7. The method of claim 3, wherein in a case that the target audio and video synchronization detection algorithm is the labial-sound synchronization detection algorithm, performing the audio and video synchronization detection on the first image frames and the first audio frames based on the target audio and video synchronization detection algorithms, comprises:
obtaining one or more face identifications by performing face detection and tracking on the first image frames, and obtaining a plurality of image lists by dividing the first image frames into groups according to the face identifications;
determining respective mouth region pictures corresponding to each image list according to first image frames in each image list; and
performing the audio and video synchronization detection on the respective mouth region pictures corresponding to each image list and the first audio frames.
8. The method of claim 7, wherein performing the audio and video synchronization detection on the respective mouth region pictures corresponding to each image list and the first audio frames, comprises:
obtaining an audio feature sequence of the first audio frames by extracting audio features of the first audio frames;
obtaining a lip motion feature sequence corresponding to the image list by extracting lip motion features from the mouth region pictures corresponding to the image list;
obtaining a labial-sound similarity corresponding to the image list according to the lip motion feature sequence corresponding to the image list and the audio feature sequence; and
performing the audio and video synchronization detection on the video according to the labial-sound similarities corresponding to the image lists.
9. The method of claim 8, wherein obtaining the lip motion feature sequence corresponding to the image list by extracting the lip motion features from the mouth area pictures corresponding to the image list, comprises:
obtaining a mouth region picture sequence by ranking the mouth region pictures corresponding to the image list according to timestamps; and
obtaining the lip motion feature sequence corresponding to the image list by extracting lip motion features of the mouth region picture sequence.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 6-12, 17-18, 22-23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kinoshita (US 2015/0373414) in view of Siagian et al. (US 11871068), and in further view of Mathews (US 2022/0269922).
Claim 1, Kinoshita teaches An audio and video synchronization detection method, comprising:
extracting, based on the preset interval, image frames from the video segment of the target length (i.e. facial detection of a set amount of frames) (p. 0175-0179), and
generating an image data list comprising the image frames (i.e. candidate images) (p. 0037-0048);
determining, from each image frame, one or more face images each having a face area greater than a preset threshold (i.e. candidate image successfully detects a matching facial image in the video) (p. 0037-0048);
obtaining one or more face identifiers by tracking the one or more face images in each image frame based on face features (i.e. assigning an ID to the facial image and tracking movement) (p. 0096, 0147-0148);
marking, with each of the one or more face identifiers, image frames that contain a face image of the face identifier (i.e. selecting a face and the faces in the video are marked with box 205) (fig. 5B; p. 0131-0133);
obtaining, for each of the one or more face identifiers, a respective face list by grouping image frames that contain the face identifier (i.e. candidate selection process p. 0058-0062), wherein each face list corresponds to a respective person (i.e. each person is assigned an ID) (p. 0147-0148).
Kinoshita is silent regarding the specific feature of:
extracting, based on a preset interval, audio frames from a video segment of a target length and generating an audio data list comprising the audio frames;
traversing, for each face list, face images in the face list, extracting mouth features of the face images,
generating a mouth feature list for the face list based on the mouth features, and determining the mouth feature list containing mouth features corresponding to opening and closing changes;
obtaining, for each face list, a labial-sound similarity based on (i) the mouth feature list containing mouth features corresponding to opening and closing changes, and (ii) an audio feature sequence of the audio data list;
Siagian teaches the specific feature of:
extracting, based on a preset interval (i.e. transition points are known), audio frames from a video segment of a target length and generating an audio data list comprising the audio frames (i.e. audio signal durations for quiet or loud are used to determine dialogue) (col. 5-6, lines 30-58);
traversing, for each face list, face images in the face list, extracting mouth features of the face images (i.e. deep learning network is training to recognize open and closed mouths) (col. 8, lines 4-42, col. 14, lines 35-58),
generating a mouth feature list for the face list based on the mouth features (i.e. utilzing deep learning neural network), and determining the mouth feature list containing mouth features corresponding to opening and closing changes (i.e. open or closed mouth) (col. 8, lines 4-42, col. 14, lines 35-58).
“performing, on the video segment, audio and video synchronization detection for each face list to determine a synchronization result of the video segment” (i.e. determining a synchronization error) (col. 2, lines 17-55).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided transition points in video content as taught by Siagian to the system of Kinoshita to identify dialogue scenes (col. 5-6, lines 30-58).
Matthews teaches the specific feature of:
obtaining, for each face list, a labial-sound similarity based on (i) the mouth feature list containing mouth features corresponding to opening and closing changes, and (ii) an audio feature sequence of the audio data list (i.e. neural network training for speech sounds corresponding to mouth movements) (p. 0044).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided speech sound detection as taught by Mathews to the system of Siagian to provide video analysis of speech segments (p. 0044); and
Claim 6, Kinoshita is silent regarding The method of claim 1, wherein a process of determining the mouth feature list containing the mouth feature corresponding to the opening and closing change, comprises:
determining the mouth feature list containing the mouth feature corresponding to the opening and closing change by extracting lip movement features from the mouth feature list
Siagian teaches The method of claim 1, wherein a process of determining the mouth feature list containing the mouth feature corresponding to the opening and closing change, comprises:
determining the mouth feature list containing the mouth feature corresponding to the opening and closing change by extracting lip movement features from the mouth feature list (i.e. open or closed mouths training data of neural network) (col. 8, lines 4-42, col. 14, lines 35-58).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided synchronization in video content as taught by Siagian to the system of Kinoshita to correct audio in dialogue scenes (col. 5-6, lines 30-58).
Claim 7, Kinoshita is silent regarding the method of claim 1, wherein determining the audio feature sequence, comprises:
obtaining an audio frame sequence by ranking the audio frames based on time stamps of the audio frames in the audio data list; and
obtaining the audio feature sequence by extracting audio features from the audio frame sequence.
Mathews teaches the method of claim 5, wherein determining the audio feature sequence, comprises:
obtaining an audio frame sequence by ranking the audio frames based on time stamps of the audio frames in the audio data list (i.e. classification score and timestamps within metadata of frames (p. 0030); and
obtaining the audio feature sequence by extracting audio features from the audio frame sequence (i.e. obtaining from report) (p. 0030).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided speech sound detection as taught by Mathews to the system of Kinoshita to provide video analysis of speech segments (p. 0044).
Claim 8, Kinoshita is silent regarding the method of claim 5, wherein obtaining the labial-sound similarity comprises:
inputting the mouth feature list containing the mouth feature corresponding to the opening and closing change and the audio feature sequence into a pre-trained labial-sound synchronization detection model; and
obtaining the labial-sound similarity corresponding to the face image list by performing cross-modal similarity calculation on the mouth feature list containing the mouth feature corresponding to the opening and closing change and the audio feature sequence through the labial-sound synchronization detection model.
Mathews teaches the specific features of:
inputting the mouth feature list containing the mouth feature corresponding to the opening and closing change and the audio feature sequence into a pre-trained labial-sound synchronization detection model (i.e. neural network training for speech sounds corresponding to mouth movements) (p. 0044); and
obtaining the labial-sound similarity corresponding to the face image list by performing cross-modal similarity calculation (i.e. using neural network) on the mouth feature list containing the mouth feature corresponding to the opening and closing change and the audio feature sequence through the labial-sound synchronization detection model (i.e. neural network training for speech sounds corresponding to mouth movements) (p. 0044).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided speech sound detection as taught by Mathews to the system of Kinoshita to provide video analysis of speech segments (p. 0044).
Claim 9, Kinoshita is silent regarding the method of claim 1, wherein performing, on the video segment, audio and video synchronization detection based on the labial-sound similarity for each face image list to determine the synchronization result of the video segment performing audio and video, comprises:
determining a preset similarity threshold;
determining whether the labial-sound similarity corresponding to the face list is greater than the preset similarity threshold;
obtaining a statistical count of face image lists each with the labial-sound similarity greater than or equal to the preset similarity threshold;
performing audio and video synchronization detection on the video segment based on the statistical count, to determine the synchronization result of the video segment.
Siagian teaches the method of claim 1, “wherein performing, on the video segment, audio and video synchronization detection to determine the synchronization result of the video segment performing audio and video”, comprises
performing audio and video synchronization detection on the video segment based on the statistical count, to determine the synchronization result of the video segment (i.e. errors in synchronization) (col. 2, lines 17-55).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided transition points in video content as taught by Siagian to the system of Kinoshita to identify dialogue scenes (col. 5-6, lines 30-58).
Mathews teaches the specific features of:
“detection based on the labial-sound similarity for each face image list” (i.e. neural network training for speech sounds corresponding to mouth movements) (p. 0044);
determining a preset similarity threshold (i.e. threshold similarity values) (p. 0021);
determining whether the labial-sound similarity corresponding to the face list is greater than the preset similarity threshold (i.e. speech sounds are determined by a trained neural network which inherently have preset thresholds) (p. 0044);
obtaining a statistical count of face image lists each with the labial-sound similarity greater than or equal to the preset similarity threshold (i.e. neural network inherently person a comparison between training data and input data using at least statistical counts) (p. 0021, 0030, 0044).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided speech sound detection as taught by Mathews to the system of Kinotshita to provide video analysis of speech segments (p. 0044).
Claim 10, Kinoshita is silent regarding the method of claim 9, wherein performing audio and video synchronization detection on the video segment based on the statistical count, to determine the synchronization result of the video segment, comprises:
determining a preset quantity;
in response to the statistical count being greater than or equal to the preset quantity threshold, determining that the video segment is audio-video synchronized; and
in response to the statistical count being less than the preset quantity threshold, determining that the video segment is audio-video unsynchronized synchronized.
Siagian teaches the method of claim 9, wherein performing audio and video synchronization detection on the video segment based on the statistical count, to determine the synchronization result of the video segment, comprises:
determining a preset quantity threshold (i.e. threshold for detecting synchronization errors) (col. 2, lines 17-55, col. 5-6, lines 52-21);
in response to the statistical count being greater than or equal to the preset quantity threshold, determining that the video segment is audio-video synchronized (i.e. neural networks inherently utilize statistical counts, in this case for determining synchronization errors) (col. 2, lines 17-55, col. 5-6, lines 52-21); and
in response to the statistical count being less than the preset quantity threshold, determining that the video segment is audio-video unsynchronized synchronized (i.e. neural networks inherently utilize statistical counts, in this case for determining synchronization errors) (col. 2, lines 17-55, col. 5-6, lines 52-21).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided synchronization in video content as taught by Siagian to the system of Kinoshita to correct audio in dialogue scenes (col. 5-6, lines 30-58).
Claim 11, Kinoshita is silent regarding The method of claim 1, before determining the audio feature sequence, further comprising:
extracting key points of a mouth from a mouth feature image in the mouth feature list, and tracking the key points to obtain a motion trajectory of the key points; and
determining whether the mouth exhibits an opening and closing change based on the motion trajectory.
Siagian teaches The method of claim 1, before determining the audio feature sequence, further comprising:
extracting key points (i.e. facial movements) of a mouth from a mouth feature image in the mouth feature list, and tracking the key points to obtain a motion trajectory of the key points (i.e. open or closed mouths based on facial movements) (col. 2-3, lines 57-14, col. 8, lines 4-42, col. 14, lines 35-58); and
determining whether the mouth exhibits an opening and closing change based on the motion trajectory (i.e. open or closed mouths based on facial movements) (col. 2-3, lines 57-14, col. 8, lines 4-42, col. 14, lines 35-58).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing data of the present invention to have provided synchronization in video content as taught by Siagian to the system of Kinoshita to correct audio in dialogue scenes (col. 5-6, lines 30-58).
Claims 12 and 22 is analyzed and interpreted as an apparatus of claim 1.
Claim 17 is analyzed and interpreted as an apparatus of claim 6.
Claim 18 is analyzed and interpreted as an apparatus of claim 7.
Claim 23 recites “A non-transitory computer program product comprising a computer program, wherein when the computer program is executed by a processor” to perform the steps of claim 1. Siagian teaches “A non-transitory computer program product comprising a computer program, wherein when the computer program is executed by a processor” to perform the steps of claim 1 (col. 18, lines 4-29).
Response to Arguments
Applicant’s arguments with respect to claim(s) 1, 6-12, 17-18, 22-23 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Claims 1, 6-12, 17-18, 22-23 are rejected.
Inquiries
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MUSHFIKH I ALAM whose telephone number is (571)270-1710. The examiner can normally be reached 1:00PM-9:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Nasser Goodarzi can be reached at 571-272-4195. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
MUSHFIKH I. ALAM
Primary Examiner
Art Unit 2426
/MUSHFIKH I ALAM/Primary Examiner, Art Unit 2426 8/6/2026