DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/17/2026 has been entered.
Response to Arguments
Applicant's arguments filed 06/17/2026 have been fully considered but they are not persuasive.
Applicant argues, on page 9, that Huang does not teach or suggest a scene video comprising a body model having removed face content in a specified face region. The examiner respectfully disagrees. Huang teaches replacing a face region in a first image by another face from another image. It is obvious to one of ordinary skill in the art that the replacement process include removing the face region from the first image to replace it with the other face. Huang does not teach the replaced (or removed) face region belongs to a body model, but Gavade [0014] teaches replacing the images of the face of actor 112 in movie 104-1 with images of Mary's 106. Therefore, the argued limitation of scene video comprising a body model having removed face content in a specified face region is obvious in view of the combination of Huang and Gavade.
Applicant argues, on page 10, that Huang's "pre-processing" is performed on the target face image, not on the scene video/body model by citing Fig. 6 and claim 5 of Huang. The examiner respectfully disagrees . The Examiner did not cite Fig. 6 or claim 5 of Huang for teaching any claim limitation. Applicant argues further, on page 10, that In Applicant's amended claim, the pre-processed/provided item is the scene video/body model, and the pre- processing/provision concerns the specified face region of that body model being provided with removed face content and face position information. As explained above, providing a scene video comprising a body model having removed face content in a specified face region is included the replacement process of Huang and body model of Gavade.
Applicant argues, on page 10 paragraph, that Huang's "target region" is not Applicant's specified face region with removed face content. The examiner respectfully disagrees. Huang’s teaching of replacing a target region as a face region includes removing the content of the target region to replace it with the desire face. Applicant argues, on page 11regarding this point, that It is not a body-model specified face region in a scene video that has removed face content. However, as explained above Gavade teaches replacing the face of a body model.
Applicant argues, on page 11, that Huang does not teach face area information associated with the specified face region. The examiner respectfully disagrees. Huang teaches face area information including hair in Fig.2 and face contour in [0030] .
Applicant argues, on page12, that Huang does not teach fitting modified user face images with hair, wearables, or objects covering the specified face region . The examiner respectfully disagrees. Huang Fig. 2 teaches fitting modified user face images with hair covering the specified face region.
Applicant argues, on page 12, that Garrido does not cure Huang's deficiencies by not teaching a scene video comprising a body model of a person with removed face content in a specified face region. As explained above, the combination of Huang and Gavade teaches this argued limitation.
Applicant argues, on page 12, that Gavade does not cure Huang's deficiencies by not teaching a scene video comprising a body model of a person with removed face content in a specified face region. As explained above, the combination of Huang and Gavade teaches this argued limitation
Applicant argues, on page 15, that the cited combination does not teach receiving a scene video comprising a body model of a person. The examiner respectfully disagrees As explained above Gavade [0014] teaches a scene video comprising a body model of a person.
Applicant argues, on page 15 that the cited combination does not teach using a specified face region corresponding to a face location in different frames. The examiner respectfully disagrees. Huang [0025] teaches a face location in different frames.
Applicant argues, on page 15 that the cited combination does not teach generating modified first face images corresponding to the specified face region. The examiner respectfully disagrees. Garrido Section 6 teaches generating modified first face images.
Applicant argues, on page 15 that the cited combination does not teach receiving face area information associated with the specified face region. The examiner respectfully disagrees. Huang [0039] teaches this limitation.
Applicant argues, on page 15 that the cited combination does not teach processing the modified first face images and scene video frames using both face position information and face area information. The examiner respectfully disagrees Garrido teaches this limitation.
Applicant argues, on page 15 that the cited combination does not teach conditionally fitting the modified first face images with indicated hair/wearables/objects when such elements are present. The examiner respectfully disagrees. Huang Fig. 2 teaches this limitation.
Applicant argues, on page 15 that the cited combination does not teach replacing the removed face content with the modified first face images. The examiner respectfully disagrees. Garrido teaches this limitation.
Applicant argues, on page 16, that There is no articulated reason to combine the references in the claimed manner. The examiner respectfully disagrees. The examiner explained articulated reasons obvious to one of ordinary skill in the art to combine the references in the claimed manner.
Applicant argues, on page 16, that The proposed combination would require substantial redesign of Huang. The examiner respectfully disagrees. The proposed combination are obvious reasonable modification of Huang.
Applicant argues, on page 17, that The rejection relies on hindsight reconstruction. The examiner respectfully disagrees. In response to applicant's argument that the examiner's conclusion of obviousness is based upon improper hindsight reasoning, it must be recognized that any judgment on obviousness is in a sense necessarily a reconstruction based upon hindsight reasoning. But so long as it takes into account only knowledge which was within the level of ordinary skill at the time the claimed invention was made, and does not include knowledge gleaned only from the applicant's disclosure, such a reconstruction is proper. See In re McLaughlin, 443 F.2d 1392, 170 USPQ 209 (CCPA 1971).
Applicant argues, on page 18 with respect to Sharma And Seidel, that they do not teach previous argued limitation discussed above and the examiner explained that this limitation is obvious in view of Huang and Garrido.
Applicant argues, on page 19 with respect to Dreessen that it does not teach previous argued limitation discussed above and the examiner explained that this limitation is obvious in view of Huang and Gavade.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1, 3-10, 12, 14-18, and 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitation "the modified first face images" in line 15. There is insufficient antecedent basis for this limitation in the claim. Claims 3-10, 12, 14-18, and 20 are rejected for depending from claim 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6, 10-15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over HUANG (US 20190005305 A1) in view of Garrido et al. (Automatic Face Reenactment, 2014 IEEE Conference on Computer Vision and Pattern Recognition), and in further view of Gavade (US 20120005595 A1)
Regarding claim 1, Huang teaches A method for generating a video using user image/video comprising:
providing a first face image/video (Huang [0032] the target face image is an image designated by a user) based on user input to a processor (Huang [0023] a device for processing a video);
detecting and extracting the user face region and face position information by the processor from the first face image/video frames (obvious from Huang [0025] the first face image in the M frames is replaced with a target face image. Note: It is obvious to one of ordinary skill in the art before the effective filing date of the current application that the replacement requires prior detection and extraction of the target face image);
receiving a scene video and a face position information associated with a specified face region of a person in different frames of the scene video (Huang [0030] face features of each frame are extracted, Huang [0039] determine positions of facial features (such as eyes, eyebrows, the nose, the mouth, the outer contour of the face) based on the face recognition) from a data storage by the processor, (Huang [0005] performing target recognition on each frame in an input video to obtain M frames containing a first face image);
wherein the scene video is pre-processed or provided with removed face content in the specified face region and face position information for the different frames, such that the specified face region is replaceable by the modified first face images (obvious from Huang [0025] the first face image in the M frames is replaced with a target face image. Note: It is obvious to one of ordinary skill in the art before the effective filing date of the current application that the replacement includes removing face content in the face region to replace it with the modified first face image);
receiving face area information associated with the specified face region (Huang [0039] the outer contour of the face) in the different frames of the scene video, wherein the face area information comprises information of at least one of a boundary of the specified face region (Huang [0039] the outer contour of the face), a neck portion, a nearby portion around the specified face region, and, when present, hair, a head wearable, a neck wearable, a face wearable, or an object covering the specified face region (Huang Fig. 2 first output frame image showing head and hair);
wherein, when the face area information indicates hair, the head wearable, the neck wearable, the face wearable, or the object covering the specified face region, the processor processes the modified first face images to fit with the indicated hair, head wearable, neck wearable, face wearable, or object (Huang Fig. 2 first output frame image showing head and hair);
Whereas face region is the space for face with/without neck portion (Huang [0039] the outer contour of the face) and/or hair in scene person's image or video frame (Hung Fig. 2 showing first output frame image showing head and hair
Whereas face images is the space for face in person's image or video frame (Huang [0039] determine positions of facial features (such as eyes, eyebrows, the nose, the mouth, the outer contour of the face)), and
Wherein the face position information comprises a boundary of face region (Huang [0039] the outer contour of the face) and optionally at least one of tilt of face, orientation of face, geometrical location of face region, (Huang [0039] determine positions of facial features (such as eyes, eyebrows, the nose, the mouth, the outer contour of the face)) and zoom of the face region.
Huang et al.do not explicitly teach
wherein the scene video comprises a body model of a person ,and wherein the specified face region corresponds to a face location of the person in the different frames of the scene video;
modifying the first face image/ video frames to generate modified first face images corresponding to specified face region in different frames of scene video by the processor for maintaining orientation and scaling to match with the orientation and scaling of the specified face region of the scene video frames;
processing, by the processor, the modified first face images and the frames of the scene video using the face position information and the face area information to place the modified first face images at the specified face region of the body model;
replacing the removed face content in the specified face region of the body model in the scene video frames with the modified first face images with/without hair by the processor in a way to generate seamless single person image in each frame and subsequently applying the process for all frames to create a seamless video.
In a similar endeavor, Gavade teaches
wherein the scene video comprises a body model of a person (Gavade [0014] One scene of original movie 104-1 includes an actor 112 , Gavade [0024] Original content DB 212 may include a server to store content (e.g., video content) into which users may insert images and/or video of themselves (e.g., "original" content), ) and wherein the specified face region corresponds to a face location of the person in the different frames of the scene video (Gavade actor 112 in scene 104-1 in Fig. 1, Gavade [0014] the mixing engine may replace the images of the face of actor 112 in movie 104-1 with images of Mary's 106).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the examined application to have modified Huang et al. face replacement by incorporating Gavade et al. face replacement of an actor in scene video to arrive at the invention.
The motivation of doing so would have provided a user as an actor in a new movie..
The combination of Huang et al. and Gavade et al. does not teach
modifying the first face image/ video frames to generate modified first face images corresponding to specified face region in different frames of scene video by the processor for maintaining orientation and scaling to match with the orientation and scaling of the specified face region of the scene video frames;
processing, by the processor, the modified first face images and the frames of the scene video using the face position information and the face area information to place the modified first face images at the specified face region of the body model;
replacing the removed face content in the specified face region of the body model in the scene video frames with the modified first face images with/without hair by the processor in a way to generate seamless single person image in each frame and subsequently applying the process for all frames to create a seamless video.
In a similar endeavor, Garrido et al. teach
modifying the first face image/ video frames to generate the modified first face images corresponding to the second face images in different frames of scene video by the processor for maintaining a continuous facial expression along with orientation and scaling to match with the orientation and scaling of the second face images of the scene video frames (Garrido Section 6 we employ a 2D warping approach which combines global and local transformations to produce a natural shape deformation of the user’s face that matches the actor in the target sequence.);
processing, by the processor, the modified first face images and the frames of the scene video using the face position information and the face area information to place the modified first face images at the specified face region of the body model Garrido section 2 the new face with different identity needs to be inserted as naturally as possible in the original video, Garrido section 3 step 3 The target head pose is transferred to the selected source frames by warping the facial landmarks. A smooth transition is created by synthesizing in-between frames, and blending the source face into the target sequence using seamless cloning);
replacing the removed face content in the specified face region of the body model in the scene video frames with the modified first face images with/without hair by the processor in a way to generate seamless single person image in each frame and subsequently applying the process for all frames to create a seamless video (Garrido section 2 the new face with different identity needs to be inserted as naturally as possible in the original video, Garrido section 3 step 3 The target head pose is transferred to the selected source frames by warping the facial landmarks. A smooth transition is created by synthesizing in-between frames, and blending the source face into the target sequence using seamless cloning).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the examined application to have modified the combination of Huang et al. and Gavade et al. by incorporating Garrido et al. seamless swap to arrive at the invention.
The motivation of doing so would have produced seamless target video with the new face.
Regarding claim 3, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, wherein extracting the face image comprises extracting the face cropped with neck from user image/video (Huang Fig. 2 first output frame image showing head and neck).
Regarding claim 4, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, wherein extracting the face image comprises extracting the face cropped with hair from user image/video (Huang Fig. 2 first output frame image showing head and hair).
Regarding claim 5, The combination of Huang et al., Gavade et al., and Garrido teaches the method according of to claim 1, wherein extracting the face image comprises extracting the region from user image, which includes face (Huang [0032] the face features of the each of the M frames are replaced with the face features of the target face image).
Regarding claim 6, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, wherein extracting the face image based on an extraction input, wherein the extraction input comprises selection of at least one of face, hair, neck (Huang Fig. 2 first output frame image showing head, hair, and neck), region around face.
Regarding claim 10, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, wherein the video with user image/s is processed with background image/video to generate processed video (Huang Fig. 3 showing first output frame image comprising the target image background)
Regarding claim 12, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, wherein at least a set of user images are provided which show same face in two slightly different perspective and are processed with a set of video which shows same scene in slightly different perspective to generate a set of video which show user image with body of person in scene video in slightly different perspective (Garrido page 1 col 2 we adapt the head pose and face shape of the selected source frames to match those of the target).
The motivation of doing so would have produced a proper composite of the user’s face in the target video sequence.
Regarding claim 13, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, comprising:
- receiving a face area Information (Huang [0044] a face region image corresponding to the target feature point set is obtained)comprising information of at least an area showing hair, head and/ or neck wearable, face wearable and an object covering face of body model, in different frames of scene video (Huang [0039 feature point locating is performed on the first face image in a first frame in the M frames to obtain a first feature point set).. the points in FIG. 4 represent locations of feature points of the face image, in which each feature point corresponds to one feature value);
- processing the extracted face image and the frames of the scene video using the information of face position information and the face area information (Huang [0044] face region image is synthesized with the first frame after the face-swap to obtain M frames after face-swap).
Regarding claim 14, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, wherein the scene video frames comprises at least a vehicle, a background, a helmet, hair (Garrido page 1 col 1 conserving the hair, face outline, and skin color, as well as the background and illumination of the target video).
The motivation of doing so would have produced a proper composite of the user’s face in the target video sequence.
Regarding claim 15, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, wherein providing a group photo or a single person video or group video based on user input, and selecting a face based on selection user input (Huang [0051] the preset face image base includes a plurality of types of face images, and at least one target face image may be selected from the preset face image base), processing the group photo or the single person video or the group video based on the selection user input to generate the user image/video with face. (Huang [p0051] When a plurality of target face images are determined, an instruction for designating an image for face-swap may be received, such that a target face image to be finally converted is determined, or the plurality of target face images may be all converted and then provided to the user for selecting).
Regarding claim 17, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, comprising:
- processing the scene video to elect the face region of body model of the person in scene video by the processor from the scene video frames (Huang [0038] feature point locating is performed on the first face image in a first frame in the M frames to obtain a first feature point set); and
- generate face position information (Huang [0039] determine positions of facial features (such as eyes, eyebrows, the nose, the mouth, the outer contour of the face)).
Regarding claim 18, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1 but does not teach comprising:
- receiving skin tone input related to skin tone (Garrido page 1 col 1 replace the actor’s inner face region, while conserving the hair, face outline, and skin color) or detecting skin tone information from the face of the user image/video;
- providing the video frames of the body model in matching skin tone either from a database based on the skin tone input/ skin tone information or by processing the body model skin colour based on the skin tone input/ skin tone information in scene video frames (Garrido section 6.2 The lighting of the target sequence, and the skin appearance and hair of the target actor, should be preserved).
The motivation of doing so would have produced a proper composite of the user’s face in the target video sequence.
Regarding claim 19, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1 comprises:
- merging of extracted face image/s with the body model of person at neck (Huang [0039] the outer contour of the face) in the scene video frame/s (Huang [0070] performing the face-swap between any frame containing the first face image and the target face image to obtain the first output frame, and performing the image synthesis on each of the extracted target feature point).
Regarding claim 20, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1 comprising:
- processing the extracted face in different frame/s of scene video with at least one of environment lighting, (Garrido section 6.2 The lighting of the target sequence, and the skin appearance and hair of the target actor, should be preserved), shading, overlay glass effect on/around the face
The motivation of doing so would have produced a proper composite of the user’s face in the target video sequence.
Claims 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over HUANG in view of Garrido et al, in further view of Gavade., and in further view of Seidel et al. (US 20130330060 A1).
Regarding claim 7, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, but does not teach
wherein the scene video is provided as per the body shape and/ or size information provided by the user.
In a similar endeavor, Seidel et al. teach
wherein the scene video is provided as per the body shape and/ or size information provided by the user (Seidel [0041] The inventive reshaping interface allows the user to generate a desired 3D target shape, Seidel [0062] simulate the desired appearance of the actor on screen, even if his true body shape and proportions do not match the desired look).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the examined application to have modified combination of Huang et al., Gavade et al., and Garrido by incorporating Seidel et al. body reshape to arrive at the invention
The motivation of doing so would have satisfied the user desire.
Regarding claim 8, The combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, but does not teach
wherein the person's body in scene video is reshaped to be in different shape and size.
In a similar endeavor Seidel et al. teach
wherein the person's body in scene video is reshaped to be in different shape and size (Seidel [0048] able to perform a large range of semantically guided body reshaping operations on video data of many different formats that are typical in movie and video production).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the examined application to have modified the combination of Huang et al., Gavade et al., and Garrido by incorporating Seidel et al. body reshape to arrive at the invention
The motivation of doing so would have satisfied the user desire.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Huang et al. in view of Gavade et al, in further view of Garrido, and in further view of Dreessen (US 20180357472 A1)
Regarding claim 9, the combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, but does not teach
wherein the scene video comprises a background, and processing the scene video to remove the background.
In a similar endeavor, Dreessen teaches
wherein the scene video comprises a background, and processing the scene video to remove the background ([0194] Editing the target video for comparison can include identifying and removing all or part of a background).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the examined application to have modified the combination of Huang et al., Gavade et al., and Garrido by incorporating Dreessen background removal to arrive at the invention.
The motivation of doing so would have produced a better-quality video.
. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over HUANG in view of Gavade et al, in further view of Garrido., and in further of Sharma et al. (US 7734070 B1).
Regarding claim 16, the combination of Huang et al., Gavade et al., and Garrido teaches the method according to claim 1, but does not teach
wherein the scene video comprises more than one body model of persons, the method comprising:
- providing one or more user image/video having one or more faces and selecting faces for body models based on user selection input .
In a similar endeavor, Sharma et al. teach
wherein the scene video comprises more than one body model of persons (Sharma col 5 lines 45-46 selection menu for the available replaceable actors’ images through said interaction interfaces), the method comprising:
- providing one or more user image/video having one or more faces and selecting faces (Sharma col 5 lines 48 the replacing users' images are decided) for body models based on user selection input (Sharma col 5 lines 43-45 the person can manually select the replaceable actor’s images).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the examined application to have n modified the combination of Huang et al., Gavade et al., and Garrido by incorporating Garrido et al. conserving environment lighting to arrive at the invention.
The motivation of doing so would have allowed the user to select desirable images.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAID M ELNOUBI whose telephone number is (571)272-9732. The examiner can normally be reached Monday-Friday 9:30AM to 6:00PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kathy Wang-Hurst can be reached at 571-270-5371. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SAID M ELNOUBI/ Examiner, Art Unit 2644