DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 5, 7-8, 10, 13-15, 17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Jain (US 20250005824 A1) and in view of Yu (US 20210303912 A1).
Regarding to claim 1, Jain discloses a method for generating an image ([0056]: these methods takes a source image of a person and a target pose guidance as inputs to produce an edited output; an image processing apparatus with a robust deep learning framework for jointly modeling different components of a human; Fig. 4; [0062]: the multi-view network performs flow-based warping and visibility prediction 410 on the source images to warp the source images to conform to the target pose of the pose image; [0064]: produce warped images and a visibility map; [0070]: obtain multiple warped images and visibility maps; [0078]: warp the plurality of body part images to obtain a plurality of warped body part images), the method comprising:
receiving a source image including a source object having appearance features, wherein the appearance features include unique visual characteristics indicative of a visual identity of the source object (Fig. 3; [0056]: these methods takes a source image of a person and a target pose guidance as inputs to produce an edited output; Fig. 3; [0057]: an image processing apparatus receives a first image 310 of a face, a second image 315 of a top, and a third image 320 of a bottom as inputs;
PNG
media_image1.png
188
120
media_image1.png
Greyscale
; Fig. 4; [0062]: the multi-view network obtains and receive source images each depicting at least one body part and a pose image depicting a target pose;
PNG
media_image2.png
176
402
media_image2.png
Greyscale
);
receiving a target image including a target object having a target motion (Fig. 3; [0056]: these methods takes a source image of a person and a target pose guidance as inputs to produce an edited output; Fig. 3; [0058]: the target pose of the pose image 325 is a target motion at time t;
PNG
media_image3.png
176
106
media_image3.png
Greyscale
; Fig. 3; [0059]: the image processing apparatus receives a pose image 360 depicting a target pose, i.e. a target motion at time t, as input;
PNG
media_image4.png
182
112
media_image4.png
Greyscale
; Fig. 4; [0062]: the multi-view network obtains and receive source images each depicting at least one body part and a pose image depicting a target pose;
PNG
media_image2.png
176
402
media_image2.png
Greyscale
; [0068]: frontward facing; backward facing);
generating, based on the appearance features and the target motion, a flow field indicative of spatial transformations between the appearance features and the target motion ([0019]: virtual try-on to three-dimensional (3D) human and scene regeneration; [0066]: the warping stage takes the source image, the source pose, and the target pose as inputs and, predicts and generates two flow fields—F.sub.v and F.sub.t. F.sub.v represents displacements for regions of the source image that would remain visible when the pose changes to the target.);
warping, using the flow field, the appearance features to generate a warped image including a warped object having warped appearance features that mimic the target motion and preserve the visual identity of the source object ([0019]: virtual try-on to three-dimensional (3D) human and scene regeneration; Fig. 3; [0057]: the image processing apparatus generates the human image 335, i.e. a warped image, based on the first image 310, the second image 315, the third image 320, the pose image 325, and the selection mask 330;
PNG
media_image5.png
280
680
media_image5.png
Greyscale
; Fig. 4; [0062]: the multi-view network performs flow-based warping and visibility prediction 410 on the source images to warp the source images to conform to the target pose of the pose image and to identify visible regions of each of the warped images in a corresponding source image; the flow-based warping may be performed by a warping module 415; [0086]: the image processing apparatus warp each of the source images and generate texture embeddings for each of the source images; the image processing apparatus generates the composite image based on the texture embeddings and the pose encodings); and
providing the warped image as an output image (Fig. 3; [0057]: the image processing apparatus generates the human image 335, i.e. a warped image, based on the first image 310, the second image 315, the third image 320, the pose image 325, and the selection mask 330;
PNG
media_image5.png
280
680
media_image5.png
Greyscale
; Fig. 4; [0063]: generate a human image, i.e. warped image, with mixed and matched body parts from the source images based on the texture embeddings and the pose encodings;
PNG
media_image6.png
212
510
media_image6.png
Greyscale
; [0086]: the image processing apparatus warp each of the source images and generate texture embeddings for each of the source images; the image processing apparatus generates the composite image based on the texture embeddings and the pose encodings; Fig. 6; [0088]: the image processing apparatus use a machine learning model 625 to generate the composite image 635;
PNG
media_image7.png
140
134
media_image7.png
Greyscale
; Fig. 9; [0106]: generate and provide composite image 950 based on the selection mask and the body part images;
PNG
media_image8.png
158
774
media_image8.png
Greyscale
).
Jain fails to explicitly disclose a three-dimensional (3D) flow field.
In same field of endeavor, Yu teaches: a three-dimensional (3D) flow field ([0042]: a scene flow estimation engine 320 performs 3D scene flow estimation; [0043]: the estimated 3D scene flow field; [0051]: the scene flow estimation engine 320 estimates a 3D scene flow field; [0053]: the original feature maps of the historical 3D frames are warped to the reference frame, according to the 3D scene flow field for each of the historical 3D frames).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Jain to include a three-dimensional (3D) flow field as taught by Yu. The motivation for doing so would have been to warp the original feature maps of the historical 3D frames to the reference frame, according to the 3D scene flow field; to estimate 3D scene flow field of the corresponding frame; to obtain the warped 3D frame corresponding to each of frame; to produce an aggregated feature map, based on the reference frame and the estimated 3D scene flow field for each of the one or more historical 3D frames and perform the 3D semantic segmentation based on the aggregated feature map as taught by Yu in paragraphs [0053], [0054], and [0108].
Regarding to claim 3, Jain in view of Yu discloses the method of claim 1, wherein warping, using the 3D flow field, the appearance features to generate the warped image including a warped object having warped appearance features that mimic the target motion and preserve the visual identity of the source object (same as rejected in claim 1) comprises:
extracting, from the source image, a 3D feature volume of the appearance features (Yu; [0051]: the original feature map of frame.sub.i may be extracted from frame.sub.i by a 3D feature extract engine; [0053]: represent a 3D feature extract sub-network for extracting an original feature map for each frame);
warping, using the 3D flow field, the 3D feature volume of the appearance features to generate a warped 3D volume of the appearance features (Yu; [0053]: the feature maps are warped from historical 3D frame (frame.sub.j) to the reference frame; Fig. 5; [0051]: a warped feature map corresponding to each of frame.sub.i-1, . . . , frame.sub.i-k, is obtained by the feature warping engine 330 based on the estimated 3D scene flow field for the corresponding frame.); and
converting the warped 3D feature volume into the warped image (Yu; [0051]: 3D semantic segmentation is performed based on the aggregated feature map by the semantic segmentation engine 350;
PNG
media_image9.png
92
476
media_image9.png
Greyscale
; [0157]: predict a displacement of each point in point cloud data for the one or more historical 3D frames, based on the estimated 3D scene flow field for each of the one or more historical 3D frames; aggregate the warped feature map for each of the one or more historical 3D frames to an original feature map of the current 3D frame).
Same motivation of claim 1 is applied here.
Regarding to claim 5, Jain in view of Yu discloses the method of claim 1, further comprising:
refining the warped image to generate a refined warped image including a refined warped object having refined warped physical characteristic that mimic the target motion and preserve the visual identity of the source object (Jain; Fig. 3; [0057]: the image processing apparatus generates the human image 335 based on the first image 310, the second image 315, the third image 320, the pose image 325, and the selection mask 330;
PNG
media_image5.png
280
680
media_image5.png
Greyscale
; [0056]: these methods takes a source image of a person and a target pose guidance as inputs to produce an edited output; an image processing apparatus with a robust deep learning framework for jointly modeling different components of a human; deep learning framework refines warped images; Fig. 4; [0063]: generate a human image with mixed and matched body parts from the source images based on the texture embeddings and the pose encodings;
PNG
media_image6.png
212
510
media_image6.png
Greyscale
; GAN-based rendering refines warped images as illustrated in Fig. 4); and
providing the refined warped image as the output image (Jain; Fig. 3; [0057]: the image processing apparatus generates and prides the human image 335 based on the first image 310, the second image 315, the third image 320, the pose image 325, and the selection mask 330;
PNG
media_image5.png
280
680
media_image5.png
Greyscale
; Fig. 4; [0063]: generate and provide a human image with mixed and matched body parts from the source images based on the texture embeddings and the pose encodings;
PNG
media_image6.png
212
510
media_image6.png
Greyscale
).
Regarding to claim 7, Jain in view of Yu discloses the method of claim 1, wherein the source object is a source individual associated with a source identity, wherein the target object is a target individual associated with a target identity (Jain; Fig. 4; [0062]: the multi-view network obtains and receive source images each depicting at least one body part and a pose image depicting a target pose;
PNG
media_image2.png
176
402
media_image2.png
Greyscale
), and wherein the source identity and the target identity are at least one of:
a matching identity, or a different identity (or is optional; Jain; [0074]: combine images of different people to produce a new human image; a person's identity, e.g., face and hair; [0075]: a virtual try-on, and changing an ID result in identity swapping; Fig. 9; [0105]: a target image 905, a first source image 910, a second source image 915, and a third source image 920 as inputs;
PNG
media_image10.png
156
750
media_image10.png
Greyscale
).
Regarding to claim 8, Jain discloses an image generation system (Fig. 2; [0033]: the memory unit 210; [0035]: a memory unit 210 include random access memory (RAM), read-only memory (ROM), or a hard disk; Fig. 2; [0034]: the processor unit 205 executes computer-readable instructions stored in a memory to perform various functions; [0056]: these methods takes a source image of a person and a target pose guidance as inputs to produce an edited output; an image processing apparatus with a robust deep learning framework for jointly modeling different components of a human; Fig. 4; [0062]: the multi-view network performs flow-based warping and visibility prediction 410 on the source images to warp the source images to conform to the target pose of the pose image; [0064]: produce warped images and a visibility map; [0070]: obtain multiple warped images and visibility maps; [0078]: warp the plurality of body part images to obtain a plurality of warped body part images), comprising:
one or more memories (Fig. 2; [0033]: the memory unit 210; [0035]: a memory unit 210 include random access memory (RAM), read-only memory (ROM), or a hard disk); and
one or more processors, communicably coupled to the one or more memories, configured to (Fig. 2; [0034]: the processor unit 205 executes computer-readable instructions stored in a memory to perform various functions):
the rest claim limitations are similar to claim limitations recited in claim 1. Therefore, same rational used to reject claim 1 is also used to reject claim 8.
Regarding to claim 10, Jain in view of Yu discloses the image generation system of claim 8, wherein the one or more processors,
The rest claim limitations are similar to claim limitations recited in claim 3. Therefore, same rational used to reject claim 3 is also used to reject claim 10.
Regarding to claim 13, Jain in view of Yu discloses the image generation system of claim 8, wherein the one or more processors are configured to:
The rest claim limitations are similar to claim limitations recited in claim 5. Therefore, same rational used to reject claim 5 is also used to reject claim 13.
Regarding to claim 14, Jain in view of Yu discloses the image generation system of claim 8,
The rest claim limitations are similar to claim limitations recited in claim 7. Therefore, same rational used to reject claim 7 is also used to reject claim 14.
Regarding to claim 15, Jain discloses a non-transitory computer-readable medium storing a set of instructions (Fig. 2; [0033]: the memory unit 210; [0035]: a memory unit 210 include random access memory (RAM), read-only memory (ROM), or a hard disk; Fig. 2; [0034]: the processor unit 205 executes computer-readable instructions stored in a memory to perform various functions; [0056]: these methods takes a source image of a person and a target pose guidance as inputs to produce an edited output; an image processing apparatus with a robust deep learning framework for jointly modeling different components of a human; Fig. 4; [0062]: the multi-view network performs flow-based warping and visibility prediction 410 on the source images to warp the source images to conform to the target pose of the pose image; [0064]: produce warped images and a visibility map; [0070]: obtain multiple warped images and visibility maps; [0078]: warp the plurality of body part images to obtain a plurality of warped body part images), the set of instructions comprising:
one or more instructions that, when executed by one or more processors of an image generation system, cause the image generation system to (Fig. 2; [0033]: the memory unit 210; [0035]: a memory unit 210 include random access memory (RAM), read-only memory (ROM), or a hard disk; Fig. 2; [0034]: the processor unit 205 executes computer-readable instructions stored in a memory to perform various functions):
the rest claim limitations are similar to claim limitations recited in claim 1. Therefore, same rational used to reject claim 1 is also used to reject claim 15.
Regarding to claim 17, Jain in view of Yu discloses the non-transitory computer-readable medium of claim 15, wherein the one or more instructions that cause the image generation system to
The rest claim limitations are similar to claim limitations recited in claim 3. Therefore, same rational used to reject claim 3 is also used to reject claim 17.
Regarding to claim 20, Jain in view of Yu discloses the non-transitory computer-readable medium of claim 15, wherein the one or more instructions further cause the image generation system to:
The rest claim limitations are similar to claim limitations recited in claim 5. Therefore, same rational used to reject claim 5 is also used to reject claim 20.
Claims 2, 9, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Jain (US 20250005824 A1) in view of Yu (US 20210303912 A1), and further in view of Ha (US 20220284663 A1).
Regarding to claim 2, Jain in view of Yu discloses the method of claim 1, wherein the target image is a target frame of a target video, and wherein the output image is an output frame of an output video (Yu; Fig. 3; [0041]: the time-ordered sequence of 3D frames, i.e. video; a point cloud data for the rear of a car; Fig. 3; [0045]: the points on framei belongs to a car;
PNG
media_image11.png
128
314
media_image11.png
Greyscale
; [0058]: Flow-Guided Feature Aggregation for Video Object Detection which is incorporated herein by reference in its entirety; [0119]: video and audio; [0121]: video camera). Same motivation of claim 1 is applied here.
Jain in view of Yu fails to explicitly disclose: face reenactment.
In same field of endeavor, Ha teaches face reenactment ([0055]: video conference and face reenactment; [0058]: the object is a face; a face image includes a face region; [0064]: various types of moving objects; [0068]: perform a backward warping operation on each of the albedo data and the depth data based on the target shape deformation value; [0075]: when the object is a face, the deformation may correspond to a facial expression).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Jain in view of Yu to include face reenactment as taught by Ha. The motivation for doing so would have been to output video conference and face reenactment; to warp original images into the canonical space which corresponds to a normalized pose space; to perform a backward warping operation on each of the albedo data and the depth data based on the target shape deformation value as taught by Ha in paragraphs [0055], [0059] and [0068].
Regarding to claim 9, Jain in view of Yu discloses the image generation system of claim 8,
The rest claim limitations are similar to claim limitations recited in claim 2. Therefore, same rational used to reject claim 2 is also used to reject claim 9.
Regarding to claim 16, Jain in view of Yu discloses the non-transitory computer-readable medium of claim 15,
The rest claim limitations are similar to claim limitations recited in claim 2. Therefore, same rational used to reject claim 2 is also used to reject claim 16.
Claims 4, 6, 11-12, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Jain (US 20250005824 A1) in view of Yu (US 20210303912 A1), and further in view of Lin (US 20250225662 A1).
Regarding to claim 4, Jain in view of Yu discloses the method of claim 1, further comprising:
Jain in view of Yu fails to explicitly disclose:
regularizing motion estimation of the 3D flow field using a cyclic warp loss.
In same field of endeavor, Lin teaches:
regularizing motion estimation of the 3D flow field using a cyclic warp loss ([0079]: the reverse warping engine 404 is configured to reverse warp the feature map; [0080]: the reverse warping engine 404 generates warped optical flow information; [0081]: the warped optical flow information F′.sub.2,i and the feature map F.sub.1,0 are provided to the correction engine 406 for reverse optical flow error correction; [0095]: apply cyclic warping in both directions in an image pair; cyclic warping generates two estimated optical flows; the estimated optical flows is used to identify if an object is occluded; [0100]: to improve the detection based on additional criteria, e.g., progression by the progression engine 541, sparsity by the sparsity engine 542, flow magnitude by the flow-guided engine 543, etc.; Fig. 7C; [0103]: FIG. 7C illustrates the same optical flow of the two consecutive images using the same optical flow engine, but with additional reverse optical flow error correction).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Jain in view of Yu to include regularizing motion estimation of the 3D flow field using a cyclic warp loss as taught by Lin. The motivation for doing so would have been to reverse optical flow error correction; to improve the detection of defective optical flows; to apply cyclic warping in both directions in an image pair; to generate two estimated optical flows by cyclic warping; improve the detection based on additional criteria, e.g., progression by the progression engine 541, sparsity by the sparsity engine 542, flow magnitude by the flow-guided engine 543, etc. as taught by Lin in paragraphs [0081], [0084], [0095], and [0100].
Regarding to claim 6, Jain in view of Yu discloses the method of claim 1, wherein the appearance features are included in a source foreground region of the source image, wherein the source image includes a source background region (Jain; Fig. 3; [0057]: an image processing apparatus receives a first image 310 of a face, a second image 315 of a top, and a third image 320 of a bottom as inputs;
PNG
media_image1.png
188
120
media_image1.png
Greyscale
; Fig. 4; [0062]: the multi-view network obtains and receive source images each depicting at least one body part and a pose image depicting a target pose;
PNG
media_image2.png
176
402
media_image2.png
Greyscale
; [0089]: high-quality person images with plain backgrounds), and wherein the method further comprises:
separating the source foreground region from the source background region (Jain; [0089]: high-quality person images with plain backgrounds; Fig. 4; [0074]: the foreground regions include a source of a person's identity, e.g., face and hair, I.sub.id, upper clothing I.sub.upp, and lower clothing I.sub.low in that order;
PNG
media_image12.png
182
150
media_image12.png
Greyscale
);
adding the source background region to the warped image (Jain; Fig. 4; [0063]: , a GAN-based rendering component 430 may generate a human image with mixed and matched body parts from the source images based on the texture embeddings and the pose encodings; [0089]: high-quality person images with plain backgrounds;); and
Jain in view of Yu fails to explicitly disclose:
inpainting blank spaces created via translation between the appearance features and the warped appearance features.
In same field of endeavor, Lin teaches:
inpainting blank spaces created via translation between the appearance features and the warped appearance features (Fig. 7A; Fig. 7B; Fig. 7C; [0103]: a region 710 is missing flow information that is present in FIG. 7A and corresponds to at least one region of the optical flow information including defective optical flows; in FIG. 7C, the region 710 includes flow information that is missing in FIG. 7B; the reverse optical flow error correction prevents the introduction of errors that can cascade and compound downstream in the optical information).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Jain in view of Yu to include inpainting blank spaces created via translation between the appearance features and the warped appearance features as taught by Lin. The motivation for doing so would have been to reverse optical flow error correction; to improve the detection of defective optical flows; to apply cyclic warping in both directions in an image pair; to generate two estimated optical flows by cyclic warping; improve the detection based on additional criteria, e.g., progression by the progression engine 541, sparsity by the sparsity engine 542, flow magnitude by the flow-guided engine 543 as taught by Lin in paragraphs [0081], [0084], [0095], and [0100].
Regarding to claim 11, Jain in view of Yu discloses the image generation system of claim 8, wherein the one or more processors are configured to:
The rest claim limitations are similar to claim limitations recited in claim 4. Therefore, same rational used to reject claim 4 is also used to reject claim 11.
Regarding to claim 12, Jain in view of Yu discloses the image generation system of claim 8,
The rest claim limitations are similar to claim limitations recited in claim 12. Therefore, same rational used to reject claim 6 is also used to reject claim 12.
Regarding to claim 18, Jain in view of Yu discloses the non-transitory computer-readable medium of claim 15, wherein the one or more instructions further cause the image generation system to:
The rest claim limitations are similar to claim limitations recited in claim 4. Therefore, same rational used to reject claim 4 is also used to reject claim 18.
Regarding to claim 19, Jain in view of Yu discloses the non-transitory computer-readable medium of claim 15,
The rest claim limitations are similar to claim limitations recited in claim 6. Therefore, same rational used to reject claim 6 is also used to reject claim 19.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Hai Tao Sun whose telephone number is (571)272-5630. The examiner can normally be reached 9:00AM-6:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 5712727642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HAI TAO SUN/Primary Examiner, Art Unit 2616