Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
[1] Claims 1, 10-12, 16, 17, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Sullivan et al. (“Sullivan”) [US 7573475 B2] in view of Menendez Gonzalez et al. (“Menendez Gonzalez”) [NPL titled, “SaiNet: Stereo aware inpainting behind objects with generative networks”].
Regarding claim 1, Sullivan discloses the following claim limitations:
1. A method (i.e. a method of creating a complementary stereoscopic image pair), comprising [Col. 1, lines 29-30]: receiving a first image (i.e. receives a first 2D image) of an object (i.e. objects in an image), the first image having a plurality of pixels representing the object that include a first pixel (i.e. for each pixel located at (X1, Y1)) [Col. 1, line 56; Col. 16, line 67; Col. 6, line 1]; generating a depth image from the first image (i.e. a depth map can be created and is based on a given image), the depth image including a depth value for a pixel corresponding to the first pixel (i.e. a 2D pixel at (X1, Y1) can use the depth value Z1 retrieved from the depth map indexed at (X1, Y1)) [Col. 5, lines 60-64; Col. 6, lines 7-10]; generating a reprojected image (i.e. projecting the first 2D image and then rendering a second 2D image that is stereoscopically complementary to the first 2D image) by moving at least content of the first pixel to a second pixel of the first image (i.e. a corresponding pixel location in the right camera view is computed) [Col. 3, lines 29-30; Col. 6, lines 34-35], the second pixel being identified based on the depth value (i.e. each pixel can have a corresponding depth value Z1), wherein the moving (i.e. a corresponding pixel location in the right camera view is computed) [Col. 6, lines 4-9; Col. 6, lines 34-35] (i.e. the first 2D image and the second 2D image can be viewed together) providing a three-dimensional representation (i.e. as a stereoscopic image pair) of the object (i.e. corresponding to objects) to a user [Col. 3, lines 33-36; Col. 3, line 25].
17. A computer program product (i.e. a computer program product) comprising a nontransitory storage (i.e. machine-readable storage device) medium [Col. 14, lines 61-63], the computer program product including (i.e. the computer program product includes instructions) code that, when executed (i.e. that are executed by a programmable processor) by processing circuitry [Col. 2, lines 1-2; Col. 14, lines 64-67], causes the processing circuitry to perform a method (i.e. method steps can be performed by a programmable processor executing a program of instructions), the method comprising [Col. 14, lines 65-67]: receiving a first image of an object (i.e. receives a first 2D image), the first image having a plurality of pixels (i.e. for each pixel) representing the object (i.e. corresponding to objects) that include a first pixel (i.e. located at (X1, Y1)) [Col. 1, line 56; Col. 6, line 1; Col. 16, line 67]; generating a depth image from the first image (i.e. a depth map can be created and is based on a given image), the depth image including a depth value for a pixel corresponding (i.e. each pixel can have a corresponding depth value Z1) to the first pixel [Col. 5, lines 60-64; Col. 6, lines 4-9]; generating a reprojected image (i.e. projecting the first 2D image and then rendering a second 2D image that is stereoscopically complementary to the first 2D image) by moving at least content of the first pixel to a second pixel (i.e. a corresponding pixel location in the right camera view is computed) of the first image [Col. 3, lines 29-30; Col. 6, lines 34-35], the second pixel being identified based on the depth value (i.e. each pixel can have a corresponding depth value Z1) wherein the moving (i.e. a corresponding pixel location in the right camera view is computed) [Col. 6, lines 4-9; Col. 6, lines 34-35] (i.e. the first 2D image and the second 2D image can be viewed together as a stereoscopic image pair) of the object (i.e. corresponding to objects) to a user [Col. 3, lines 33-36; Col. 3, line 25].
19. An apparatus, comprising: memory (i.e. memory 1220); and a processor (i.e. processor 1210) coupled to the memory (i.e. memory 1220 and processor 1210 are interconnected using a system bus 1250), the processor being configured (i.e. processor capable of processing instructions for execution) to [Fig. 12; Col. 14, lines 32-33; Col. 14, lines 33-35]: receive a first image of an object (i.e. receives a first 2D image), the first image having a plurality of pixels (i.e. for each pixel) representing the object (i.e. objects in an image) that include a first pixel (i.e. located at (X1, Y1)) [Col. 1, line 56; Col. 16, line 67; Col. 6, line 1]; generate a depth image from the first image (i.e. a depth map can be created and is based on a given image), the depth image including a depth value (i.e. a 2D pixel at (X1, Y1) can use the depth value Z1 retrieved from the depth map indexed at (X1, Y1)) for a pixel corresponding to the first pixel [Col. 5, lines 60-64]; generate a reprojected image (i.e. projecting the first 2D image and then rendering a second 2D image that is stereoscopically complementary to the first 2D image) by moving at least content of the first pixel to a second pixel of the first image (i.e. a corresponding pixel location in the right camera view is computed) [Col. 3, lines 29-30; Col. 6, lines 34-35], the second pixel being identified based on the depth value (i.e. each pixel can have a corresponding depth value Z1) [Col. 6, lines 4-9], wherein the moving (i.e. a corresponding pixel location in the right camera view is computed) [Col. 6, lines 34-35] (i.e. the first 2D image and the second 2D image can be viewed together as a stereoscopic image pair) of the object (i.e. corresponding to objects) to a user [Col. 3, lines 33-36; Col. 3, line 25].
Sullivan does not explicitly disclose the following claim limitations:
1. produces a mask defined by the plurality of pixels
and generating a second by inpainting the mask
17. produces a mask defined by the plurality of pixels; and generating a second image by inpainting the mask,
19. wherein the moving produces a defined by the plurality of pixels; and generate a second image by inpainting the mask,
However, in the same field of endeavor Menendez Gonzalez discloses the deficient claim limitations, as follows:
1. produces a mask (i.e. create a geometrically-meaningful bank of masks for inpainting) defined by the plurality of pixels (i.e. helps preserve the inpainting consistency at a pixel level) [Section 3.2, para. 2; Section 2, para. 4]
and generating a second image (i.e. Completed image) by inpainting the mask (i.e. the region behind the object (hole) that the network will inpaint) [Figure 2; Section 3.1, para. 1]
17. produces a mask (i.e. create a geometrically-meaningful bank of masks for inpainting) defined by the plurality of pixels (i.e. helps preserve the inpainting consistency at a pixel level) [Section 3.2, para. 2; Section 2, para. 4]; and generating a second image (i.e. Completed image) by inpainting the mask (i.e. the region behind the object (hole) that the network will inpaint) [Figure 2; Section 3.1, para. 1],
19. wherein the moving produces a mask (i.e. create a geometrically-meaningful bank of masks for inpainting) defined by the plurality of pixels (i.e. helps preserve the inpainting consistency at a pixel level) [Section 3.2, para. 2; Section 2, para. 4]; and generate a second image (i.e. Completed image) by inpainting the mask (i.e. the region behind the object (hole) that the network will inpaint) [Figure 2; Section 3.1, para. 1],
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teachings of Sullivan with Menendez Gonzalez to use the inpainting technique of Menendez Gonzalez to generate the second 2D image of Sullivan, the reasoning being to use meaningful and geometrically-consistent object masks to improve view synthesis in media production [Menendez Gonzalez: Section 1, para. 2].
Regarding claim 10, Sullivan discloses the following claim limitations:
10. The method as in claim 1, wherein mapping at least the content of the first pixel to the second pixel (i.e. remapping pixel information from the left image using a right-to-left offset map) based on the depth value (i.e. a 2D pixel at (X1, Y1) can use the depth value Z1 retrieved from the depth map indexed at (X1, Y1)) includes [Col. 12, lines 4-5; Col. 5, lines 60-64]: combining color values (i.e. a set of pixels that neighbor the location can be used to compute a pixel color for the right camera pixel) of pixels neighboring the second pixel and at least the content of the (i.e. for the location of (p, q), the neighboring pixels at (p+1, q), (p, q+1), and (p+1, q+1) can be used to calculate the right camera pixel color) first pixel [Col. 12, lines 28-30; Col. 12; lines 30-33].
Regarding claim 11, Menendez Gonzalez discloses the following claim limitations:
11. The method as in claim 10, wherein generating the reprojected image includes: determining a representation of a boundary of the object (i.e. generate context and synthesis regions from object boundaries) in the first image [Figure 3, caption]; and generating, as the mask (i.e. generate geometrically-valid masks), a set of pixels of the reprojected image (i.e. context and synthesis areas of an object) adjacent to the representation of the boundary (i.e. where the context area is the background surrounding the object and the synthesis area is a region behind the object) of the object [Section 3.1, para. 1; Section 3.1, para. 1].
Regarding claim 12, Menendez Gonzalez discloses the following claim limitations:
12. The method as in claim 1, wherein inpainting the mask (i.e. inpainting masks that represent real image occlusions) includes [Section 3.2, para. 1]: using an inpainting model to generate content (i.e. inpainting masks that represent real image occlusions) for pixels in the mask [Section 3.2, para. 1], the content being consistent with content of pixels outside of the mask (i.e. realistic and geometrically consistent fore ground object masks to explore inpainting behind objects) [Section 2, para. 4], the inpainting model being based on the reprojected image (i.e. stereo-context image is added to the input) [Section 3.1, para. 1].
Regarding claim 16, Menendez Gonzalez discloses the following claim limitations:
16. The method as in claim 1, wherein the mask (i.e. create a geometrically-meaningful bank of masks for inpainting) includes a set of pixels adjacent to the plurality of pixels (i.e. by detecting depth discontinuities along object boundaries) [Section 3.2, para. 2].
[2] Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Sullivan in view of Menendez Gonzalez further in view of Zobel et al. (“Zobel”) [US 20230216999A1].
Regarding claim 2, Menendez Gonzalez discloses the following claim limitations:
Sullivan and Menendez Gonzalez do not explicitly disclose the following claim limitations:
2. The method as in claim 1, (i.e. generating context and synthesis regions from object boundaries) [Section 3.2, para. 2];
However, in the same field of endeavor Zobel discloses the deficient claim limitations, as follows:
2. further comprising performing a postprocessing operation (i.e. post processing can be applied to clean up the depth values from the depth sensor to provide higher quality depth values) on the depth image by [para. 0201]:
and aligning pixels of the depth image with the representation of the edge of the (i.e. in depth based alignment 2710, the depth data and image data for each object is aligned) object [para. 0204].
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teaching of Sullivan and Menendez Gonzalez with Zobel to apply the post processing technique of Zobel to the depth data of Sullivan and Menendez Gonzalez, the motivation being to improve video frame quality using optical flow [Zobel: para. 0081].
[3] Claims 3, 4, 18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Sullivan in view of Menendez Gonzalez and further in view of Bhat et al. (“Bhat”) [NPL titled, “ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth”].
Regarding claims 3, 18, and 20, Sullivan in view of Menendez Gonzalez meets the claim limitations as set forth in claims 1, 17, and 19.
Sullivan discloses the following claim limitations:
3. The method as in claim 1, wherein the first image and the depth image have a first resolution (i.e. depth map can be a two-dimensional array that has the same resolution as the current image scene) [Col. 5, lines 66-67]; and wherein generating the depth image includes: (i.e. depth map can be a two-dimensional array that has the same resolution as the current image scene) [Col. 5, lines 66-67];
18. The computer program product as in claim 17, wherein the first image and the depth image have a first resolution (i.e. depth map can be a two-dimensional array that has the same resolution as the current image scene) [Col. 5, lines 66-67]; and wherein generating the depth image (i.e. a depth map can be created and is based on a given image) includes: (i.e. a depth map can be created and is based on a given image)
20. The apparatus as in claim 19, wherein the first image and the depth image have a first resolution (i.e. depth map can be a two-dimensional array that has the same resolution as the current image scene) [Col. 5, lines 66-67]; wherein the processor configured to generate the depth image is further (i.e. processor capable of processing instructions for execution) [Col. 14, lines 33-35] configured to:
Sullivan and Menendez Gonzalez do not explicitly disclose the following claim limitations:
3. generating a resized image by resizing the first image to have a second; using a first model to generate a first depth image from the first, the first depth image representing relative distances from a camera
and using a second to generate a second depth image from the resized image, the second depth image representing metric distances from the camera and having the second resolution; and wherein the depth image is generated based on the first depth image and the second depth image.
18. generating a resized image by resizing the first image to have a second resolution; using a first model to
the first depth image representing a function of relative distance from a camera
and using a second model to
the second depth image representing a function of metric distance from the camera and having the second resolution; and wherein the depth image is generated based on the first depth image and the second depth image.
20. generate a resized image by resizing the first image to have a second resolution; use a first model to generate a first depth image from the first image, the first depth image representing a function of relative distance from a camera
and use a second model to generate a second depth image from the resized image, the second depth image representing a function of metric distance from the camera and having the second resolution and wherein the depth image is generated based on the first depth image and the second depth image.
However, in the same field of endeavor Bhat discloses the deficient claim limitations, as follows:
3. generating a resized image by resizing the first image to have a second resolution (i.e. resizing the input to the training resolution) [A.1, para. 2]; using a first model (i.e. model used in the first stage) to generate a first depth image from the first image (i.e. MiDaS Depthmap generated based on image input) from the first image, the first depth image representing relative distances from a camera (i.e. relative depth estimation in the first stage) [Section 1, para. 3; Figure 2]
and using a second model (i.e. using metric depth estimation in the second stage) to generate a second depth image from the resized image (i.e. resizing the input to the training resolution), the second depth image representing metric distances from the camera (i.e. performing metric depth estimation in the second stage) and having the second resolution (i.e. training resolution) [Section 1, para. 3; A.1, para. 2]; and wherein the depth image (i.e. Zoe Depthmap) is generated based on the first depth image and the second depth image (i.e. generated based on the MiDaS Depthmap (Relative Depth) with Metric Bins) [Figure 2].
18. generating a resized image by resizing the first image (i.e. input is resized to the training resolution) to have a second resolution [A.1, para. 2]; using a first model (i.e. model used in the first stage) to generate a first depth image from the first image (i.e. MiDaS Depthmap generated based on image input) from the first image [Section 1, para. 3; Figure 2], the first depth image representing a function of relative distance from a camera (i.e. using relative depth estimation in the first stage) [Section 1, para. 3]
and using a second model (i.e. model used in the second stage) to generate a second depth image from the resized (i.e. input is resized to the training resolution) image [Section 1, para. 3; A.1, para. 2], the second depth image representing a function of metric distance from the camera (i.e. performing metric depth estimation in the second stage) and having the second resolution (i.e. training resolution) [Section 1, para. 3; A.1, para. 2]; and wherein the depth image (i.e. Zoe Depthmap) is generated based on the first depth image and the second depth image (i.e. generated based on the MiDaS Depthmap (Relative Depth) with Metric Bins) [Figure 2].
20. generate a resized image by resizing the first image to have a (i.e. input is resized to the training resolution) second resolution [A.1, para. 2]; use a first model (i.e. model used in the first stage) to generate a first depth image (i.e. MiDaS Depthmap generated based on image input) from the first image [Section 1, para. 3; Figure 2], the first depth image representing a function of relative distance from a camera (i.e. using relative depth estimation in the first stage) [Section 1, para. 3]
and use a second model (i.e. using metric depth estimation in the second stage) to generate a second depth image from the resized image (i.e. resizing the input to the training resolution) [Section 1, para. 3; A.1, para. 2], the second depth image representing a function of metric distance from the camera (i.e. performing metric depth estimation in the second stage) and having the second resolution (i.e. training resolution) [Section 1, para. 3; A.1, para. 2]; and wherein the depth image (i.e. Zoe Depthmap) is generated based on the first depth image and the second depth image (i.e. generated based on the MiDaS Depthmap (Relative Depth) with Metric Bins) [Figure 2].
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teachings of Sullivan and Menendez Gonzalez with Bhat to use the combined relative and metric depth estimation models of Bhat to generate the depth map in Sullivan and Mendez Gonzalez, the reasoning being so that the model used to generate the depth map has excellent generalization performance while maintaining a metric scale [Section 1, para. 3].
[4] Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Sullivan in view of Menendez Gonzalez further in view of Bhat and further in view of Wei et al. (“Wei”) [NPL titled, “FS-Depth: Focal-and-Scale Depth Estimation from a Single Image in Unseen Indoor Scene”].
Regarding claim 4, Bhat discloses the claim limitations, as follows:
4. The method as in claim 3, wherein generating the depth image further includes [Figure 2]: computing, via the second model (i.e. using metric depth estimation in the second stage), a normalized metric distance (i.e. decompose metric depth into normalized depth) [Section 1, para. 3; Section 3, para. 1]
Sullivan, Menendez Gonzalez, and Bhat do not explicitly disclose the claim limitations:
from the camera based on a focal length of the camera; and aligning the first depth image to the second depth image using the normalized metric distance from the camera.
However, in the same field of endeavor Wei discloses the deficient claim limitations, as follows:
from the camera based on a focal length (i.e. take focal length as input of the base model to well handle the focal-ambiguous problem) of the camera [Section 2.3, para. 1]; and aligning the first depth image to the second depth image (i.e. combining relative and absolute depth estimation) using the normalized metric distance from the camera (i.e. using normalized depth decoders) [Section 2.3, para. 1].
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teachings of Sullivan, Menendez Gonzalez, and Bhat with Wei so that the depth estimation model of Menendez Gonzalez incorporates Wei’s focal length variations in the network, the motivation being to alleviate serious deformation in 3D view [Wei: Figure 1, caption].
[5] Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Sullivan in view of Menendez Gonzalez further in view of Bhat further in view of Wei and further in view of Wofk et al. (“Wofk”) [US 20220343521 A1].
Regarding claims 5, Sullivan in view of Menendez Gonzalez, Bhat, and Wei meets the claim limitations as set forth in claim 4.
Sullivan, Menendez Gonzalez, and Bhat do not explicitly disclose the following claim limitations:
5. The method as in claim 4, wherein aligning the first depth image to the second depth image includes: determining a value of a scale parameter and a value of a shift parameter; and using the value of the scale parameter and the value of the shift parameter in aligning the first depth image to the second depth image.
However, in the same field of endeavor Wei discloses the deficient claim limitations, as follows:
5. The method as in claim 4, wherein aligning (i.e. global alignment) the first depth image to the second depth image includes [Fig. 1]: determining a value of a scale parameter and a value of a shift parameter (i.e. estimation for global scale and shift) [para. 0039]; and using the value of the scale parameter and the value of the shift parameter in aligning the first depth image to the second depth (i.e. global scale and shift alignment, where monocular depth estimates are fitted to metric sparse depth in a least-squares manner) image [paras. 0037, 0039].
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teachings of Sullivan, Menendez Gonzalez, Bhat and Wei with Wofk to use the estimated global scale and shift parameters of Wofk to align the first depth image and second depth image generated in Sullivan, the motivation being to make highly generalizable affine-invariant depth models more practical for integration into applications such as real-world sensor fusion, augmented reality and/or virtual reality (AR/VR) [Wofk: para. 0037].
[6] Claims 6 and 7 are rejected under 35 U.S.C. 103 as being unpatentable over Sullivan in view of Menendez Gonzalez further in view of Bhat and further in view of Ranftl et al. (“Ranftl”) [US 20220012848 A1].
Regarding claims 6 and 7, Sullivan in view of Menendez Gonzalez and Bhat meets the claim limitations as set forth in claim 3.
Bhat discloses the following claim limitations:
6. The method as in claim 3, wherein the first model includes an encoder and a decoder (i.e. common encoder-decoder architecture in the first stage) [Section 1, para. 3],
7. The method as in claim 3, wherein the second model includes an encoder and a decoder (i.e. encoder-decoder architecture in the second stage) [Section 1, para. 3],
Bhat does not explicitly disclose the following claim limitations:
6. the encoder configured to transform a portion of the first image into a token, the decoder configured to derive a portion of the first depth image from the token.
7. the encoder being configured to transform a portion of the resized image into a token, the decoder being configured to derive a portion of the second depth image from the token.
However, in the same field of endeavor Ranftl discloses the deficient claim limitations, as follows:
6. the encoder configured to transform (i.e. transformer encoder) a portion (i.e. image divider 202 divides the images into patches) of the first image into a (i.e. the readout token generator 210 generates special tokens 114ST) token [Fig. 2A; para. 0034; para. 0035], the decoder configured (i.e. a convolutional decoder) to derive a portion of the first depth image (i.e. reassembles the bag-of-words representation into image-like feature representations at various resolutions) from the token [para. 0026].
7. the encoder being configured to transform (i.e. transformer encoder) a portion of the resized image into a (i.e. image divider 202 divides the images into patches) token [Fig. 2A; para. 0034; para. 0035], the decoder being configured (i.e. convolutional decoder) to derive a portion of the second depth image (i.e. reassembles the bag-of-words representation into image-like feature representations at various resolutions) from the token [para. 0026].
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teachings of Sullivan, Menendez Gonzalez, and Bhat with Ranftl to use the encoder-decoder architecture of Bhat implements the methods of Raftl to transform the image patches into tokens with the encoder and reassemble them with the decoder, the reasoning being to mitigate the loss of feature resolution and granularity in the deeper stages of the dense prediction model [Ranftl: para. 0026]
[7] Claims 8 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Sullivan in view of Menendez Gonzalez and further in view of Li et al. (“Li”) [NPL titled, “Enforcing Temporal Consistency in Video Depth Estimation”].
Regarding claim 8, Sullivan in view of Menendez Gonzalez meets the claim limitations as set forth in claim 1.
Sullivan and Menendez Gonzalez do not explicitly disclose the following claim limitations:
8. The method as in claim 1, wherein the first image is an initial frame of a sequence of frames and the depth image is an initial depth frame corresponding to the initial frame; and wherein the method further comprises: receiving a next frame of the sequence of frames representing a next time step from the initial frame; and generating a next depth frame corresponding to the next frame based on the initial frame, the initial depth frame, and the next frame.
However, in the same field of endeavor Li discloses the deficient claim limitations, as follows:
8. The method as in claim 1, wherein the first image is an initial frame of a sequence of frames (i.e. Xi(i = 1, ..., n) is a video with a sequence of frames) and the depth image is an initial depth frame corresponding to the initial frame (i.e. with a corresponding sequence of depth predictions Di(i = 1, ..., n)) [Section 3.1, para. 2]; and wherein the method further comprises: receiving a next frame of the sequence of frames representing a next time step (i.e. considering any two consecutive frames Xi, Xi+1) from the initial frame [Section 3.1, para. 2]; and generating a next depth frame (i.e.
D
^
i
+
1
) corresponding to the next frame based on the initial frame (i.e. Xi), the initial depth frame (i.e. Di), and the next frame (i.e. Xi+1) [Figure 2].
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teaching of Sullivan and Menendez Gonzalez with Li to incorporate Li’s temporal consistency techniques to the image sequence processing of Sullivan and Menendez Gonzalez, the motivation being to reduce flickering while maintaining high depth quality [Li: Figure 1, caption].
Regarding claim 9, Li discloses the deficient claim limitations, as follows:
9. The method as in claim 8, further comprising: generating an optical flow (i.e. calculating the optical flow between two consecutive frames) based on the initial frame (i.e. Xi), the next frame (i.e. Xi+1), the initial depth frame (i.e. Di), and the next depth frame (i.e. Di+1) [Figure 2]; generating a warped next frame based (i.e. calculate the optical flow between two consecutive frames and use the flow vectors to warp one frame to align to another) on the optical flow [Figure 2, caption]; and combining the next frame and the warped next frame to produce a smoothed next depth frame (i.e. use warping operation w(·) to warp Di+1 as
D
^
i
+
1
=
w
(
D
^
i
+
1
,
f
i
+
1
→
i
)
so that it is now spatially aligned with Di) [Section 3.1, para. 2].
[8] Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Sullivan in view of Menendez Gonzalez and further in view of Suin et al. (“Suin”) [NPL titled, “Distillation-guided Image Inpainting”].
Regarding claim 13, Menendez Gonzalez discloses the following claim limitations:
13. The method as in claim 12, (i.e. generating the completed image) [Figure 2].
Sullivan and Menendez Gonzalez do not explicitly disclose the following claim limitations:
13. The method as in claim 12, wherein the inpainting model includes a knowledge distillation model configured to reduce latency in generating the second image.
However, in the same field of endeavor Suin discloses the deficient claim limitations, as follows:
13. The method as in claim 12, wherein the inpainting model includes a knowledge distillation model (i.e. using a knowledge distillation to make the inpainting network (IN) learn from the ideal features) configured to reduce latency (i.e. increases efficiency because the proposed training strategy does not require multiple generators or progressive refinement at inference time) [Section 1, para. 4; Section 1, para. 7]
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teaching of Sullivan and Menendez Gonzalez with Suin so that the inpainting model of Menendez Gonzalez includes the knowledge distillation of Suin, the motivation being to help the inpainting network to learn better while not requiring multiple generators or progressive refinement at inference time so that efficiency is improved [Suin: Section 1, para. 7].
[9] Claims 14 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Sullivan in view of Menendez Gonzalez and in further view of Chang et al. (“Chang”) [NPL titled, “VORNet: Spatio-temporally Consistent Video Inpainting for Object Removal”].
Regarding claim 14, Sullivan in view of Menendez Gonzalez meets the claim limitations as set forth in claim 12.
Sullivan and Menendez Gonzalez do not explicitly disclose the following claim limitations:
14. The method as in claim 12, wherein the reprojected image is a current reprojected frame of a sequence of reprojected frames; and wherein the method further comprises: receiving a set of previous reprojected frames of the sequence of reprojected frames, the set of previous reprojected frames having a set of corresponding masks; and generating the second image based on the set of previous reprojected frames and an inpainted reprojected image.
However, in the same field of endeavor Chang discloses the deficient claim limitations, as follows:
14. The method as in claim 12, wherein the reprojected image is a current reprojected frame of a sequence of reprojected frames (i.e. takes as input the video-with-target frames {It | t = 1...n}) [Section 3; para. 1]; and wherein the method further comprises: receiving a set of previous reprojected frames of the sequence of reprojected frames (i.e. input frames It and It-k received as inputs), the set of previous reprojected frames having a set of corresponding masks (i.e. each input frame has a corresponding foreground mask Mt) [Section 3.1, para. 2; Figure 2]; and generating the second image based on the set of previous reprojected frames (i.e. combine the information from previous frames and generated result in current frame) and an inpainted reprojected image (i.e. the inpainting network uses a generative model to create a possible image according to its surroundings) [Section 1, para. 4; Figure 2, caption].
It would have been obvious to one with ordinary skill in the art before the effective filing date of the invention to modify the teaching of Sullivan and Menendez Gonzalez with Chang to process the sequence of reprojected frames with the combines optical flow warping and image-based inpainting model of Chang, the motivation being to improve the perceptual quality and temporal stability [Chang: Section 1, para. 5].
Regarding claim 15, Chang discloses the following claim limitations:
15. The method as in claim 14, wherein generating the second image (i.e. generating the final output frame Ot) based on the set of previous reprojected frames (i.e. takes as input the video-with-target frames {It | t = 1...n}) and the inpainted reprojected image (i.e. inpainted frame Pt) includes [Section 3.3, para. 1; Section 3; para. 1; Figure 2]: computing an optical flow between the set of previous reprojected frames and the current reprojected frame (i.e. the warping network estimates the motion (optical flow) between the It and its kth previous frame It-k) [Figure 3, caption]; generating a set of warped previous reprojected frames (i.e. warped frame Wt-k →t) based on the optical flow (i.e. raw optical flow Fraw t-k →t) [Figure 2]; and combining the inpainted reprojected image (i.e. inpainted frame Pt) and the set of warped previous reprojected frames (i.e. warped frame Wt-k →t) to produce the second image (i.e. output frame Ot is generated using candidates from the warping network and the inpainting network) [Figure 2; Section 3.3, para. 1].
Conclusion
[10] Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jose B. Rodriguez whose telephone number is (571) 270-0829. The examiner can normally be reached Monday - Thursday 7:30 a.m. - 5 p.m., Friday 7:30 a.m. - 4 p.m. ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sath V. Perungavoor can be reached at (571) 272-7455. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/J.B.R./Examiner, Art Unit 2488 July 10, 2026
/SATH V PERUNGAVOOR/Supervisory Patent Examiner, Art Unit 2488