Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/26/2026 has been entered.
Response to Amendment
The amendment filed on 6/26/2026 has been entered and made of record. Claims 1, 12 and 21 are amended. Claims 1-21 are pending.
Response to Arguments
Applicant’s arguments with respect to claims 1, 12 and 21 have been considered but they are not persuasive.
Applicant contends that The cited references, alone or in combination, do not teach these limitations "generating, without generating a depth map for each
image, a feature map from the one or more input images, wherein the feature map comprises abstract features representing depth of one or more objects in the real-world environment; generating an occlusion mask from the feature map and the depth map for the virtual object but without any depth map corresponding to any of the one or more input images, ... , and wherein generating the occlusion mask employs temporal smoothing across timestamps of the one or more input images." (p. 1 of Remarks).
Examiner notices that applicant discloses “The feature network 410 inputs the images 405 to output a feature map 415 of the images. The feature map 415 encodes abstract features representing depth of one or more objects” in [0051]; “For example, the feature map 415 may be of lower dimensionality, which encodes the abstract depth features of the input images” in [0052]. It seems that the above argument is directed to generated feature map including abstract depth information without generating a depth map for each image.
Zhang discloses “At operation 630, the image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515” in [0070]; “The image encoder 720 can generate features that include information indicating which object is the occluding object and which object is the occluded object that would be completed” in [0076], see also [0091]; “At operation 1130, the image features can be decoded to obtain an occlusion mask for the occluding and occluded objects” in [0111]. Here, the detection of the objects with their positions can have abstract depth information to be used to determine which object is occluded without generating a depth map for each image, while an occlusion mask is generated based on feature map and the abstract depth information (which object is visible) without any depth map corresponding to any input image.
Applicant also alleges that Zhang and Daly are incompatible for combination and Veges is nonanalogous art (p. 10-11 of Remarks).
Examiner notices that Zhang teaches “The image encoder 720 can generate features that include information indicating which object is the occluding object and which object is the occluded object that would be completed” in [0076], and generating an occlusion mask for the overlapping visible region of the objects in [0071]. Here, Zhang teaches encoder generates feature map with abstract depth information without depth map generation for each input image. At the same field of image processing, Daly further teaches generating a depth map of a captured image in [0029] and an occlusion mask based on the depth in [0028]. The addition of Daly is merely the combination of known elements to perform their known function (generating a depth map of an input image). "The combination of familiar elements according to known methods is likely to be obvious when it does no more than yield predictable results." KSR Int 'l Co. v. Teleflex Inc., 550 U.S. 398,416 (2007). Further, Daly discloses the system may perform without spatial or temporal limitation in [0080], but Daly doesn’t directly teach temporal smoothing process. At the field of occluding technique, Veges further discloses temporal smoothing for 3D object pose estimation and localization for occluded object in Title, and solve the problem of smoothing coordinates for visible poses only and not handle occlusions at p. 3. Therefore, reference Veges is an analogous art at the field of occlusion processing application.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-21 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 2024/0169541 A1) in view of Sui et al. (CN 113205018 A), Gonzalez Morin et al. (US 2023/001444 A1, hereinafter Morin) and Veges et al. (Temporal Smoothing for 3D Human Pose Estimation and Localization for Occluded People, https://arxiv.org/abs/2011.00250)
As to Claim 1, Zhang teaches A computer-implemented method for generating a composite image including a virtual object placed in an image of a real-world environment, the method comprising:
receiving one or more input images captured by a camera assembly of a client device of the real-world environment (Zhang discloses “According to some aspects, an image generator 200 receives an original image including original content having an occluded portion and an occluding portion” in [0039]; “The image may be provided by a user and/or previously generated by an image generation model…” in [0109]. It is obvious that the input image can be an image captured by a camera of a user device.);
generating, without generating a depth map for each image, a feature map from the one or more input images, wherein the feature map comprises abstract features representing depth of one or more objects in the real-world environment (Zhang discloses “At operation 630, the image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515” in [0070], see also [0110]; “The image encoder 720 can generate features that include information indicating which object is the occluding object and which object is the occluded object that would be completed” in [0076], see also [0091]. Here, it is obvious that the generated feature map may include abstract depth information to be used to determine which object is occluded. For example, Sui discloses “The method based on the depth coding-decoding network has been widely applied in the building extraction. The coding part of this kind of network is mainly used for extracting depth abstract features…” at p. 3.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Zhang with the teaching of Sui so as to explain a depth coding-decoding network for extracting depth abstract features.
Zhang and Sui don’t directly teach depth map of virtual object. The combination of Morin further teaches following limitations:
accessing instructions for rendering augmented reality content inclusive of a virtual object to be placed into images captured of the real-world environment, wherein the instructions include a depth map for placement of the virtual object, wherein the depth map for the virtual object indicates a depth of each pixel of the virtual object (Zhang discloses “In some aspects, the output image combines additional content in a manner consistent with the original content” in [0049]; Morin further discloses “For example, a two-dimensional (2D) color and 2D depth image representation of the virtual content may be obtained. This can be obtained directly from an AR/MR application, e.g. from Unity. In the case of a 2D depth image, an (x,y) value may represent the pixel location, while a z value may represent the depth value for the corresponding pixel. FIGS. 5A and 5B illustrate a representation of a three-dimensional (3D) color depth map of the virtual content (FIG. SA) and a greyscale representation of the 2D depth image (FIG. 5B)” in [0043].);
generating an occlusion mask from the feature map and the depth map for the virtual object but without any depth map corresponding to any of the one or more input images, wherein the occlusion mask indicates one or more pixels of the virtual object that are occluded by an object in the real-world environment (Zhang discloses “At operation 1120, the image can be encoded to obtain image features and feature maps” in [0110]; “At operation 1130, the image features can be decoded to obtain an occlusion mask for the occluding and occluded objects” in [0111]. Here, the occluding or occluded objects can be either real object or virtual object. Morin further discloses “FIGS. 5A and 5B illustrate a representation of a three-dimensional (3D) color depth map of the virtual content (FIG. 5A)… Referring to FIG. SB, an example of information input to an algorithm is illustrated where the pixels of FIG. 5B mask the pixels from an RGB/depth input that may be further processed” in [0043]; “The operations further include rendering (906) a final composition of an augmented reality image containing the virtual object occluded by the occluding object based on applying the alpha mask to pixels in the at least one pixel classification image” in [0048], see also [0091]. Here, the detection of the objects with their positions can have abstract depth information to be used to determine which object is occluded without generating a depth map for each image, while an occlusion mask is generated based on feature map and the abstract depth information (which object is visible) without any depth map corresponding to any input image.);
generating the composite image based on a first input image at a current timestamp, the virtual object, and the occlusion mask; storing the composite image for subsequent display on an electronic display of the client device (Zhang discloses “and generating an output image including the original content from the original image and the additional content in the region using a diffusion model 1360 that takes the embedding vector as input” in [0125]. Morin also discloses “The device can perform further operations rendering a final composition of an augmented reality image containing the virtual object occluded by the occluding object based on applying the alpha mask to pixels in the at least one pixel classification image” in [0003]; see also Fig 4.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Zhang and Sui with the teaching of Morin so as to obtain a depth map of virtual content and use the depth map for rendering a final composition of an augmented reality image containing virtual object occluded by the occluding object based on occlusion mask (Morin, [0003]).
Zhang, Sui and Morin are silent on temporal smoothing. The combination of Veges further teaches following limitation:
wherein generating the occlusion mask employs temporal smoothing across timestamps of the one or more input images (Zhang discloses “An occlusion mask can be generated for the overlapping visible region of the occluding object 512 and occluded object 515” in [0071]. Veges further discloses “While temporal methods can still predict a reasonable estimation for a temporarily disappeared pose using past and future frames… We present an energy minimization approach to generate smooth, valid trajectories in time, bridging gaps in visibility” in Abstract; “Our second contribution is an energy minimization based smoothing function, targeting specifically those frames where a person became temporarily invisible. It adaptively smoothes the prediction stronger at frames where the pose is occluded and weaker when the pose is visible” at p.2.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Zhang, Sui and Morin with the teaching of Veges so as to use temporal methods to predict a reasonable estimation for a temporarily disappeared pose using past and future frames to generate smooth, valid trajectories in time, bridging gaps in visibility (Veges, Abstract).
As to Claim 2, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 1, wherein the one or more input images are frames from video data captured by the camera assembly (Morin discloses the usage of depth cameras and RGB cameras for a dynamic environment in [0026].)
As to Claim 3, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 1, wherein a dimensionality of the feature map is the same as a dimensionality of the one or more input images (Zhang, [0061, 0080]).
As to Claim 4, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 3, wherein the feature map is a matrix comprising features across a plurality of input images (Zhang discloses “In various embodiments, an input RGB image can be split into non-overlapping image patches by a patch splitting module. Each image patch is treated as a "token", where a feature is set as a concatenation of the raw pixel RGB values. For example, with a patch size of 4x4, the feature dimension of each patch would be 4x4x3=48 for the RGB image.” in [0080].)
As to Claim 5, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 1, wherein generating the feature map from the one or more input images comprises applying a trained feature network to the one or more input features to generate the feature map (Zhang discloses “image features generated by an encoder (i.e., latent diffusion)” in [0052]; “At operation 630, the image encoder can generate feature maps for the image 510, where the feature maps represent each object instance 512, 515.” in [0070]; “The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130…The image encoder 720 can generate one or more feature maps 740” in [0077].)
As to Claim 6, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 5, wherein the trained feature network is a neural network (Zhang discloses “The image encoder 720 can include a plurality of convolutional neural network (CNN) layers that forms a backbone of the mask network of the image processing system 130…The image encoder 720 can generate one or more feature maps 740” in [0077].)
As to Claim 7, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 1, wherein generating the occlusion mask from the feature map and the depth map for the virtual object comprises applying a mask predictor to the feature map and the depth map for the virtual object to generate the occlusion mask (Zhang discloses “In various embodiments, a mask network is used to provide an instance mask prediction and a rough occlusion prediction based on the features generated by an image encoder in the first step. The outputs of the first step are then provided to the diffusion model to perform amodal mask completion” in [0069]; see also Fig 6 & 9.)
As to Claim 8, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 7, wherein the mask predictor is a multi-layer perceptron (Zhang discloses multilayer perceptron 755 in Fig 7.)
As to Claim 9, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 1, wherein generating the occlusion mask comprises performing temporal smoothing with a previous occlusion mask generated for a second input image at a prior timestamp before the current timestamp (Zhang discloses “The instance mask 820 can be based on the object detection and feature masks previously generated by the mask network. An occluded region 785 can be identified in the image 510” in [0105]. Veges further discloses “While temporal methods can still predict a reasonable estimation for a temporarily disappeared pose using past and future frames… We present an energy minimization approach to generate smooth, valid trajectories in time, bridging gaps in visibility” in Abstract; “Our second contribution is an energy minimization based smoothing function, targeting specifically those frames where a person became temporarily invisible. It adaptively smoothes the prediction stronger at frames where the pose is occluded and weaker when the pose is visible” at p.2.)
As to Claim 10, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 1, wherein generating the composite image comprises: applying the occlusion mask to the virtual object to determine a portion of the virtual object that is in view; and placing the portion of the virtual object into the first input image to generate the composite image (Zhang discloses applying occlusion mask to two objects in Fig 7-9. Here, the occluding object can be a virtual object. For example, Morin discloses “Referring to FIG. 5B, an example of information input to an algorithm is illustrated where the pixels of FIG. 5B mask the pixels from an ROB/depth input that may be further processed” in [0043].)
As to Claim 11, Zhang in view of Sui, Morin and Veges teaches The computer-implemented method of claim 1, wherein the occlusion mask is generated further based on a depth map for a second virtual object, and wherein the composite image further includes the second virtual object (Morin discloses “2D depth image representation of the virtual content may be obtained… Referring to FIG. 5B, an example of information input to an algorithm is illustrated where the pixels of FIG. 5B mask the pixels from an ROB/depth input that may be further processed” in [0043]; a plurality of virtual objects in [0112-0117].)
Claim 12 recites similar limitations as claim 1 but in a computer readable storage medium form. Therefore, the same rationale used for claim 1 is applied.
Claim 13 is rejected based upon similar rationale as Claim 2.
Claim 14 is rejected based upon similar rationale as Claim 3.
Claim 15 is rejected based upon similar rationale as Claim 4.
Claim 16 is rejected based upon similar rationale as Claim 5.
Claim 17 is rejected based upon similar rationale as Claim 7.
Claim 18 is rejected based upon similar rationale as Claim 9.
Claim 19 is rejected based upon similar rationale as Claim 10.
Claim 20 is rejected based upon similar rationale as Claim 11.
Claim 21 recites similar limitations as claim 1 but in a system form. Therefore, the same rationale used for claim 1 is applied.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEIMING HE whose telephone number is (571)270-1221. The examiner can normally be reached on Monday-Friday, 8:30am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached on 571-272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WEIMING HE/
Primary Examiner, Art Unit 2611