Prosecution Insights
Last updated: October 01, 2026
Application No. 19/077,050

Method and Device for Multi-Camera Hole Filling

Non-Final OA §103
Filed
Mar 11, 2025
Priority
Jun 29, 2020 — provisional 63/045,394 +3 more
Examiner
GE, JIN
Art Unit
Tech Center
Assignee
Apple Inc.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
440 granted / 552 resolved
+19.7% vs TC avg
Strong +19% interview lift
Without
With
+18.8%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
26 currently pending
Career history
572
Total Applications
across all art units

Statute-Specific Performance

§101
10.6%
-29.4% vs TC avg
§103
62.0%
+22.0% vs TC avg
§102
11.0%
-29.0% vs TC avg
§112
9.5%
-30.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 552 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on 03/11/2025, 06/09/2025, 12/30/2025, and 07/15/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5, 9-13, 17-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2019/0101758 to Zhu et al. in view of U.S. PGPubs 2020/0302688 to Hosfield et al.. Regarding claim 1, Zhu et al. teach a method comprising (par 0006): obtaining a first image from a first image sensor communicatively coupled to a computing system with non-transitory memory and one or more processors (par 0009, “the left camera is used to capture a raw left image of the surrounding environment and the right camera is used to capture a raw right image of the surrounding environment”, par 0039, “a HMD is configured with a stereo camera pair which comprises a left and right camera. The stereo camera pair is used to capture images of a surrounding environment, wherein a center-line perspective of an image captured by the left camera has a non-parallel alignment with respect to a center-line perspective of an image captured by the right camera (i.e. the cameras are angled with respect to one another so that their fields of view are not parallel). The left camera captures a raw left image of the surrounding environment and the right camera captures a raw right image of the surrounding environment”, Fig 1, par 0046-0047, “The computer system 100 may also be connected (via a wired or wireless connection) to external sensors 130 (e.g., one or more remote cameras, accelerometers, gyroscopes, acoustic sensors, magnetometers, etc.). Further, the computer system 100 may also be connected through one or more wired or wireless networks 135 to remote systems(s) 140 that are configured to perform any of the processing described with regard to computer system 100.”, par 0067, “ the left camera 605 and the right camera 615 are able to record the surrounding environment”), wherein the first image sensor has a first perspective of an environment (par 0061-0067, “these components operate to reconstruct a perspective captured by a camera image so that the captured perspective matches a perspective of the user ….FIG. 6 presents an abstract view of an inside-out head tracking system. As shown, a head-mounted device 600 includes a stereo camera pair (i.e. the left camera 605 and the right camera 615). Here, the left camera 605 has a “field of view” 610. It will be appreciated that a camera's “field of view” is the amount of area that can be captured by the camera's lens and that is then included in a camera image. The left camera 605, in some instances, may be a wide-angle camera such that the field of view 610 is a wide-angle field of view”); obtaining a second image from a second image sensor communicatively coupled to the computing system (par 0009, “the left camera is used to capture a raw left image of the surrounding environment and the right camera is used to capture a raw right image of the surrounding environment”, par 0039, “a HMD is configured with a stereo camera pair which comprises a left and right camera. The stereo camera pair is used to capture images of a surrounding environment, wherein a center-line perspective of an image captured by the left camera has a non-parallel alignment with respect to a center-line perspective of an image captured by the right camera (i.e. the cameras are angled with respect to one another so that their fields of view are not parallel). The left camera captures a raw left image of the surrounding environment and the right camera captures a raw right image of the surrounding environment”, Fig 1, par 0046-0047, “these components operate to reconstruct a perspective captured by a camera image so that the captured perspective matches a perspective of the user …The computer system 100 may also be connected (via a wired or wireless connection) to external sensors 130 (e.g., one or more remote cameras, accelerometers, gyroscopes, acoustic sensors, magnetometers, etc.). Further, the computer system 100 may also be connected through one or more wired or wireless networks 135 to remote systems(s) 140 that are configured to perform any of the processing described with regard to computer system 100.”, par 0067, “ the left camera 605 and the right camera 615 are able to record the surrounding environment”), wherein the second image sensor has a second perspective of the environment (Fig 1, par 0046-0047, “Similar to the left camera 605, the right camera 615 also has a corresponding field of view 620. Furthermore, the right camera 615, in some instances, is a wide-angle camera. As also shown in FIG. 6, the field of view 610 and the field of view 620 include an overlapping region (i.e. overlapping FOV 630)”); generating a first display image for a first eye of a user of the computing system in the environment based at least in part on the first image, the second image, the first perspective of the environment, a first eye perspective of the environment from the first eye of the user, and the second perspective of the environment (Fig 13, par 0108, “By performing a reprojection operation 1300, the images captured by the left camera 1305A are altered. In particular, the images are altered in such a manner so that it appears as though the camera was actually located at a different position. This different position corresponds to the user's pupil locations 1315A and 1315B. To clarify, pupil location 1315A corresponds to a location of the user's left pupil while pupil location 1315B corresponds to a location of the user's right pupil. The distance 1320 corresponds to the user's interpupil distance (i.e. the distance between the pupil location 1315A and the pupil location 1315B). In essence, the images captured by the left camera 1305A are altered so that they appear as though they were actually captured by a camera (i.e. the “simulated,” or rather “reprojected,” left camera 1310) that was located near (i.e. a predetermined distance from, or rather in front of) pupil location 1315A”, par 0110-0114, “ during the initial image capture, the cameras were oriented at angles with respect to one another. As a consequence, the center-line perspectives were not in parallel with one another. In order to generate a user-friendly passthrough visualization, however, it is desired to alter the center-line perspectives so that the center-line perspectives match the user's perspective, which is determined by the user's pupil locations and distances. As clarified earlier, while the disclosure makes reference to “resulting location of the reprojected” cameras (or camera image), the physical location of the cameras is not being altered. Instead, various transformations are being imputed onto the image data to make it seem as though the image data was captured by a camera situated at a different location than where it actually is. Accordingly, the following equation is used to select the location of the resulting reprojected camera (or camera image). C.sub.syn(λ)=(1−λ)*C.sub.L+λ*C.sub.R Where C.sub.syn is the resulting location for the reprojected left camera 1310, C.sub.L is the actual position of the left camera 1305A, and C.sub.R is the actual position of the right camera 1305B. λ was defined earlier. To summarize, this equation is used to select the resulting location of the reprojected cameras“). But Zhu et al. keep silent for teaching generating a first display image for a first eye of a user of the computing system in the environment based at least in part on the first image, the second image, a first perspective difference between the first perspective of the environment and a first eye perspective of the environment from the first eye of the user, and a second perspective difference between the first perspective and the second perspective of the environment. PNG media_image1.png 340 652 media_image1.png Greyscale In related endeavor, Hosfield et al. teach generating a first display image for a first eye of a user of the computing system in the environment based at least in part on the first image, the second image, a first perspective difference between the first perspective of the environment and a first eye perspective of the environment from the first eye of the user, and a second perspective difference between the first perspective and the second perspective of the environment (par 0068, “The position and orientation of the camera 802 relative to the subject corresponds to the position and orientation of the camera relative to the subject for one of the source images. In FIG. 8B, vertices of the mesh are shown as having been re-projected in accordance with the location of a virtual camera (shown at position ‘x’) corresponding to the render camera location. In FIG. 8B, the re-projected portions of the mesh are shown as re-projected portions 804B. As can be seen in FIG. 8B, the vertices of these portions have been re-projected so as to be parallel to the view plane (and perpendicular to the view direction of the virtual camera). This re-projection may involve, for example, re-aligning at least some of the triangles in the polygonal mesh such that, in screen space, the x and y values of the triangles are preserved but the z values are set to be the closest value to the viewpoint of the three vertices, or alternatively to a median or mean of the three vertices. This process of re-projecting portions of the mesh to face the virtual camera may be repeated for each mesh, based on the pose of the camera for each source image and the pose of the virtual camera”, Figs 10A-10B, par 0076-0081, “n FIG. 10A, the virtual camera 1002 is shown at a given pose relative to the subject 1004. A line joining the virtual camera to the subject is indicated via dashed arrow 1006, which provides an indication of the distance of the subject in the z-direction. The poses of the source camera(s) relative to the subject are indicated at poses 1008. In FIGS. 10A and 10B, the source cameras are arranged in a co-planar arrangement …..The weighting associated with the distorted images may be determined by calculating an intersection between the line joining the virtual camera to the subject and a camera array plane. In FIGS. 10A and 10B, this point of intersection is marked as ‘x’ for the three rightmost cameras. The camera array plane may be defined as the plane joining at least three neighboring cameras”, Fig 12, par 0101-0105, “he image processor 1206 is configured to distort each source image based on the difference in pose of the source camera associated with that image and the pose of a virtual camera through which the object is viewable. The distortion of the source images may be performed in accordance with any of the methods described above in relation to FIGS. 3-11. … The system further comprises an image combiner 1208 configured to determine a weighting associated with each distorted image based on a similarity between the pose of the camera associated with that image and the pose of the viewer's head relative to the object (or a reconstruction thereof). The image combiner 1208 is further configured to combine the distorted images together in accordance with the associated weightings to form an image of the subject for display at the display element. That is, the image combiner 1208 may be configured to blend the distorted images together, to generate a final image of the object from the viewpoint of the viewer (which in turn, corresponds to the viewpoint of a virtual camera through which the object is made viewable). The final, blended image may correspond may be superimposed with a real or virtual environment that is viewable to the viewer” …. render left image for left eye). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Zhu et al. to include generating a first display image for a first eye of a user of the computing system in the environment based at least in part on the first image, the second image, a first perspective difference between the first perspective of the environment and a first eye perspective of the environment from the first eye of the user, and a second perspective difference between the first perspective and the second perspective of the environment as taught by Hosfield et al. to blended captured image together in accordance with the weightings to form an image of the subject from the viewpoint of the virtual camera based on a similarity between the pose of the corresponding source camera and the pose of the virtual camera to provide full immersion VR experience in a HMD device. Regarding claim 2, Zhu et al. as modified by Hosfield et al. teach all the limitation of claim 1, and Zhu et al. further teach wherein the computing system is head-mountable, the first image sensor is near the first eye of the user, and the second image sensor is near a second eye of the user (par 0008, “The disclosed HMDs include a stereo camera pair comprising a left camera and a right camera. The stereo camera pair is used to capture images of a surrounding environment”). Regarding claim 3, Zhu et al. as modified by Hosfield et al. teach all the limitation of claim 2, and further teach wherein the computing system further includes a first display for presenting the first display image to the first eye of the user and a second display for presenting a second display image to the second eye of the user (Zhu et al.: par 0140-0141, “the two passthrough visualizations are displayed using the head-mounted device (act 1840), such as can be performed with the reprojection component 425 and the display devices rendering the passthrough visualizations”, Hosfield et al.: par 0041, par 0087-0088, “The distorted images generated for the left images may be blended together to form the image for display at the left-eye display element of the HMD. Similarly, the distorted images generated for the right images may be blended together to form the image for display at the right-eye display element of the HMD”). Regarding claim 4, Zhu et al. as modified by Hosfield et al. teach all the limitation of claim 1, and further teach further comprising: generating a second display image for a second eye of the user based at least in part on the second image, the first image, a third perspective difference between the second perspective of the environment and a second eye perspective of the environment from the second eye of the user, and the second perspective difference between the first perspective and the second perspective of the environment (par 0068, “The position and orientation of the camera 802 relative to the subject corresponds to the position and orientation of the camera relative to the subject for one of the source images. In FIG. 8B, vertices of the mesh are shown as having been re-projected in accordance with the location of a virtual camera (shown at position ‘x’) corresponding to the render camera location. In FIG. 8B, the re-projected portions of the mesh are shown as re-projected portions 804B. As can be seen in FIG. 8B, the vertices of these portions have been re-projected so as to be parallel to the view plane (and perpendicular to the view direction of the virtual camera). This re-projection may involve, for example, re-aligning at least some of the triangles in the polygonal mesh such that, in screen space, the x and y values of the triangles are preserved but the z values are set to be the closest value to the viewpoint of the three vertices, or alternatively to a median or mean of the three vertices. This process of re-projecting portions of the mesh to face the virtual camera may be repeated for each mesh, based on the pose of the camera for each source image and the pose of the virtual camera”, Figs 10A-10B, par 0076-0081, “n FIG. 10A, the virtual camera 1002 is shown at a given pose relative to the subject 1004. A line joining the virtual camera to the subject is indicated via dashed arrow 1006, which provides an indication of the distance of the subject in the z-direction. The poses of the source camera(s) relative to the subject are indicated at poses 1008. In FIGS. 10A and 10B, the source cameras are arranged in a co-planar arrangement …..The weighting associated with the distorted images may be determined by calculating an intersection between the line joining the virtual camera to the subject and a camera array plane. In FIGS. 10A and 10B, this point of intersection is marked as ‘x’ for the three rightmost cameras. The camera array plane may be defined as the plane joining at least three neighboring cameras”, Fig 12, par 0101-0105, “he image processor 1206 is configured to distort each source image based on the difference in pose of the source camera associated with that image and the pose of a virtual camera through which the object is viewable. The distortion of the source images may be performed in accordance with any of the methods described above in relation to FIGS. 3-11. … The system further comprises an image combiner 1208 configured to determine a weighting associated with each distorted image based on a similarity between the pose of the camera associated with that image and the pose of the viewer's head relative to the object (or a reconstruction thereof). The image combiner 1208 is further configured to combine the distorted images together in accordance with the associated weightings to form an image of the subject for display at the display element. That is, the image combiner 1208 may be configured to blend the distorted images together, to generate a final image of the object from the viewpoint of the viewer (which in turn, corresponds to the viewpoint of a virtual camera through which the object is made viewable). The final, blended image may correspond may be superimposed with a real or virtual environment that is viewable to the viewer” …. render right image for right eye). Regarding claim 5, Zhu et al. as modified by Hosfield et al. teach all the limitation of claim 4, and further teach wherein generating the first display image and generating the second display image are performed in parallel (Zhu et al.: par 0140-0141, “the two passthrough visualizations are displayed using the head-mounted device (act 1840), such as can be performed with the reprojection component 425 and the display devices rendering the passthrough visualizations”, Hosfield et al.: par 0041, par 0087-0088, “The distorted images generated for the left images may be blended together to form the image for display at the left-eye display element of the HMD. Similarly, the distorted images generated for the right images may be blended together to form the image for display at the right-eye display element of the HMD” ….left and right images are display at same time after cameras captured left and right images). Regarding claim 9, Zhu et al. teach a computing system comprising: an interface for communicating with a first image sensor and a second image sensor; one or more processors; a non-transitory memory; and one or more programs stored in the non-transitory memory, which, when executed by the one or more processors, cause the computing system to (Fig 1, par 0009, par 0044). The remaining limitations of the claim are similar in scope to claim 1 and rejected under the same rationale. Regarding claims 10-13, Zhu et al. as modified by Hosfield et al. teach all the limitation of claim 9, the claims 10-13 are similar in scope to claims 2-5 and are rejected under the same rational. Regarding claim 17, Zhu et al. teach a non-transitory memory storing one or more programs, which, when executed by one or more processors of a computing system with an interface for communicating with a first image sensor and a second image sensor, cause the computing system to (Fig 1, par 0009, par 0044-0045). The remaining limitations of the claim are similar in scope to claim 1 and rejected under the same rationale. Regarding claims 18-20, Zhu et al. as modified by Hosfield et al. teach all the limitation of claim 17, the claims 18-20 are similar in scope to claims 2-4 and are rejected under the same rational. Claim(s) 6-7 and 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2019/0101758 to Zhu et al. in view of U.S. PGPubs 2020/0302688 to Hosfield et al., further in view of Lai et al. (CHUN-JUI LAI et al., "Exploring Manipulation Behavior on Video See-Through Head-Mounted Display with View Interpolation," Computer Vision-ACCV 2016 Workshops: ACCV 2016 International Workshops, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part III 13. Springer International Publishing, 2017). Regarding claim 6, Zhu et al. as modified by Hosfield et al. teach all the limitation of claim 1, and further teach further comprising: performing a warping operation on the first image to generate a first warped image to account for the perspective difference between the first perspective of the environment and the first eye perspective of the environment from the first eye of the user (Zhu et al.: Fig 13, par 0108, “By performing a reprojection operation 1300, the images captured by the left camera 1305A are altered. In particular, the images are altered in such a manner so that it appears as though the camera was actually located at a different position. This different position corresponds to the user's pupil locations 1315A and 1315B. To clarify, pupil location 1315A corresponds to a location of the user's left pupil while pupil location 1315B corresponds to a location of the user's right pupil. The distance 1320 corresponds to the user's interpupil distance (i.e. the distance between the pupil location 1315A and the pupil location 1315B). In essence, the images captured by the left camera 1305A are altered so that they appear as though they were actually captured by a camera (i.e. the “simulated,” or rather “reprojected,” left camera 1310) that was located near (i.e. a predetermined distance from, or rather in front of) pupil location 1315A”, par 0110-0114, “ during the initial image capture, the cameras were oriented at angles with respect to one another. As a consequence, the center-line perspectives were not in parallel with one another. In order to generate a user-friendly passthrough visualization, however, it is desired to alter the center-line perspectives so that the center-line perspectives match the user's perspective, which is determined by the user's pupil locations and distances. As clarified earlier, while the disclosure makes reference to “resulting location of the reprojected” cameras (or camera image), the physical location of the cameras is not being altered. Instead, various transformations are being imputed onto the image data to make it seem as though the image data was captured by a camera situated at a different location than where it actually is. Accordingly, the following equation is used to select the location of the resulting reprojected camera (or camera image). C.sub.syn(λ)=(1−λ)*C.sub.L+λ*C.sub.R Where C.sub.syn is the resulting location for the reprojected left camera 1310, C.sub.L is the actual position of the left camera 1305A, and C.sub.R is the actual position of the right camera 1305B. λ was defined earlier. To summarize, this equation is used to select the resulting location of the reprojected cameras“, Hosfield et al.: par 0068, “The position and orientation of the camera 802 relative to the subject corresponds to the position and orientation of the camera relative to the subject for one of the source images. In FIG. 8B, vertices of the mesh are shown as having been re-projected in accordance with the location of a virtual camera (shown at position ‘x’) corresponding to the render camera location. In FIG. 8B, the re-projected portions of the mesh are shown as re-projected portions 804B. As can be seen in FIG. 8B, the vertices of these portions have been re-projected so as to be parallel to the view plane (and perpendicular to the view direction of the virtual camera). This re-projection may involve, for example, re-aligning at least some of the triangles in the polygonal mesh such that, in screen space, the x and y values of the triangles are preserved but the z values are set to be the closest value to the viewpoint of the three vertices, or alternatively to a median or mean of the three vertices. This process of re-projecting portions of the mesh to face the virtual camera may be repeated for each mesh, based on the pose of the camera for each source image and the pose of the virtual camera”, Figs 10A-10B, par 0076-0081, “n FIG. 10A, the virtual camera 1002 is shown at a given pose relative to the subject 1004. A line joining the virtual camera to the subject is indicated via dashed arrow 1006, which provides an indication of the distance of the subject in the z-direction. The poses of the source camera(s) relative to the subject are indicated at poses 1008. In FIGS. 10A and 10B, the source cameras are arranged in a co-planar arrangement …..The weighting associated with the distorted images may be determined by calculating an intersection between the line joining the virtual camera to the subject and a camera array plane. In FIGS. 10A and 10B, this point of intersection is marked as ‘x’ for the three rightmost cameras. The camera array plane may be defined as the plane joining at least three neighboring cameras”, Fig 12, par 0101-0105, “he image processor 1206 is configured to distort each source image based on the difference in pose of the source camera associated with that image and the pose of a virtual camera through which the object is viewable. The distortion of the source images may be performed in accordance with any of the methods described above in relation to FIGS. 3-11. … The system further comprises an image combiner 1208 configured to determine a weighting associated with each distorted image based on a similarity between the pose of the camera associated with that image and the pose of the viewer's head relative to the object (or a reconstruction thereof). The image combiner 1208 is further configured to combine the distorted images together in accordance with the associated weightings to form an image of the subject for display at the display element. That is, the image combiner 1208 may be configured to blend the distorted images together, to generate a final image of the object from the viewpoint of the viewer (which in turn, corresponds to the viewpoint of a virtual camera through which the object is made viewable). The final, blended image may correspond may be superimposed with a real or virtual environment that is viewable to the viewer”), but keep silent for teaching generating an occlusion mask based on the first warped image indicating a plurality of holes in the first warped image. In related endeavor, Lai et al. teach further comprising: performing a warping operation on the first image to generate a first warped image to account for the perspective difference between the first perspective of the environment and the first eye perspective of the environment from the first eye of the user (Fig 2, section 3.1, “we can use camera parameter to warp the triangles from cameras to virtual views for left and right eyes according to follow formula”); and generating an occlusion mask based on the first warped image indicating a plurality of holes in the first warped image (section 3.2, “If the depth of any vertices on the triangle is 1.05 times greater or 0.95 times smaller than any other two vertices, we consider the triangle as an occlusion area because the difference of depth is too large.”, Figs 5-6, section 3.3 and 3.4, “because the cameras are in front of the user, we can ensure that the occlusion area occurred when background is occluded by foreground. Therefore, for each pixel in occlusion area, we consider the horizontal line and vertical line cross the pixel and boundary of occlusion area.”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Zhu et al. as modified by Hosfield et al. to include generating an occlusion mask based on the first warped image indicating a plurality of holes in the first warped image as taught by Lai et al. to use depth image-based rendering algorithm to re-compute the true distance of the scene to remove the error due to the distance between cameras and users to reduce the occlusion areas and render the correct image to the user. Regarding claim 7, Zhu et al. as modified by Hosfield et al. and Lai et al. teach all the limitation of claim 6, and Lai et al. further teach wherein the occlusion mask is determined based at least in part on a first set of depth values relative to the first perspective associated with the first image sensor and a second set of depth values relative to the second perspective associated with the first eye (Fig 2, section 3.1, “we can use camera parameter to warp the triangles from cameras to virtual views for left and right eyes according to follow formula”); and generating an occlusion mask based on the first warped image indicating a plurality of holes in the first warped image (section 3.2, “If the depth of any vertices on the triangle is 1.05 times greater or 0.95 times smaller than any other two vertices, we consider the triangle as an occlusion area because the difference of depth is too large.”, Figs 5-6, section 3.3 and 3.4, “because the cameras are in front of the user, we can ensure that the occlusion area occurred when background is occluded by foreground. Therefore, for each pixel in occlusion area, we consider the horizontal line and vertical line cross the pixel and boundary of occlusion area.”, also show in Fig 11). Regarding claims 14-15, Zhu et al. as modified by Hosfield et al. teach all the limitation of claim 9, the claims 14-15 are similar in scope to claims 6-7 and are rejected under the same rational. Claim(s) 8 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over U.S. PGPubs 2019/0101758 to Zhu et al. in view of U.S. PGPubs 2020/0302688 to Hosfield et al., further in view of Lai et al. (CHUN-JUI LAI et al., "Exploring Manipulation Behavior on Video See-Through Head-Mounted Display with View Interpolation," Computer Vision-ACCV 2016 Workshops: ACCV 2016 International Workshops, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part III 13. Springer International Publishing, 2017), further in view of U.S. PGPubs 2012/0120192 to Alregib et al.. Regarding claim 8, Zhu et al. as modified by Hosfield et al. and Lai et al. teach all the limitation of claim 6, but keep silent for teaching further comprising: filling a first set of the plurality of holes of the occlusion mask based on the second image; and generating a diffused first image by performing a pixelwise diffusion process to fill a second set of the plurality of holes of the occlusion mask, wherein generating the first display image for the first eye of the user includes compositing the diffused first image with rendered extended reality (XR) content based at least in part on depth information relative to at least one of the first image sensor and the first eye of the user. In related endeavor, Alregib et al. teach further comprising: filling a first set of the plurality of holes of the occlusion mask based on the second image; and generating a diffused first image by performing a pixelwise diffusion process to fill a second set of the plurality of holes of the occlusion mask, wherein generating the first display image for the first eye of the user includes compositing the diffused first image with rendered extended reality (XR) content based at least in part on depth information relative to at least one of the first image sensor and the first eye of the user (par 0017-0020, “Embodiments of the present invention comprise at least two new approaches for error and disocclusion removal in depth image-based rendering ("DIBR"). These approaches can include hierarchical hole-filling ("HHF") and depth adaptive hierarchical hole-filling ("depth adaptive HHF") and can eliminate the need for additional smoothing or filtering of a depth map …. In some embodiments of the HHF approach, lower-resolution estimates of a 3D wrapped image can be produced by a pseudo Gaussian plus zero-canceling filter. This operation can be referred to as a Reduce operation and, in some embodiments, can be repeated until there are no longer holes remaining in the image. In some embodiments, the lowest resolution image can then be expanded in an Expand operation. The expanded pixels can then be averaged and used to fill the holes in the second most reduced image. The second most reduce image can then be expanded …. After the 3D wrapped image is preprocessed to increase the weight of the background, the HHF process can be applied to the preprocessed image and the pixels from the resulting image are then used to fill holes in the 3D wrapped image”, par 0076, “the smoothing function can be a pseudo Gaussian plus zero-canceling filter. In other embodiments, the holes can be non-zero values that are, for example and not limitation, saturated at the highest luminance value or other forms of non-zero holes. In such cases, the smoothing function can be a pseudo Gaussian plus K-canceling filter, where K is the value assigned for the holes in an image. The smoothing filter can use, for example, non-zero values in an [X.times.X] block of pixels”). It would have been obvious to a person of ordinary skill in the art at the time before the effective filing data of the claimed invention to modified Zhu et al. as modified by Hosfield et al. and Lai et al. to include further comprising: filling a first set of the plurality of holes of the occlusion mask based on the second image; and generating a diffused first image by performing a pixelwise diffusion process to fill a second set of the plurality of holes of the occlusion mask, wherein generating the first display image for the first eye of the user includes compositing the diffused first image with rendered extended reality (XR) content based at least in part on depth information relative to at least one of the first image sensor and the first eye of the user as taught by Alregib et al. to provide depth adaptive hierarchical hole-filling method to preprocessing the depth map of a 3D wrapped image that contains holes, reducing the preprocessed image, expanding the reduced image, and filling the holes in the 3D wrapped image with data obtained from the expanded image to efficiently reduce errors in images and produce 3D images from a 2D images and/or depth map information. Regarding claim 16, Zhu et al. as modified by Hosfield et al. and Lai et al. teach all the limitation of claim 14, the claim 16 is similar in scope to claim 8 and is rejected under the same rational. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jin Ge whose telephone number is (571)272-5556. The examiner can normally be reached 8:00 to 5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at (571)272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. JIN . GE Examiner Art Unit 2619 /JIN GE/Primary Examiner, Art Unit 2619
Read full office action

Prosecution Timeline

Mar 11, 2025
Application Filed
Sep 15, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749262
THREE-DIMENSIONAL MODEL GENERATION METHOD, THREE-DIMENSIONAL MODEL GENERATION DEVICE, AND NON-TRANSITORY COMPUTER READABLE MEDIUM
3y 4m to grant Granted Sep 29, 2026
Patent 12743815
ON COMPRESSION OF A MESH WITH MULTIPLE TEXTURE MAPS
2y 1m to grant Granted Sep 22, 2026
Patent 12731176
Method for creating digital art from photos and videos of coins and various methods of presenting the art to be viewed.
3y 8m to grant Granted Sep 08, 2026
Patent 12731352
Video System with Scene-Based Object Insertion Feature
3y 0m to grant Granted Sep 08, 2026
Patent 12718500
DELIVERING VIRTUALIZED CONTENT
2y 4m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
98%
With Interview (+18.8%)
2y 6m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 552 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month