Got DETAILED ACTION
Response to Amendment
Applicant’s amendments filed on 17 March 2026 have been entered. Claims 1, 19 and 20 have been amended. Claims 1-20 are still pending in this application, with claims 1, 19 and 20 being independent.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4, 6, 7, 19 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hayasaka et al. (US 20230107179 A1), referred herein as Hayasaka in view of Philip et al. (US 20230237718 A1), referred herein as Philip and Kar et al. (US 20220335638 A1), referred herein as Kar.
Regarding Claims 1, 19 and 20, Hayasaka in view of Philip and Kar teaches a method, device and non-transitory memory storing one or more programs comprising (Hayasaka Abst: an information processing apparatus and method as well as a program that allow editing of a depth map; [0062] The application program executed by the computer 100 can be recorded, for example, to the removable recording medium 121 as a package medium or the like for application):
at a device including an image sensor, a display, one or more processors, and a non- transitory memory (Hayasaka [0029] a captured image 10 illustrated in A of FIG. 1 is an image captured of a person 11, an object 12, and the like; [0062] The application program executed by the computer 100 can be recorded, for example, to the removable recording medium 121 as a package medium or the like for application; FIG. 2):
Hayasaka in view of Philip teaches
capturing, using the image sensor, an image of a physical environment (Hayasaka [0069] the file acquisition section 152 acquires an image (e.g., captured image); Philip [0050] The digital image 106 is configurable to capture physical environments);
obtaining a first depth map including a plurality of depths respectively associated with a plurality of pixels of the image of the physical environment (Hayasaka [0030] A depth map 20 illustrated in B of FIG. 1 is a depth map generated from the captured image 10 or the like and basically includes depth values corresponding to the captured image 10. In the depth map 20, the smaller the depth value of each pixel (that is, the more forward (the more on the camera side) the pixel is located), the closer to white the pixel is represented as being, and the larger the depth value (that is, the more backward the pixel is located), the closer to black the pixel is represented as being; [0034] Here, the basic depth map is a depth map that includes depth values of a subject of a certain image);
generating a second depth map (Hayasaka [0033] an assist depth map is generated that includes depth data to be added to (merged with) a basic depth map.The use of the UI in this manner allows an instruction associated with generation of the assist depth map to be input with ease) by aligning one or more portions of the first depth map based on fi)_a control signal associated with the image of the physical environment and (ii) the image of the physical environment (Kar [0033] The depth estimation system 100 is configured to convert the depth map 120 having the first scale to the depth map 138 having the second scale. The depth maps 138 with the second scale may be used to control augmented reality, robotics, natural user interface technology, gaming, or other applications; [0049] The depth map transformer 126 is configured to execute a parameter estimation algorithm to solve an optimization problem (e.g., an objective function) which minimizes an objective of aligning the depth estimates 108 with the depth map 120. In other words, the depth map transformer 126 is configured to minimize an objective function of aligning the depth estimates 108 with the depth map 120 to estimate the affine parameters 132; [0072] Operation 508 includes transforming the depth map 120 to a depth map 138 (e.g., a second depth map) using the depth estimates 108, where the depth map 138 has a second scale);
transforming, using the one or more processors, the image of the physical environment based on the second depth map (Hayasaka [0078] the display image generation section 158 can further superimpose an assist depth map or an overwrite depth map on the superimposed image. That is, the display image generation section 158 can function as a superimposition processing section; Philip [0036] The image processing system 104 is implemented at least partially in hardware of the computing device 102 to process and transform a digital image 106); and
displaying, via the display, the transformed image (Hayasaka [0148] a superimposition processing section adapted to superimpose the basic depth map and an image corresponding to the basic depth map one on the other and output the resultant image to the display section as a superimposed image; Philip [0057] The edited digital image 234 is then output (block 412), e.g., to be saved in the storage device 108, rendered in the user interface 110, and so forth).
Philip discloses a directional propagation editing techniques, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Hayasaka to incorporate the teachings of Philip, and apply the edited image of a real-world environment using a depth map to the depth map editing interface.
Doing so would be able to improve image editing accuracy in both virtual and real-world environments.
Kar discloses a method for depth estimation includes receiving image data from a sensor system, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Hayasaka to incorporate the teachings of Kar, and apply the transforming of the first depth map to a second depth map using the depth estimates to the depth map editing interface.
Doing so would provide an efficient method for estimating a depth map from a single image well suited for mobile applications.
Regarding Claim 2, Hayasaka in view of Philip and Kar teaches the method of claim 1, and further teaches further comprising: obtaining the control signal associated with texture alignment, wherein generating the second depth map includes aligning the one or more portions of the first depth map based on at least one texture or texture transition within the image of the physical environment (Hayasaka [0051] In the case of such a depth map, there may be an invalid value in portions with a depth value having low reliability (e.g., portions with little texture in the frame). The setting of a depth value to such portions with use of an assist depth map allows a more reliable depth map to be generated).
Regarding Claim 3, Hayasaka in view of Philip and Kar teaches the method of claim 1, and further teaches further comprising: obtaining the control signal associated with color alignment, wherein generating the second depth map includes aligning the one or more portions of the first depth map based on at least one color or color transition within the image of the physical environment (Hayasaka [0040] Also, a basic depth map and an image (RGB image which will be described later) corresponding to the basic depth map may be displayed in a superimposed manner. In the information processing apparatus, for example, a superimposition processing section may superimpose a basic depth map and an image corresponding to the basic depth map one on the other and output the resultant image to the display section as a superimposed image. In the case of a depth map, the regions having the same depth value are represented as being the same color).
Regarding Claim 4, Hayasaka in view of Philip and Kar teaches the method of claim 1, and further teaches further comprising: obtaining the control signal associated with luminous intensity alignment, wherein generating the second depth map includes aligning the one or more portions of the first depth map based on at least one luminous intensity or luminous intensity transition within the image of the physical environment (Hayasaka [0087] the depth map display section 201 can display an RGB image and a depth map in a superimposed manner (in a translucent manner). Here, the RGB image is, for example, an image including luminance components, color components, and the like which is not but corresponds to a depth map).
Regarding Claim 6, Hayasaka in view of Philip and Kar teaches the method of claim 1, and further teaches wherein the control signal indicates one or more particular objects within the image of the physical environment, and wherein the one or more portions of the first depth map correspond to the one or more particular objects (Hayasaka [0029] For example, a captured image 10 illustrated in A of FIG. 1 is an image captured of a person 11, an object 12, and the like. The person 11 is located forward (camera side). White represents a background and is the farthest from the camera (e.g., infinity). Portions indicated by diagonal lines, parallel lines, a hatched pattern, and the like are located between the person 11 and the background as seen from the camera; [0096] Next, when the user specifies a pixel in the region 332 by operating the cursor 341 in the edge depth map 330 as illustrated in C of FIG. 7, the depth value (i.e., 20) of the pixel is set in (copied to) the shape region 342 of the assist depth map 340 as illustrated in D of FIG. 7).
Regarding Claim 7, Hayasaka in view of Philip and Kar teaches the method of claim 1, and further teaches wherein the control signal indicates one or more edges within the image of the physical environment, and wherein the one or more portions of the first depth map correspond to the one or more edges (Hayasaka [0052] An in-frame position whose depth value is relatively highly reliable is an edge portion of a subject in the image corresponding to the basic depth map, and the basic depth map may be an edge depth map that includes the depth value of the edge portion).
Claim(s) 5 and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hayasaka et al. (US 20230107179 A1), referred herein as Hayasaka in view of Philip et al. (US 20230237718 A1), referred herein as Philip, Kar et al. (US 20220335638 A1), referred herein as Kar and Price et al. (US 20210358155 A1), referred herein as Price.
Regarding Claim 5, Hayasaka in view of Philip and Kar teaches the method of claim 1, but does not teach the limitations here.
However, Price teaches wherein the one or more portions of the first depth map are aligned using a joint bilateral filter (Price [0117] Thus, FIG. 5D depicts a filtering operation 555A being performed on upsampled left depth map 535B. In some instances, the filtering operation 555A is an edge-preserving filtering operation that uses a guidance image, such as, by way of nonlimiting example, a joint bilateral filter, a guided filter, a bilateral solver, or any other suitable filtering technique).
Price discloses systems and methods for temporally consistent depth map generation, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Hayasaka to incorporate the teachings of Price, and apply the joint bilateral filter to the depth map editing technique.
Doing so would be able to implement reconstruction filtering to prevent artifacts.
Regarding Claim 15, Hayasaka in view of Philip and Kar teaches the method of claim 1, but does not teach all the limitations here.
However, in viewing of Price, the prior art teaches wherein the first depth map corresponds to a temporally stable depth map including a plurality of depths respectively associated with a plurality of pixels of the image of the physical environment (Price [0058] As used herein, a “depth map” details the positional relationship and depths relative to objects in the environment. Consequently, the positional arrangement, location, geometries, contours, and depths of objects relative to one another can be determined. From the depth maps (and possibly the raw images), a 3D representation of the environment can be generated; [0064] Based on these pixel disparities, the embodiments are able to determine depths for objects located within the overlapping region (i.e. “stereoscopic depth matching,” “stereo depth matching,” or simply “stereo matching”). As such, the visible light camera(s) 210 can be used to not only generate passthrough visualizations, but they can also be used to determine object depth; [0117] the 3D geometric transforms rely on depth computations in which the objects in the HMD 300's environment are mapped out to determine their depths as well as the pose 375; [0162] Pursuant to providing temporally consistent consecutively generated depth maps, FIG. 7C illustrates generating a reprojected P1 depth map 715B by performing a reprojection operation 765 on the P1 depth map 715A), wherein the temporally stable depth map excludes depths based on distances between the image sensor and dynamic objects (Hayasaka [0088] The time line 212 is a GUI that horizontally indicates a video sequence. That is, the RGB image, the edge depth map, the assist depth map, and the overwrite depth map may be a single-frame still image or a video that includes a plurality of frames).
Price discloses systems and methods for temporally consistent depth map generation, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Hayasaka to incorporate the teachings of Price, and apply the temporally consistent depth map generation to the depth map editing technique.
Doing so would be able to avoid discrepancies, or temporal inconsistencies that can give rise to artifacts.
Claim(s) 8-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hayasaka et al. (US 20230107179 A1), referred herein as Hayasaka in view of Philip et al. (US 20230237718 A1), referred herein as Philip, Kar et al. (US 20220335638 A1), referred herein as Kar and TOMARU et al. (US 20190102948 A1), referred herein as TOMARU.
Regarding Claim 8, Hayasaka in view of Philip and Kar teaches the method of claim 1, but does not teach the limitations here.
However, TOMARU teaches wherein the first depth map includes, for a particular pixel at a particular pixel location representing a dynamic object in the physical environment, a particular depth corresponding to a distance between the image sensor and a static object in the physical environment behind the dynamic object (TOMARU [0077] The object information acquisition unit 23 reads and acquires the navigation data 41 stored in the storage 13, which is information on the object existing around the moving body 100; [0069] In Embodiment 1, the depth map generation unit 21 generates the depth map by a stereo method. Specifically, the depth map generation unit 21 finds a pixel capturing the same object in images captured by the two cameras, and determines a distance of the pixel found by triangulation. The depth map generation unit 21 generates a depth map by calculating distances for all the pixels. The depth map generated from the image illustrated in FIG. 4 is as illustrated in FIG. 5, and each pixel indicates the distance from the camera to the subject. In FIG. 5, a value is smaller as it is closer to the camera, and is larger as it is farther from the camera, so that the closer side is shown by denser hatching, and the farther side is shown by thinner hatching).
TOMARU discloses a technique for displaying an object around a moving body by superimposing the object on a scenery around the moving body, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Hayasaka to incorporate the teachings of TOMARU, and apply the depth map generation unit 21 generates the depth map by a stereo method for a moving body to the depth map editing technique.
Doing so would make it easy to see the necessary information while maintaining the sense of reality by switching presence or absence of the shielding according to an importance of the object.
Regarding Claim 9, Hayasaka in view of Philip, Kar and TOMARU teaches the method of claim 8, and further teaches wherein obtaining the first depth map includes determining the particular depth via interpolation using depths of locations surrounding the particular pixel location (Philip [0056] To do so, the feature system 22 employs a feature generation module 226 to generate the feature volume 224 by sampling color features 228 from the digital image 106 and corresponding depth features 230 from the depth map 208 along the direction 218 (block 408). The color features 228 describe color of the respective pixels and thus when collected for each of the pixels provides information regarding a relationship of a pixel to local neighbors of the pixel. The depth feature 230 also describes the depth of those pixels).
Regarding Claim 10, Hayasaka in view of Philip, Kar and TOMARU teaches the method of claim 8, and further teaches wherein obtaining the first depth map includes determining the particular depth at a time the dynamic object was not represented at the particular pixel location (Hayasaka [0045] a basic depth map may correspond to a still image or to a video. In a case where the basic depth map corresponds to a video, the basic depth map may include a plurality of frames in a chronological order (i.e., group of basic depth maps corresponding to respective frame images)).
Claim(s) 11, 12 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hayasaka et al. (US 20230107179 A1), referred herein as Hayasaka in view of Philip et al. (US 20230237718 A1), referred herein as Philip, Kar et al. (US 20220335638 A1), referred herein as Kar, TOMARU et al. (US 20190102948 A1), referred herein as TOMARU and ADKINSON et al. (US 20210295599 A1), referred herein as ADKINSON.
Regarding Claim 11, Hayasaka in view of Philip, Kar and TOMARU teaches the method of claim 8, but does not teach the limitations here.
However, ADKINSON teaches wherein obtaining the first depth map includes determining the particular depth based on a three-dimensional model of the physical environment (ADKINSON [0008] FIG. 4 is a flowchart of the operations of an example method for recreating depth information and recovering scale for a reconstructed 3D model; [0058] The depth estimation network may be trained on sets of various images with corresponding depth maps that provide actual (real-world) metric scale on objects within the various images).
ADKINSON discloses a technique of reconstruction of a 3D model (or “digital twin”) and associated depth and camera data, and scale estimation from the reconstructed model and data, from a remote video feed, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Hayasaka to incorporate the teachings of ADKINSON, and apply the depth and camera data from reconstructed of 3D model to the depth map editing technique.
Doing so would be capable of supporting a remote video session with which users can interact via AR objects in real-time.
Regarding Claim 12, Hayasaka in view of Philip, Kar, TOMARU and ADKINSON teaches the method of claim 11, and further teaches wherein determining the particular depth based on a three-dimensional model includes at least one of: rasterizing the three-dimensional model; or ray tracing based on the three-dimensional model (Philip [0061] FIG. 6 depicts an example implementation 600 of a ray-marching technique used by the feature generation module 226 to generate the features 502 as including the color 506 and depth 510).
Regarding Claim 14, Hayasaka in view of Philip, Kar, TOMARU and ADKINSON teaches the method of claim 11, and further teaches wherein the three-dimensional model of the physical environment excludes dynamic objects (Hayasaka [0088] The time line 212 is a GUI that horizontally indicates a video sequence. That is, the RGB image, the edge depth map, the assist depth map, and the overwrite depth map may be a single-frame still image or a video that includes a plurality of frames).
Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hayasaka et al. (US 20230107179 A1), referred herein as Hayasaka in view of Philip et al. (US 20230237718 A1), referred herein as Philip, Kar et al. (US 20220335638 A1), referred herein as Kar, TOMARU et al. (US 20190102948 A1), referred herein as TOMARU, ADKINSON et al. (US 20210295599 A1), referred herein as ADKINSON and Price et al. (US 20210358155 A1), referred herein as Price.
Regarding Claim 13, Hayasaka in view of Philip, Kar, TOMARU and ADKINSON teaches the method of claim 11, but does not teach the limitations here.
However Price teaches wherein the three-dimensional model of the physical environment corresponds to a temporally stable three-dimensional model of the physical environment (Price [0117] the 3D geometric transforms rely on depth computations in which the objects in the HMD 300's environment are mapped out to determine their depths as well as the pose 375; [0162] Pursuant to providing temporally consistent consecutively generated depth maps, FIG. 7C illustrates generating a reprojected P1 depth map 715B by performing a reprojection operation 765 on the P1 depth map 715A).
Price discloses systems and methods for temporally consistent depth map generation, which is analogous to the present patent application.
It would have been obvious for a person of ordinary skill in the art before the effective filing date of the claimed invention to have modified Hayasaka to incorporate the teachings of Price, and apply the temporally consistent depth map generation to the depth map editing technique.
Doing so would be able to avoid discrepancies, or temporal inconsistencies that can give rise to artifacts.
Claim(s) 16-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hayasaka et al. (US 20230107179 A1), referred herein as Hayasaka in view of Philip et al. (US 20230237718 A1), referred herein as Philip, Kar et al. (US 20220335638 A1), referred herein as Kar and ADKINSON et al. (US 20210295599 A1), referred herein as ADKINSON.
Regarding Claim 16, Hayasaka in view of Philip, Kar and TOMARU teaches the method of claim 1, but does not teach all the limitations here.
However, in viewing of Price, the prior art teaches wherein obtaining the first depth map includes obtaining a three-dimensional model of the physical environment and generating the first depth map (ADKINSON [0008] FIG. 4 is a flowchart of the operations of an example method for recreating depth information and recovering scale for a reconstructed 3D model; [0058] The depth estimation network may be trained on sets of various images with corresponding depth maps that provide actual (real-world) metric scale on objects within the various images), based on the three-dimensional model, the first depth map including a plurality of depths respectively associated with a plurality of pixels of the image of the physical environment (Philip [0026] A depth map is also obtained. Continuing with the first scenario above, the depth map is also captured that measures depths at corresponding points (e.g., pixels) within a physical environment).
Regarding Claim 17, Hayasaka in view of Philip, Kar and ADKINSON teaches the method of claim 16, and further teaches wherein the three-dimensional model is based on objects in the physical environment determined to be static (Hayasaka [0088] The time line 212 is a GUI that horizontally indicates a video sequence. That is, the RGB image, the edge depth map, the assist depth map, and the overwrite depth map may be a single-frame still image or a video that includes a plurality of frames. In the time line 212, for example, the frame to be processed is selected by specifying the position of a pointer 213 (horizontal position)).
Regarding Claim 18, Hayasaka in view of Philip, Kar and ADKINSON teaches the method of claim 16, and further teaches further comprising:
generating the three-dimensional model, at least in part by:
determining that one or more points correspond to one or more static objects in the physical environment (ADKINSON [0039] one or more anchor points may be identified from the video feed. These anchor points serve as locations within the environment around the capturing device that can be repeatably and consistently identified when the point moves out of and back into frame. These anchor points can be identified, tagged, or otherwise associated with corresponding objects within the 3D model/digital twin); and
adding the one or more points to the three-dimensional model at one or more locations in a three-dimensional coordinate system of the physical environment corresponding to the one or more static objects in the physical environment (ADKINSON [0050] As the video stream or sequence of frames is processed, additional identified points form additional depth maps, which are added to the sparse reconstruction as more consecutive or temporally proximate frames are registered, until all frames of the video stream or sequence of frames intended to be used for the reconstruction have been processed).
Response to Arguments
Applicant’s arguments with respect to claim(s) 1, 19 and 20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Samantha (Yuehan) Wang whose telephone number is (571)270-5011. The examiner can normally be reached Monday-Friday, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at (571)272-7440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Samantha (YUEHAN) WANG/
Primary Examiner
Art Unit 2617