DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1-2, 9 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Hazeghi U.S. Patent Application 20180130255 in view of Lee U.S. Patent Application 20190379842, and further in view of Wen U.S. Patent Application 20260148405.
Regarding claim 11, Hazeghi discloses a stereo image display system, comprising:
a stereo display device (display 62); and
at least one processor (CPU/GPU 42), coupled to the stereo display device (paragraph [0020]: the memory of the client-side device stores instructions that, when executed by the processor of the client-side device, cause the processing system to: control the acquisition system to capture images of an object; paragraph [0065]: Stereoscopic algorithms exploit this property of the disparity. These algorithms achieve 3D reconstruction by matching points (or features) detected in the left and right views, which is equivalent to estimating disparities), and configured to:
obtain images; perform depth estimation on the images respectively to obtain depth maps; generate depth information (paragraph [0094]: In operation 530, the scanning system computes depth images or depth maps from the images captured by the acquisition system 20. The depth images may capture different views of the object (e.g., the acquisition system 20 may have different poses, with respect to the object, as it captures images of the object, and the images captured at different poses can be used to generate depth maps of the object from the viewpoints of the different poses); paragraph [0065]: Stereoscopic algorithms exploit this property of the disparity. These algorithms achieve 3D reconstruction by matching points (or features) detected in the left and right views, which is equivalent to estimating disparities);
generate 3D mesh of the images based on the depth maps, generate a 3D scene mesh according to the stereo depth information, the left side 3D mesh and the right side 3D mesh (paragraph [0095]: In operation 550, the depth maps are combined to generate a 3D point cloud (e.g., each of the depth maps may be a point cloud that can be merged with other point clouds); paragraph [0096]: In operation 570, a 3D mesh model is generated from the combined 3D point cloud; paragraph [0065]: Stereoscopic algorithms exploit this property of the disparity. These algorithms achieve 3D reconstruction by matching points (or features) detected in the left and right views, which is equivalent to estimating disparities); and
output 3D scene content according to the 3D scene mesh (paragraph [0111]: The mesh and texture maps are sent back to the client (e.g., the interaction system 60) for display of the resulting model (e.g., on the display device 62 of the interaction system 60); paragraph [0020]: display, on the display device, a 3D mesh model generated from the images).
Hazeghi discloses all the features with respect to claim 11 as outlined above. However, Hazeghi fails to disclose obtaining a left eye image and a right eye image from a stereo image pair; performing monocular depth estimation, generating a stereo depth information according to disparity information between the left eye image and the right eye image.
Lee discloses obtaining a left eye image and a right eye image from a stereo image pair; generating a stereo depth information according to disparity information between the left eye image and the right eye image (paragraph [0028]: after the depth map generator 405 receives the left-eye temporary image and the right-eye temporary image, because there is a disparity between the left-eye temporary image and the right-eye temporary image, the depth map generator 405 can generate the depth map (as shown in FIG. 7) according to the left-eye temporary image, the right-eye temporary image and the disparity. It is obvious to those of ordinary skill in the art that the depth map generator 405 generates the depth map according to the left-eye temporary image, the right-eye temporary image and the disparity; Lee’s teaching of left and right eye images can be combined with Hazeghi’s device, such that to generate left and right side 3D mesh based on the left and right eye depth maps).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi’s to use left and right eye images as taught by Lee, to process images efficiently.
Hazeghi as modified by Lee discloses all the features with respect to claim 11 as outlined above. However, Hazeghi as modified by Lee fails to disclose performing monocular depth estimation.
Wen discloses performing monocular depth estimation to obtain depth map (paragraph [0054]: the second STA 108 utilizes a monocular depth estimation model 154 and a side-tuning CNN 156 to generate a left feature map; paragraph [0036]: the STA may use the monocular depth estimation model 154… given a pair of left and right images Il, Ir∈RH×W×3, the side-tuning CNN 156 (e.g., an EdgeNeXt-S model) may be used within the STA (e.g., the STAs 106 and 108) to extract multi-level pyramid features, where the ¼ level feature is equipped with the monocular depth estimation model's 154 features).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi and Lee’s to perform monocular depth estimation as taught by Wen, to perform stereo matching efficiently.
Claim 1 recites the functions of the apparatus recited in claim 11 as method steps. Accordingly, the mapping of the prior art to the corresponding functions of the apparatus in claim 11 applies to the method steps of claim 1.
Regarding claim 2, Hazeghi as modified by Lee and Wen discloses the 3D scene visualization method as claimed in claim 1, wherein the step of performing the monocular depth estimation on the left eye image and the right eye image respectively to obtain the left eye depth map and the right eye depth map comprises:
inputting the left eye image into a monocular depth estimation model to obtain the left eye depth map; and inputting the right eye image into the monocular depth estimation model to obtain the right eye depth map (Wen’s paragraph [0036]: the STA may use the monocular depth estimation model 154… given a pair of left and right images Il, Ir∈RH×W×3, the side-tuning CNN 156 (e.g., an EdgeNeXt-S model) may be used within the STA (e.g., the STAs 106 and 108) to extract multi-level pyramid features, where the ¼ level feature is equipped with the monocular depth estimation model's 154 features; paragraph [0054]: the second STA 108 utilizes a monocular depth estimation model 154 and a side-tuning CNN 156 to generate a left feature map).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi’s to use left and right eye images as taught by Lee, to process images efficiently; and combine Hazeghi and Lee’s to perform monocular depth estimation as taught by Wen, to perform stereo matching efficiently.
Regarding claim 9, Hazeghi as modified by Lee and Wen discloses the 3D scene visualization method as claimed in claim 1, wherein the step of outputting 3D scene content according to the 3D scene mesh comprises:
generating a two-dimensional rendered screen of a single viewpoint according to the 3D scene mesh; and displaying the two-dimensional rendered screen using a display device (Hazeghi’s paragraph [0158]: FIGS. 8A, 8B, and 8C are screenshots depicting user interfaces displayed on the display device 62... the user interface shown on the display device 62 of the interaction system is configured to show a live view of the current view from the color camera 22. The user interface may also show the current point cloud (3D scene mesh) if the user would like to resume. The point cloud is constituted by a set of three-dimensional points which can be visualized in the two-dimensional frame of the user interface by applying rendering techniques which account for occlusions, normal, and the orientation of the current image acquired by the color camera with respect to the point cloud itself).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi’s to use left and right eye images as taught by Lee, to process images efficiently; and combine Hazeghi and Lee’s to perform monocular depth estimation as taught by Wen, to perform stereo matching efficiently.
Claim 3-5 are rejected under 35 U.S.C. 103 as being unpatentable over Hazeghi U.S. Patent Application 20180130255 in view of Lee U.S. Patent Application 20190379842, in view of Wen U.S. Patent Application 20260148405, and further in view of Skrypnyk U.S. Patent Application 20240355064.
Regarding claim 3, Hazeghi as modified by Lee and Wen discloses generating the left side 3D mesh of the left eye image based on the left eye depth map, and generating the right side 3D mesh of the right eye image based on the right eye depth map (Hazeghi's paragraph [0095]: In operation 550, the depth maps are combined to generate a 3D point cloud (e.g., each of the depth maps may be a point cloud that can be merged with other point clouds); paragraph [0096]: In operation 570, a 3D mesh model is generated from the combined 3D point cloud; Lee’s paragraph [0028]: the depth map generator 405 generates the depth map according to the left-eye temporary image, the right-eye temporary image and the disparity). However, Hazeghi as modified by Lee and Wen fails to disclose mapping a plurality of pixels of the images to a plurality of first 3D coordinates in a 3D coordinate system according to camera intrinsic parameters and the depth maps, to construct the 3D mesh of the images.
Skrypnyk discloses mapping a plurality of pixels of the images to a plurality of first 3D coordinates in a 3D coordinate system according to camera intrinsic parameters and the depth maps, to construct the 3D mesh of the images (paragraph [0182]: uses the depth populated image template and the color populated image template to create a point cloud, which is a collection of 3D points representing the spatial information of the scene by mapping the pixel coordinates of the color image to 3D coordinates using the depth values and the camera's intrinsic parameters (focal length, principal point, etc.). The interaction system 100 reconstructs a 3D mesh from the point cloud by connecting neighboring points to form triangles or other polygonal faces).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi, Lee and Wen’s to map pixel to 3D coordinate system as taught by Skrypnyk, to overlay visual content using model adaption.
Regarding claim 4, Hazeghi as modified by Lee, Wen and Skrypnyk discloses the 3D scene visualization method as claimed in claim 3, wherein the mapping of the left side 3D mesh and the right side 3D mesh is executed based on a preset depth range (Skrypnyk’s paragraph [0182]: uses the depth populated image template and the color populated image template to create a point cloud, which is a collection of 3D points representing the spatial information (preset depth range) of the scene by mapping the pixel coordinates of the color image to 3D coordinates using the depth values and the camera's intrinsic parameters (focal length, principal point, etc.). The interaction system 100 reconstructs a 3D mesh from the point cloud by connecting neighboring points to form triangles or other polygonal faces).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi, Lee and Wen’s to map pixel to 3D coordinate system as taught by Skrypnyk, to overlay visual content using model adaption.
Regarding claim 5, Hazeghi as modified by Lee, Wen and Skrypnyk discloses the 3D scene visualization method as claimed in claim 1, wherein the step of generating the 3D scene mesh according to the stereo depth information, the left side 3D mesh and the right side 3D mesh comprises:
mapping the left side 3D mesh to a world coordinate system to obtain a plurality of left side world coordinate points; mapping the right side 3D mesh to the world coordinate system to obtain a plurality of right side world coordinate points; and combining the plurality of left side world coordinate points and the plurality of right side world coordinate points based on matching relationship between the plurality of left side world coordinate points and the plurality of right side world coordinate points, to generate the 3D scene mesh (Skrypnyk’s paragraph [0182]: uses the depth populated image template and the color populated image template to create a point cloud, which is a collection of 3D points representing the spatial information (preset depth range) of the scene by mapping the pixel coordinates of the color image to 3D coordinates using the depth values and the camera's intrinsic parameters (focal length, principal point, etc.). The interaction system 100 reconstructs a 3D mesh from the point cloud by connecting neighboring points to form triangles or other polygonal faces; Hazeghi’s paragraph [0095]: In operation 550, the depth maps are combined to generate a 3D point cloud (e.g., each of the depth maps may be a point cloud that can be merged with other point clouds); paragraph [0096]: In operation 570, a 3D mesh model is generated from the combined 3D point cloud; Lee’s paragraph [0028]: the depth map generator 405 generates the depth map according to the left-eye temporary image, the right-eye temporary image and the disparity).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi, Lee and Wen’s to map pixel to 3D coordinate system as taught by Skrypnyk, to overlay visual content using model adaption.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Hazeghi U.S. Patent Application 20180130255 in view of Lee U.S. Patent Application 20190379842, in view of Wen U.S. Patent Application 20260148405, in view of Skrypnyk U.S. Patent Application 20240355064, and further in view of Kim U.S. Patent Application 20240233326.
Regarding claim 6, Hazeghi as modified by Lee, Wen and Skrypnyk discloses all the features with respect to claim 5 as outlined above. However, Hazeghi as modified by Lee, Wen and Skrypnyk fails to disclose calibrating depth deviation between the plurality of world coordinate points using the stereo depth information.
Kim discloses calibrating depth deviation between the plurality of world coordinate points using the stereo depth information (paragraph [0021]: correcting, by a stereo matching device, a deviation of a depth value between the first image and the second image in the overlapping region while an object is detected in the first image and the second image by using an object detection network).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi, Lee, Wen and Skrypnyk’s to correct depth deviation as taught by Kim, to display more realistic images.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Hazeghi U.S. Patent Application 20180130255 in view of Lee U.S. Patent Application 20190379842, in view of Wen U.S. Patent Application 20260148405, and further in view of Park U.S. Patent Application 20050012757.
Regarding claim 7, Hazeghi as modified by Lee and Wen discloses obtaining a normal vector of a mesh surface corresponding to each vertex in the 3D scene mesh (Hazeghi’s paragraph [0100]: Surface normals, which can be estimated from the point cloud, can also be used for this purpose (with the idea that for two well-aligned point clouds, the normal vectors at nearby points should also be aligned.)). However, Hazeghi as modified by Lee and Wen fails to disclose determining pixel texture information for each vertex according to the normal vector corresponding to each vertex, preset viewpoints.
Park discloses determining pixel texture information for each vertex according to the normal vector corresponding to each vertex, preset viewpoints (paragraph [0071]: The normal vector calculating unit 410 calculates normal vectors of the pixels of each of the depth maps in operation S520; paragraph [0075]: In operation S550, the redundant data removing unit 440 receives the pixels determined as rendering the same portion of the 3D model, abandons the pixels with low reliabilities among the received pixels, and outputs a set of simple texture images; paragraph [0016]: the reliabilities of the pixels of each of the simple texture images may be determined depending on inner projects of the normal vectors of the pixels of each of the simple texture images and a vector perpendicular to a plane of vision of a camera used for creating the simple texture images).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi, Lee and Wen’s to determine pixel texture as taught by Park, to precisely generate a 3D model and can more quickly edit the 3D model using depth information.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Hazeghi U.S. Patent Application 20180130255 in view of Lee U.S. Patent Application 20190379842, in view of Wen U.S. Patent Application 20260148405, and further in view of Cantero Clares U.S. Patent Application 20230239458.
Regarding claim 10, Hazeghi as modified by Lee and Wen discloses all the features with respect to claim 1 as outlined above. However, Hazeghi as modified by Lee and Wen fails to disclose generating a first viewpoint image and a second viewpoint image of a side-by-side image according to the 3D scene mesh; and performing a 3D display operation according to the side-by-side image using a stereo display device.
Cantero Clares discloses generating a first viewpoint image and a second viewpoint image of a side-by-side image according to the 3D scene mesh; and performing a 3D display operation according to the side-by-side image using a stereo display device (paragraph [0061]: Step S470: Capturing an output side-by-side image from the output three-dimensional mesh, wherein the output side-by-side image includes a left-eye image and a right-eye image; paragraph [0062]: Step S480: Weaving the left-eye image and the right-eye image into an output image, and displaying the output image on the stereoscopic-image display device).
Therefore, it would have been obvious before the effective filing date of the claimed invention to combine Hazeghi, Lee and Wen’s to display side-by-side image as taught by Cantero Clares, to generate stereo image more efficiently.
Allowable Subject Matter
Claim 8 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Claim 8 is about the step of determining the pixel texture information for each vertex according to the normal vector corresponding to each vertex, the preset left eye viewpoint and the preset right eye viewpoint comprises:
determining a left side texture weight according to an angle between the normal vector corresponding to a first vertex and the preset left eye viewpoint; determining a right side texture weight according to an angle between the normal vector corresponding to the first vertex and the preset right eye viewpoint; and performing weighted calculation on texture information of the left eye image and texture information of the right eye image according to the left side texture weight and the right side texture weight, to obtain the pixel texture information of the first vertex.
Hazeghi 20180130255, Lee 20190379842, Wen 20260148405, Park 20050012757 and George 20210375044 combined cannot teach these features perfectly. These limitations when read in light of the rest of the limitations in the claim and the claims to which it depends make the claim allowable subject matter.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Yi Yang whose telephone number is (571)272-9589. The examiner can normally be reached on Monday-Friday 9:00 AM-6:00 PM EST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached on 571-272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
/YI YANG/
Primary Examiner, Art Unit 2616