DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 5-10 and 13-20 are rejected under 35 U.S.C. 103 as being unpatentable over US 11,922,580 B2 to Tang et al (hereinafter ‘Tang’).
Regarding claim 1, Tang discloses a method for determining three-dimensional layout information performed by a computer device (column 1, lines 54-56, wherein devices, systems, and methods that generate floorplans and measurements using three-dimensional (3D) representations of a physical environment) and comprising: obtaining a first image and a second image of a 3D region using a first camera and a second camera simultaneously (column 18, lines 62-67, and column 7, lines 25-26, wherein the image source(s) may include a depth camera 402 that acquires depth data 404 of the physical environment, and a light intensity camera 406 (e.g., RGB camera) that acquires light intensity image data 408 (e.g., a sequence of RGB image frames), and wherein a 3D point cloud may be generated based on depth camera information received concurrently with the images); and generating three-dimensional layout information of the 3D region based on a relative position between the first camera and the second camera (column 12, lines 11-20, wherein the live preview unit 244 obtains a sequence of light intensity images from a light intensity camera (e.g., a live camera feed), a semantic 3D representation (e.g., semantic 3D point cloud) generated from the 3D representation unit 242, as generated 3d layout, and other sources of physical environment information (e.g., camera positioning information from a camera's simultaneous localization and mapping (SLAM) system) to output a 2D floorplan image that is iteratively updated with the sequence of light intensity images.), a photographing parameter of each of the first camera and the second camera (column 16, lines 57-59, wherein sources of physical environment information (e.g., camera positioning information from a camera's SLAM system)), the first image, and the second image, the three- dimensional layout information being configured for characterizing a three-dimensional spatial layout of at least one real object in the 3D region (column 21, lines 8-14, wherein at block 506, the method 500 generates a live preview of a preliminary 2D floorplan of the physical environment based on the 3D representation of the physical environment. For example, 2D top-down view of a preliminary floorplan of the physical environment 105 may be generated that includes the structures identified in the room (e.g., walls, table, door, window, etc.)). Tang does not specifically disclose relative position between the first and second camera however, Tang discloses instead of having a depth camera, the system could determine depth based on stereo imaging and therefore the positioning between the cameras are inherently considered in depth determination (Column 20, lines 50-53, wherein the one or more depth cameras can acquire depth based on structured light (SL), passive stereo (PS), active stereo (AS), time-of-flight (ToF), and the like.). Therefore, it would have been obvious to one of ordinary skill in the art to combine the stereo imaging with 3D space generation of Tang’s method so that to facilitate stereoscopic display images and/or a screen for viewing on a head-mounted display (HMD) (column 22, lines 40-41).
Regarding claim 2, Tang discloses wherein the generating three-dimensional layout information of the 3D region based on a relative position between the first camera and the second camera, a photographing parameter of each of the first camera and the second camera, the first image, and the second image comprises: generating an image feature of the first image and an image feature of the second image (column 11, lines 59-62, wherein obtain light intensity image data (e.g., RGB) and depth image data and generate a semantic 3D representation (e.g., a 3D point cloud with associated semantic labels, as the features); fusing the image feature of the first image and the image feature of the second image based on the relative position between the first camera and the second camera and the photographing parameter of each of the first camera and the second camera, to generate a three-dimensional feature of the 3D region, the three-dimensional feature being configured for characterizing spatial information of the 3D region (column 11, lines 43-51, wherein the 3D representation unit 242 is configured with instructions executable by a processor to obtain image data (e.g., light intensity data, depth data, etc.) and integrate (e.g., fuse) the image data using one or more of the techniques disclosed herein. For example, the 3D representation unit 242 fuses RGB images from a light intensity camera with a sparse depth map from a depth camera (e.g., time-of-flight sensor) and other sources of physical environment information to output a dense depth point cloud of information.); and generating the three-dimensional layout information of the 3D region based on the three-dimensional feature of the 3D region (column 19, lines 49-52, wherein the semantic representation unit 440 generates a semantically labeled 3D point cloud 447 by acquiring the 3D point cloud data 424 and the semantic segmentation 434 using a semantic 3D algorithm that fuses the 3D data and semantic labels).
Regarding claim 5, Tang discloses wherein the three-dimensional layout information of the 3D region comprises three-dimensional pose information and annotation information of the at least one real object in the 3D region (column 4, lines 65 through column 5, line 1, and column 19, lines 13-15, wherein the 3D bounding boxes may provide location, pose (e.g., location and orientation), and shape of each piece furniture and appliance in the room, and wherein the different size grey dots, inherently as annotations, in the 3D point cloud 424 represent different depth values detected within the depth data).
Regarding claim 6, Tang discloses wherein the first camera and the second camera are arranged on a same device with predefined relative positions of the first camera and the second camera (column 20, lines 50-52, wherein the one or more depth cameras can acquire depth based on structured light (SL), passive stereo (PS), active stereo (AS), inherently as cameras on the same device at a given predetermined positions).
Regarding claim 7, Tang discloses wherein the method further comprises: constructing a virtual scene or a virtual object adapted to the 3D region based on the three-dimensional layout information; and displaying the virtual scene or the virtual object in the 3D region (column 35, lines 36-40, wherein the systems described herein may include a XR unit that is configured with instructions executable by a processor to provide a XR environment that includes depictions of a physical environment including real physical objects and virtual content.).
Regarding claim 9, Tang discloses a computer device, comprising a processor and a memory, the memory having computer programs stored therein, the computer programs, when executed by the processor, causing the computer device to implement a method for determining three-dimensional layout information (column 8, lines 38-44, wherein a device includes one or more processors, a non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors and the one or more programs include instructions for performing or causing performance of any of the methods) including: Please refer to the corresponding method claim 1 for further teachings.
Regarding device claims 10 and 13-15, please refer to the corresponding method claims 2 and 5-7 above for further teachings.
Regarding claim 16, Tang discloses a non-transitory computer-readable storage medium, having computer programs stored therein, the computer programs, when executed by a processor of a computer device (column 8, lines 38-44, wherein a device includes one or more processors, a non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors and the one or more programs include instructions for performing or causing performance of any of the methods), causing the computer device to implement a method for determining three-dimensional layout information including:
Regarding storage medium claims 10 and 13-15, please refer to the corresponding method claims 2 and 5-7 above for further teachings.
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Tang in view of US 11,120,280 B2 to Hu et al (hereinafter ‘Hu’).
Regarding claim 8, Tang does not specifically disclose wherein the relative position comprises a distance, and the photographing parameter comprises a focal length. Hu discloses wherein the relative position comprises a distance, and the photographing parameter comprises a focal length (column 5, lines 56-58, wherein in which baseline is a measure of the distance between the centers of the stereo camera system, and focal length is a characteristic of the camera optical system). Tang and Hu are combinable because they both disclose 3D image processing and display. Therefore, before the effective filing date of the claimed invention, it would have been obvious to one of ordinary skill in the art to combine the distance between cameras, and the focal length of Hu’s method with Tang’s because disparity representation is better than converting to a depth representation since disparity maps already contain shape information of objects and do not suffer from the quadratically increased error issue (column 6, lines 12-17).
Allowable Subject Matter
Claims 3, 4, 11 and 12 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter: the prior art or the prior art of record specifically, Tang and US 11600039 B2 to Shandilya et al, does not disclose:
. . . the three-dimensional feature fuser being configured to fuse the image feature of the first image and the image feature of the second image based on the relative position between the first camera and the second camera and the photographing parameter of each of the first camera and the second camera, to generate the three-dimensional feature of the 3D region; and the neural network decoder being configured to generate the three-dimensional layout information of the 3D region based on the three-dimensional feature of the 3D region, of claims 3 and 11 combined with other features and elements of the claims;
Claims 4 and 12 depend from an allowable base claim and are thus allowable themselves.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHERVIN K NAKHJAVAN whose telephone number is (571)272-5731. The examiner can normally be reached Monday-Friday 9:00-12:00 PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sue Lefkowitz can be reached at (571)272-3638. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHERVIN K NAKHJAVAN/ Primary Examiner, Art Unit 2672