Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Applicant’s Arguments
Applicant’s arguments filed 06/29/2026 have been fully considered, but they are not deemed to be persuasive. Applicant amended claims to show “the user input is utilized ‘as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections’ for generating the dimensionally accurate three-dimensional map” which is obvious in view of the new reference of SINGHAL et al (US 20210357107). Specifically, Singhal teaches the claimed user input to be utilized as “a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections” (Singhal, [0056] - The interface may present a view seen by a camera and enable a user to physically mark corner points of a boundary rectangle to define boundary information. The interface may use a simultaneous location and mapping (SLAM) technique to determine a boundary width and a boundary height of the boundary rectangle). Additionally, Eder’s user input to create a new object on the virtual representation (e.g., [0183] - a user may create and spatially localize geometries comprising 3D shapes, objects, and structures in the 3D model; [0186] - a user may suggest modifications to the virtual representation through the graphical user interface) in which Eder’s new point indicating a corner (e.g., [0061] - In another example, a user can interactively indicate the sequence of corners and walls corresponding to the layout of the location to create a floor plan; [0101] – For example, the geometric reconstruction may be performed by structure-from-motion (SFM), simultaneous localization and mapping (SLAM), or other similar approaches that triangulate 3D points and estimate both camera poses up to scale and camera intrinsics) is “utilized as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections” through the layout estimate algorithm for generating the dimensionally accurate three-dimensional map (e.g., [0179] - the layout estimation algorithm is further configured to generate or update the floor plan based on constraints and heuristics, the heuristics comprising orthogonal junctions of the location, planar assumptions about a composition of the location, and other layout assumptions related to the location). Accordingly, the claimed invention as represented in claims 1, 10 and 16 does not represent a patentable distinction over the art of record.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over TOSCHI et al (Improving automated 3D reconstruction methods via vision metrology) in view of EDER et al (US 20210279957) and SINGHAL et al (US 20210357107).
As per claim 1, Toschi teaches the claimed "method" comprising: "executing a three-dimensional mapping operation via a computing system" (Toschi, Abstract - This paper aims to provide a procedure for improving automated 3D reconstruction methods via vision metrology), including: "capturing an image sequence of a plurality of physical surfaces via an image capture device" (Toschi, 3.2 Image acquisition - Figure 4 (The photogrammetric network geometry) shows the camera network actually realized); and "automated detection of intersection of the plurality of physical surfaces for generating a dimensionally accurate three-dimensional map of the plurality of physical surfaces" (Toschi, 3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates...) (Noted: Toschi's reconstruction of the 3D object based on the detected feature points on its 2D images (e.g., Toshi's SIFT (Scale-Invariant Feature Transform) points are designed to be highly distinctive, often occurring at corners, high-contrast edges within an image. While SIFT aims to identify unique features, it often detects points along edges, which are then refined to eliminate unstable, low-contrast, or edge- based points, favoring strong corner-like structures). It is noted that Toschi does not explicitly teach "user-input to identify location of a surface edge" for improving "automated detection of intersection" as claimed. However, Toschi's technique improvement of the detection of the image features by projecting fine texture with recognizable feature points onto the surface (Toschi, 3.2 Image acquisition and Figure 3 (left image) - Since the test-object is poorly textured, two Optoma Pico projectors (PK301, resolution of 854 X 480 pixels) are used to project a fine texture with recognizable feature points onto the surface. A pebble and gravel image is projected from a distance of about 1 m, so that the object is illuminated with a sharp pattern; 3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates) suggests the claimed user-assisted feature identification of the image features on the captured 2D images; furthermore, in the same endeavor, Singhal and Eder teach, in addition to automatically detect the image features, the user can assist the identification of image features by manually interact with the captured 2D images as claimed by "receiving a user input corresponding to the image sequence, the user input identifying a location of a surface edge from the plurality of physical surfaces by placing an indicator of the location of the surface edge in the image sequence" and “utilizing the user input to improve as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections for generating a dimensionally accurate three-dimensional map of the plurality of physical surfaces”(Singhal, [0056] - The interface may present a view seen by a camera and enable a user to physically mark corner points of a boundary rectangle to define boundary information. The interface may use a simultaneous location and mapping (SLAM) technique to determine a boundary width and a boundary height of the boundary rectangle); additionally, Eder’s user input to create a new object on the virtual representation (e.g., [0183] - a user may create and spatially localize geometries comprising 3D shapes, objects, and structures in the 3D model; [0186] - a user may suggest modifications to the virtual representation through the graphical user interface) in which Eder’s new point indicating a corner (e.g., [0061] - In another example, a user can interactively indicate the sequence of corners and walls corresponding to the layout of the location to create a floor plan; [0101] – For example, the geometric reconstruction may be performed by structure-from-motion (SFM), simultaneous localization and mapping (SLAM), or other similar approaches that triangulate 3D points and estimate both camera poses up to scale and camera intrinsics) is “utilized as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections” through the layout estimate algorithm for generating the dimensionally accurate three-dimensional map (e.g., [0179] - the layout estimation algorithm is further configured to generate or update the floor plan based on constraints and heuristics, the heuristics comprising orthogonal junctions of the location, planar assumptions about a composition of the location, and other layout assumptions related to the location). Thus, it would have been obvious, in view of Eder and Singhal, to configure Toschi's method as claimed by manually marking surface corners to assist the system in utilized as a constraint on automated detection of intersections of the plurality of physical surfaces (i.e., features of an image) in constructing 3D model from the captured 2D images. The motivation is to improve the reliability of image feature identification with the assisting of user-selected edges.
Claim 2 adds into claim 1 "the dimensionally accurate three-dimensional map includes a virtual recreation of an interior of a building" (Eder, [0049] - According to some embodiments, a virtual representation of a location is created. In an embodiment, the location refers to any open or closed spaces for which virtual representation may be generated. For example, the location may be a physical scene, a room, a warehouse, a classroom, an office space, an office room, a restaurant room, a coffee shop, etc.). Thus, it would have been obvious, in view of Singhal and Eder, to configure Toschi's method as claimed by interactively providing a virtual recreation of an interior of a building. The motivation is to interactively providing a visual representation of a virtual recreation of an interior of a building.
Claim 3 adds into claim 1 "generating a user interface (UI) for the image capture device; and receiving the user input via the UI for the image capture device while capturing the image sequence" (Singhal, [0056] - The interface may present a view seen by a camera and enable a user to physically mark corner points of a boundary rectangle to define boundary information. The interface may use a simultaneous location and mapping (SLAM) technique to determine a boundary width and a boundary height of the boundary rectangle; Eder, [0051] - The graphical user interface also enables a user to pause and resume data capture within the location). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by generating a UI to interactively control the cameras. The motivation is to interactively manipulate the camera's position and orientation on capturing images.
Claim 4 adds into claim 1 "receiving the user input as a position indicator within the image sequence" (Singhal, [0056] - The interface may present a view seen by a camera and enable a user to physically mark corner points of a boundary rectangle to define boundary information. The interface may use a simultaneous location and mapping (SLAM) technique to determine a boundary width and a boundary height of the boundary rectangle) (Eder, [0120] - In an embodiment, a user may use a graphical user interface to manually specify per-pixel semantic or instance labels by selecting pixels individually or in groups and assigning a desired label. This 2-dimensional object localization and annotation may also be performed by a user manually through the graphical user interface; [0181] - In an embodiment, a graphical user interface may be configured to view or modify the virtual representation. In an example, where the user uses a mouse and keyboard, this capability may be provided through a combination of clicking and holding mouse button(s) or key(s) on a keyboard. In another example, where the user uses a touch screen, this capability may be provided through a combination of touches, holds, and gestures; [0184] - In an example, a user may make temporary or permanent changes to the hue, saturation, brightness, transparency, and other aspects of the appearance of the virtual representation, including elements of its geometry, annotations, and associated metadata. For example, different annotations may be assigned a specific color or made to appear or disappear depending on the user's needs); and "storing the position indicator as metadata associated with a location in the image sequence" (Eder, [0061] - As mentioned earlier, metadata refers to a set of data that describes and gives information about other data. For example, the metadata associated with an image may include items such as a GPS coordinates of the location where the image was taken, the date and time it was taken, camera type and image capture settings, the software used to edit the image, or other information related to the image, the location or the camera. In an embodiment, the metadata may include information about elements of the locations, such as information about a wall, a chair, a bed, a floor, a carpet, a window, or other elements that may be present in the captured images or video In another example, a user can interactively indicate the sequence of corners and walls corresponding to the layout of the location to create a floor plan). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by interactively collecting the user input of locations on the 2D ages. The motivation is to interactively define the feature locations on the captured 2D images for reconstructing a 3D object.
Claim 5 adds into claim 1 "generating the dimensionally accurate three- dimensional map via application of a simultaneous localization and mapping (SLAM) algorithm" (Eder, [0101] - For example, the geometric reconstruction may be performed by structure-from-motion (SFM), simultaneous localization and mapping (SLAM), or other similar approaches that triangulate 3D points and estimate both camera poses up to scale and camera intrinsics). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by generating the dimensionally accurate three- dimensional map via application of a simultaneous localization and mapping (SLAM) algorithm. The motivation is to interactively reconstructing a 3D environment using the captured 2D images.
Claim 6 adds into claim 5 "tracking a plurality of points of light projected in a consistent artificial texture pattern between individual images in the image sequence" (Toschi, 3.2 Image acquisition and Figure 3 (left image) - Since the test-object is poorly textured, two Optoma Pico projectors (PK301, resolution of 854 X 480 pixels) are used to project a fine texture with recognizable feature points onto the surface. A pebble and gravel image is projected from a distance of about 1 m, so that the object is illuminated with a sharp pattern); and utilizing the plurality of points of light for generating the dimensionally accurate three-dimensional map" (Toschi, 3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates).
Claim 7 adds into claim 6 "providing the plurality of points of light to the SLAM algorithm as feature points in a point cloud to track between the individual images in the image sequence" (Toschi, 3.2 Image acquisition and Figure 3 (left image) - Since the test-object is poorly textured, two Optoma Pico projectors (PK301, resolution of 854 X 480 pixels) are used to project a fine texture with recognizable feature points onto the surface. A pebble and gravel image is projected from a distance of about 1 m, SO that the object is illuminated with a sharp pattern; 3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates; Eder, [0047] - The original surface can be recovered using an algorithm of the class of isosurface extraction algorithms comprising marching cubes, among others. "Structure from Motion" (SFM) refers to a class of algorithms that estimate intrinsics and extrinsic camera parameters, as well as a scene structured in the form of a sparse point cloud. SFM can be applied to both ordered image data, such as frames from a video, as well unordered data, such as random images of a scene from one or more different camera sources. Traditionally, SFM algorithms are computationally expensive and are used in an offline setting. "Simultaneous localization and mapping" (SLAM) refers to a class of algorithms that estimate both camera pose and scene structure in the form of point cloud. SLAM is applicable to ordered data, for example, a video stream. SLAM algorithms may operate at interactive rates, and can be used in online settings). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by tracking the light
points captured in the sequence of 2D images which form the reconstructed points in the point cloud using the SLAM algorithm. The motivation is to enhance the detection of feature points which form a 3D cloud point of the model.
Claim 8 adds into claim 1 "displaying the dimensionally accurate three- dimensional map as a virtual recreation of a real-world environment corresponding to the plurality of physical surfaces" (Eder, [0049] - According to some embodiments, a virtual representation of a location is created. In an embodiment, the location refers to any open or closed spaces for which virtual representation may be generated. For example, the location may be a physical scene, a room, a warehouse, a classroom, an office space, an office room, a restaurant room, a coffee shop, etc... At operation 4005, the captured data 4000 is used as input to construct a virtual representation VIR). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by displaying the dimensionally accurate three-dimensional map as a virtual recreation of a real-world environment corresponding to the plurality of physical surfaces. The motivation is to interactively provide a visual representation of real world environment on a display.
Claim 9 adds into claim 8, "displaying the virtual recreation as a virtual reality environment" (Eder, [0049] - According to some embodiments, a virtual representation of a location is created. In an embodiment, the location refers to any open or closed spaces for which virtual representation may be generated. For example, the location may be a physical scene, a room, a warehouse, a classroom, an office space, an office room, a restaurant room, a coffee shop, etc. At operation 4005, the captured data 4000 is used as input to construct a virtual representation VIR) (Noted: Eder's 3D virtual scene, e.g., an office space, can be used in a virtual reality environment in which a virtual object, e.g., a chair, is merged into the virtual scene of the office scene). The motivation is to enhance a visual representation to a user by combining multiple virtual 3D objects into a virtual 3D environment.
As per claim 10, Toschi teaches the claimed "memory device storing instructions" that, when executed, cause a processor to perform a method comprising: "executing a three-dimensional mapping operation via a computing system" (Toschi, Abstract - This paper aims to provide a procedure for improving automated 3D reconstruction methods via vision metrology), including: "capturing an image sequence of a plurality of physical surfaces via an image capture device" (Toschi, 3.2 Image acquisition - Figure 4 (The photogrammetric network geometry) shows the camera network actually realized); "tracking a plurality of points of light projected in a consistent artificial texture pattern between individual images in the image sequence" (Toschi, 3.2 Image acquisition and Figure 3 (left image) - Since the test-object is poorly textured, two Optoma Pico projectors (PK301, resolution of 854 X 480 pixels) are used to project a fine texture with recognizable feature points onto the surface. A pebble and gravel image is projected from a distance of about 1 m, so that the object is illuminated with a sharp pattern); and "utilizing the plurality of points of light to improve automated detection of intersection of the plurality of physical surfaces in the image sequence for generating a dimensionally accurate three-dimensional map of the plurality of physical surfaces" (3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates). It is noted that Toschi does not explicitly teach "utilizing the plurality of points of light and the user input, wherein the user input is applied as a constraint on to improve automated detection of intersections of the plurality of physical surfaces in the image sequence to modify determination of intersections for generating a dimensionally accurate three-dimensional map of the plurality of physical surfaces" as claimed. However, Toschi's technique for improvement of the detection of the intersection by projecting fine texture with recognizable feature points onto the surface (Toschi, 3.2 Image acquisition and Figure 3 (left image) - Since the test-object is poorly textured, two Optoma Pico projectors (PK301, resolution of 854 X 480 pixels) are used to project a fine texture with recognizable feature points onto the surface. A pebble and gravel image is projected from a distance of about 1 m, so that the object is illuminated with a sharp pattern; 3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates) suggests the claimed user-assisted feature identification of the image features on the captured 2D images; furthermore, in the same endeavor, Singhal teaches, in addition to automatically detect the image features, the user can assist the identification of image features by manually interact with the captured 2D images as claimed by "receiving a user input corresponding to the image sequence, the user input identifying a location of a surface edge from the plurality of physical surfaces by placing an indicator of the location of the surface edge in the image sequence" and “utilizing the user input to improve as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections for generating a dimensionally accurate three-dimensional map of the plurality of physical surfaces”(Singhal, [0056] - The interface may present a view seen by a camera and enable a user to physically mark corner points of a boundary rectangle to define boundary information. The interface may use a simultaneous location and mapping (SLAM) technique to determine a boundary width and a boundary height of the boundary rectangle); additionally, Eder’s user input to create a new object on the virtual representation (e.g., [0183] - a user may create and spatially localize geometries comprising 3D shapes, objects, and structures in the 3D model; [0186] - a user may suggest modifications to the virtual representation through the graphical user interface) in which Eder’s new point indicating a corner (e.g., [0061] - In another example, a user can interactively indicate the sequence of corners and walls corresponding to the layout of the location to create a floor plan; [0101] – For example, the geometric reconstruction may be performed by structure-from-motion (SFM), simultaneous localization and mapping (SLAM), or other similar approaches that triangulate 3D points and estimate both camera poses up to scale and camera intrinsics) is “utilized as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections” through the layout estimate algorithm for generating the dimensionally accurate three-dimensional map (e.g., [0179] - the layout estimation algorithm is further configured to generate or update the floor plan based on constraints and heuristics, the heuristics comprising orthogonal junctions of the location, planar assumptions about a composition of the location, and other layout assumptions related to the location). Thus, it would have been obvious, in view of Eder and Singhal, to configure Toschi's method as claimed by manually marking surface corners to assist the system in utilized as a constraint on automated detection of intersections of the plurality of physical surfaces (i.e., features of an image) in constructing 3D model from the captured 2D images. The motivation is to improve the reliability of image feature identification with the assisting of user-selected edges.
Claim 11 adds into claim 10 "generating a user interface (UI) for the image capture device; and receiving the user input via the UI for the image capture device while capturing the image sequence" (Eder, [0051] - The graphical user interface also enables a user to pause and resume data capture within the location). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by generating a UI to interactively control the cameras. The motivation is to interactively manipulate the camera's position and orientation on capturing images.
Claim 12 adds into claim 10 "receiving the user input as a position indicator within the image sequence" (Eder, [0120] - In an embodiment, a user may use a graphical user interface to manually specify per-pixel semantic or instance labels by selecting pixels individually or in groups and assigning a desired label This 2- dimensional object localization and annotation may also be performed by a user manually through the graphical user interface; [0181] - In an embodiment, a graphical user interface may be configured to view or modify the virtual representation. In an example, where the user uses a mouse and keyboard, this capability may be provided through a combination of clicking and holding mouse button(s) or key(s) on a keyboard. In another example, where the user uses a touch screen, this capability may be provided through a combination of touches, holds, and gestures; [0184] - In an example, a user may make temporary or permanent changes to the hue, saturation, brightness, transparency, and other aspects of the appearance of the virtual representation, including elements of its geometry, annotations, and associated metadata. For example, different annotations may be assigned a specific color or made to appear or disappear depending on the user's needs); and "storing the position indicator as metadata associated with a location in the image sequence" (Eder, [0061] - As mentioned earlier, metadata refers to a set of data that describes and gives information about other data. For example, the metadata associated with an image may include items such as a GPS coordinates of the location where the image was taken, the date and time it was taken, camera type and image capture settings, the software used to edit the image, or other information related to the image, the location or the camera. In an embodiment, the metadata may include information about elements of the locations, such as information about a wall, a chair, a bed, a floor, a carpet, a window, or other elements that may be present in the captured images or video In another example, a user can interactively indicate the sequence of corners and walls corresponding to the layout of the location to create a floor plan). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by interactively collecting the user input of locations on the 2D images. The motivation is to interactively define the feature locations on the captured 2D images for reconstructing a 3D object.
Claim 13 adds into claim 10 "generating the dimensionally accurate three- dimensional map via application of a simultaneous localization and mapping (SLAM) algorithm" (Eder, [0101] - For example, the geometric reconstruction may be performed by structure-from-motion (SFM), simultaneous localization and mapping (SLAM), or other similar approaches that triangulate 3D points and estimate both camera poses up to scale and camera intrinsics); and "providing the plurality of points of light to the SLAM algorithm as feature points in a point cloud to track between the individual images in the image sequence" (Toschi, 3.2 Image acquisition and Figure 3 (left image) - Since the test-object is poorly textured, two Optoma Pico projectors (PK301, resolution of 854 X 480 pixels) are used to project a fine texture with recognizable feature points onto the surface. A pebble and gravel image is projected from a distance of about 1 m, so that the object is illuminated with a sharp pattern; 3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates; Eder, [0047] - The original surface can be recovered using an algorithm of the class of isosurface extraction algorithms comprising marching cubes, among others. "Structure from Motion" (SFM) refers to a class of algorithms that estimate intrinsics and extrinsic camera parameters, as well as a scene structured in the form of a sparse point cloud. SFM can be applied to both ordered image data, such as frames from a video, as well unordered data, such as random images of a scene from one or more different camera sources. Traditionally, SFM algorithms are computationally expensive and are used in an offline setting. "Simultaneous localization and mapping" (SLAM) refers to a class of algorithms that estimate both camera pose and scene structure in the form of point cloud. SLAM is applicable to ordered data, for example, a video stream. SLAM algorithms may operate at interactive rates, and can be used in online settings). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by tracking the light points captured in the sequence of 2D images which form the reconstructed points in the point cloud using the SLAM algorithm. The motivation is to enhance the detection of feature points which form a 3D cloud point of the model.
Claim 14 adds into claim 10 "capturing movement data from a sensor of the image capture device" (Eder, [0077] - For example, the IMU sensor may include, but is not limited to, accelerometers, gyroscopes, and magnetometers that can be used to determine metric distance as the camera moves. These devices provide a metric camera position in a coordinate system scaled to correspond to the scene that enables visual-inertial simultaneous localization and mapping (VI-SLAM) or use of odometry algorithm); and "providing the movement data to the SLAM algorithm to determine a path of the image capture device while capturing the image sequence" (Eder, [0050] - the virtual representation VIR may be represented as a 3D model of the location with metadata MD1 comprising data associated images, videos, natural language, camera trajectory, and geometry.. [0077] - These devices provide a metric camera position in a coordinate system scaled to correspond to the scene that enables visual-inertial simultaneous localization and mapping (VI-SLAM) (i.e., VI-SLAM may be a particular type of SLAM algorithm that performs SLAM using both image and IMU data) or use of odometry algorithm). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by tracking the camera positions to determine a path of the image capture device while capturing the image sequence using the VI-SLAM algorithm. The motivation is to visually provide the camera path showing the locations of camera taking the images.
Claim 15 adds into claim 10 "displaying the dimensionally accurate three- dimensional map as a virtual reality recreation of a real-world environment corresponding to the plurality of physical surfaces" (Eder, [0049] - According to some embodiments, a virtual representation of a location is created. In an embodiment, the location refers to any open or closed spaces for which virtual representation may be generated. For example, the location may be a physical scene, a room, a warehouse, a classroom, an office space, an office room, a restaurant room, a coffee shop, etc At operation 4005, the captured data 4000 is used as input to construct a virtual representation VIR). Thus, it would have been obvious, in view of Singhal and Eder, to configure Toschi's method as claimed by displaying the dimensionally accurate three-dimensional map as a virtual recreation of a real-world environment corresponding to the plurality of physical surfaces. The motivation is to interactively provide a visual representation of real world environment on a display.
As per claim 16, Toschi teaches the claimed "apparatus" comprising: "configured to execute a three-dimensional mapping operation" (Toschi, Abstract - This paper aims to provide a procedure for improving automated 3D reconstruction methods via vision metrology), including: "capture an image sequence of a plurality of physical surfaces via an image capture device" (Toschi, 3.2 Image acquisition - Figure 4 (The photogrammetric network geometry) shows the camera network actually realized); “receive a user input corresponding to the image sequence, the user input identifying a location of a surface edge from the plurality of physical surfaces by placing an indicator of the location of the surface edge in the image sequence; and utilize the user input as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of the intersections for generating a dimensionally accurate three-dimensional map” (Singhal, [0056] - The interface may present a view seen by a camera and enable a user to physically mark corner points of a boundary rectangle to define boundary information. The interface may use a simultaneous location and mapping (SLAM) technique to determine a boundary width and a boundary height of the boundary rectangle); additionally, Eder’s user input to create a new object on the virtual representation (e.g., [0183] - a user may create and spatially localize geometries comprising 3D shapes, objects, and structures in the 3D model; [0186] - a user may suggest modifications to the virtual representation through the graphical user interface) in which Eder’s new point indicating a corner (e.g., [0061] - In another example, a user can interactively indicate the sequence of corners and walls corresponding to the layout of the location to create a floor plan; [0101] – For example, the geometric reconstruction may be performed by structure-from-motion (SFM), simultaneous localization and mapping (SLAM), or other similar approaches that triangulate 3D points and estimate both camera poses up to scale and camera intrinsics) is “utilized as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections” through the layout estimate algorithm for generating the dimensionally accurate three-dimensional map (e.g., [0179] - the layout estimation algorithm is further configured to generate or update the floor plan based on constraints and heuristics, the heuristics comprising orthogonal junctions of the location, planar assumptions about a composition of the location, and other layout assumptions related to the location). Thus, it would have been obvious, in view of Eder and Singhal, to configure Toschi's method as claimed by manually marking surface corners to assist the system in utilized as a constraint on automated detection of intersections of the plurality of physical surfaces (i.e., features of an image) in constructing 3D model from the captured 2D images. The motivation is to improve the reliability of image feature identification with the assisting of user-selected edges.
Claim 17 adds into claim 16 “track a plurality of points of light projected in a consistent artificial texture pattern between individual images in the image sequence; and utilize the plurality of points of light for generating the dimensionally accurate three- dimensional map of the plurality of physical surfaces” (Toschi, 3.2 Image acquisition and Figure 3 (left image) - Since the test-object is poorly textured, two Optoma Pico projectors (PK301, resolution of 854 X 480 pixels) are used to project a fine texture with recognizable feature points onto the surface. A pebble and gravel image is projected from a distance of about 1 m, so that the object is illuminated with a sharp pattern; 3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates). It is noted that Toschi's technique for improvement of the detection of the intersection's surface edges by projecting fine texture with recognizable feature points onto the surface (Toschi, 3.2 Image acquisition and Figure 3 (left image) - Since the test-object is poorly textured, two Optoma Pico projectors (PK301, resolution of 854 X 480 pixels) are used to project a fine texture with recognizable feature points onto the surface. A pebble and gravel image is projected from a distance of about 1 m, so that the object is illuminated with a sharp pattern; 3.3 Image processing - images with the projected pattern are imported into a typical SfM software, namely Agisoft Photoscan (PS). Its SIFT-like operator is exploited to automatically extract a large number of homologous points, these are then exported in the form of both image observations (2D points) and corresponding 3D coordinates) suggests optional techniques can be used to improve the process of feature identification of the image features on the captured 2D images; furthermore, in the same endeavor, Eder teaches, in addition to automatically detect the image features, the user can assist the identification of image features by manually interact with the captured 2D images as claimed by "receiving a user input corresponding to the image sequence, the user input identifying a location of a surface edge from the plurality of physical surfaces by placing an indicator of the location of the surface edge in the image sequence" (Eder, [0061] - In another example, a user can interactively indicate the sequence of corners and walls corresponding to the layout of the location to create a floor plan). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by manually selecting surface edge to assist the system in detecting of intersections of the edges (i.e., features of an image) in constructing 3D model from the captured 2D images. The motivation is to improve the reliability of image feature identification with the assisting of user-selected edges.
Claim 18 adds into claim 16 "a mobile computing device including: the processor; the image capture device; and a position sensor configured to generate movement information for the mobile computing device; and the processor is further configured to generate the dimensionally accurate three-dimensional map further based on the movement information" (Eder, [0058] - In an embodiment, the description data may be captured by a mobile computing device associated with a user and transmitted to the one or more processors with or without a first user and/or other user interaction; [0077] - For example, the IMU sensor may include, but is not limited to, accelerometers, gyroscopes, and magnetometers that can be used to determine metric distance as the camera moves. These devices provide a metric camera position in a coordinate system scaled to correspond to the scene that enables visual-inertial simultaneous localization and mapping (VI-SLAM) or use of odometry algorithm; [0078] - These devices may include, but are not limited to, traditional video cameras, DSLR cameras, multi-camera rigs, and certain smartphone cameras. Some of these devices may provide relative camera pose information along with the RGB images using a built-in visual odometry or visual simultaneous localization and mapping (SLAM) algorithm). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by displaying the dimensionally accurate three-dimensional map as a virtual recreation of a real-world environment corresponding to the plurality of physical surfaces on the mobile computer based on the captured 2D images. The motivation is to interactively provide a visual representation of real world environment on the mobile's display.
Claim 19 adds into claim 16 "a virtual reality display component; and the processor is further configured to present the dimensionally accurate three-dimensional map as a virtual reality recreation of a real-world environment corresponding to the plurality of physical surfaces using the virtual reality display component" (Eder, [0049] - According to some embodiments, a virtual representation of a location is created. In an embodiment, the location refers to any open or closed spaces for which virtual representation may be generated. For example, the location may be a physical scene, a room, a warehouse, a classroom, an office space, an office room, a restaurant room, a coffee shop, etc. At operation 4005, the captured data 4000 is used as input to construct a virtual representation VIR). Thus, it would have been obvious, in view of Eder, to configure Toschi's method as claimed by displaying the dimensionally accurate three-dimensional map as a virtual recreation of a real-world environment corresponding to the plurality of physical surfaces. The motivation is to interactively provide a visual representation of real world environment on a display.
Claim 20 adds into claim 19 "present the dimensionally accurate three- dimensional map as an interactive virtual property tour" (Eder, [0049] - According to some embodiments, a virtual representation of a location is created. In an embodiment, the location refers to any open or closed spaces for which virtual representation may be generated. For example, the location may be a physical scene, a room, a warehouse, a classroom, an office space, an office room, a restaurant room, a coffee shop, etc At operation 4005, the captured data 4000 is used as input to construct a virtual representation VIR) (Noted: Eder's 3D virtual scene, e.g., an office space, can be used in a virtual reality environment in which a virtual object, e.g., a chair, is merged into the virtual scene of the office space). The motivation is to enhance a visual representation to a user by combining multiple virtual 3D objects into a virtual 3D environment.
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1, 5, 10, 14 and 16 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12,112,431. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims of the US patent contains all features of the claimed invention of the application
Claims of the application
Claims of the US patent
1.A method comprising: executing a three-dimensional mapping operation via a computing system, including:
capturing an image sequence of a plurality of physical surfaces via an image capture device;
receiving a user input corresponding to the image sequence, the user input identifying a location of a surface edge from the plurality of physical surfaces by placing an indicator of the location of the surface edge in the image sequence; and
utilizing the user input as a constraint on automated detection of intersections of the plurality of physical surfaces to modify determination of intersections for generating a dimensionally accurate three-dimensional map of the plurality of physical surfaces
A method comprising: executing a three dimensional mapping operation via a computing system, including:
capturing an image sequence of a plurality of physical surfaces via an image capture device;
receiving a user input corresponding to the image sequence, the user input identifying a location of a surface edge from the plurality of physical surfaces by placing an indicator of the location of the surface edge in the image sequence; tracking a plurality of feature points between individual images in the image sequence based on a projected light pattern situated to create a consistent artificial texture pattern on the plurality of physical surfaces; and generating a three dimensional map of the plurality of physical surfaces based on the image sequence, the user input identifying the location of the surface edge, and the plurality of feature points.
The grounds of the rejection of the remaining claims are provided in the previous action.
.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHU K NGUYEN whose telephone number is (571)272-7645. The examiner can normally be reached M-F 8-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel F. Hajnik can be reached at (571) 272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHU K NGUYEN/Primary Examiner, Art Unit 2616