Prosecution Insights
Last updated: August 17, 2026
Application No. 18/858,256

SPACE VISUALIZATION SYSTEM AND SPACE VISUALIZATION METHOD

Final Rejection §103
Filed
Oct 18, 2024
Priority
Apr 21, 2022 — JP 2022-070193 +1 more
Examiner
LI, JAI WEI TOMMY
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Hitachi Ltd.
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
34 currently pending
Career history
27
Total Applications
across all art units

Statute-Specific Performance

§101
1.1%
-38.9% vs TC avg
§103
76.1%
+36.1% vs TC avg
§102
19.3%
-20.7% vs TC avg
§112
3.4%
-36.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The objections to the specifications and the some of the objections to the claims on informalities have been withdrawn in view of applicant’s amendments filed 06/22/2026. Claim Objections Claim(s) 1, 5, 11-13, 15-16, and 20 objected to because of the following informalities: Introduced the term “the detected object” but the term used in dependent claim include “an object”. Appropriate correction is required. Claim(s) 20 objected to because of the following informalities: Introduced the term “the three-dimensional space” but the term used in dependent claim include “three-dimensional position”. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-5, 7, 8, and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), further in view of Stauber et al. (U.S. Pub. No. 20210074077) and Forsblom et al. (U.S. Pub. No. 20130044137). Regarding claim 1, Newcombe discloses a space visualization system comprising: a calculation device configured to execute predetermined calculation processing; and a storage device accessible by the calculation device (paragraph 7, line(s) 2-4, “a 3D model of a real-world environment is generated in a 3D volume made up of voxels stored on a memory device”; also, paragraph 34, line(s) 3-4, “The mobile environment capture device also comprises one or more processors, a memory and a communications infrastructure”), wherein the calculation device includes a spatial structure recognition unit configured to construct a spatial structure based on a plurality of images (paragraph 22, line(s) 9-12, “images captured by the mobile depth camera 102 are used to form and build up a dense 3D model of the environment as the person moves about the room”; also, paragraph 24, line(s) 1-4, “The real-time camera tracking system 112 provides input to the dense 3D modeling system, in order to allow individual depth images to be built up into an overall 3D model”), calculation device includes an image semantic and spatial structure fusion unit configured to estimate a spatial position of the detected object on the spatial structure (paragraph 53, line(s) 12-15, “the real-world coordinate for the current voxel associated with the thread is perspective projected through the depth camera's projection, and can take into account the intrinsic parameters of the camera”; also, paragraph 62, line(s) 10-12, “The averaged signed distance function values can then be stored 430 at the current voxel.”), storage device stores information on the image as a source of constructing the spatial structure, the constructed spatial structure (paragraph 7, line(s) 2-4, “a 3D model of a real-world environment is generated in a 3D volume made up of voxels stored on a memory device”; also, paragraph 89, line(s) 12-15, “The memory 816 can also provide a data store 830, which can be used to provide storage for data used by the processors 802 when performing the 3D modeling techniques”), calculation device includes an image semantic recognition unit configured to detect an object included in each of the plurality of images, the calculation device includes an image semantic and spatial structure summary unit configured to summarize data to be presented to a user, and information on the detected object including a three-dimensional position and a size of the detected object estimated by projecting a two- dimensional position of the detected object onto the spatial structure using position and posture information of the image. However, in a similar field of endeavor, Choi discloses the calculation device includes an image semantic recognition unit configured to detect an object included in each of the plurality of images (paragraph 29, line(s) 4-6, “a CNN directly processes an input image and outputs bounding boxes of object detections”; also, paragraph 30, line(s) 2-3, “region proposal techniques are utilized to generate a set of object candidates from each image”; also, paragraph 17, line(s) 1-5, “the present invention proposes a method and system for solving visual object detection given an image. The goal is to identify the object of interest (e.g., car, pedestrian, etc.) and estimate the location of such object in the image space”) and information on the detected object (paragraph 7, line(s) 3-10, “The system includes a memory and a processor in communication with the memory, wherein the processor is configured to generate object region proposals from an image by a region proposal network (RPN) which utilizes subcategory information, and classify and refine the object region proposals by an object detection network (ODN) that simultaneously performs object category classification, subcategory classification, and bounding box regression”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a sequence of images, in which a spatial structure recognition unit constructs a spatial structure from a plurality of images and a storage device persists the source images and the constructed spatial structure, with the features of Choi's invention of convolutional-neural-network-based per-image object detection producing bounding boxes and class labels. A person of ordinary skill would have made this combination because Choi teaches that a convolutional neural network directly processes an input image and outputs bounding boxes of object detections, which are image-plane observations of the same form that Newcombe's projection pipeline already consumes, namely observations with image-plane locations paired to an image frame whose camera pose has been estimated, yielding the predictable result of a set of semantically labeled detections that Newcombe's pipeline can project into and persist with the reconstruction. Stauber discloses information on the detected object including a three-dimensional position and a size of the detected object estimated by projecting a two- dimensional position of the detected object onto the spatial structure using position and posture information of the image (para 33, “As shown in FIGS. 1A, 2, and 3, the application can: implement visual-inertial odometry, structure from motion, or similar techniques to derive depth information from a series of 2D frames recorded over a period of time and based on motion information (e.g., changes in position and orientation) of the mobile device tracked over this period of time; and then construct a 3D manifold or other virtual 3D representation of surfaces in the field of view of the 2D camera.”; also, para 34, “Upon detecting an object in a first 2D frame—in the set of consecutive frames recorded by the camera—the application can: calculate a 2D bounding box around this object in the first 2D frame; and project the 2D bounding box and the object depicted in two dimensions in the first 2D frame onto the 3D manifold generated from preceding 2D frames and concurrent motion of the mobile device and defined relative to the camera. The application can then implement ray casting techniques to: virtually project a first ray from the position of the camera into the 2D bounding box projected on the 3D manifold and determine whether the first ray intersects the projection of the object onto the 3D manifold. If so, the application can: calculate a distance from the position of the camera to a point on the 3D manifold at which the first ray intersects the projection of the object; and store the lateral, longitudinal, and depth positions of this intersection relative to the camera.”; also, para 35, “The application can then populate a 3D graph with a cluster of 3D points—defining lateral, longitudinal, and depth locations relative to the camera—wherein each point represents an intersection of a ray, virtually cast from the camera, on the object detected in the 2D frame and projected onto the 3D manifold. The application can subsequently calculate a 3D bounding box that encompasses this cluster of 3D points in the 3D graph and define this 3D bounding box relative to the camera.”; also, para 51, “The application can store additional characteristics in association with any object in the 3D constellation of objects, such as: characteristics of surfaces that define the object (e.g., the number of such surfaces); the structure of the object (e.g., represented as a 3D point cloud) or any other representation of this structure such as the total volume of the object; the orientation of the object; the dimensions of a 3D bounding box of the object; text, symbols, or visual patterns detected on the object; and/or colors present on the object; etc.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction as modified by Choi, in which the detected objects supplied by Choi are available as image-plane observations paired to image frames of estimated camera pose, with the features of Stauber's invention of projecting the two-dimensional bounding box of a detected object onto a three-dimensional manifold and ray casting from the tracked position and orientation of the camera to obtain the three-dimensional position and the three-dimensional bounding box of that object. A person of ordinary skill would have made this combination because Stauber teaches that projecting the two-dimensional bounding box onto a manifold generated from preceding frames and the concurrent tracked position and orientation of the device, and then casting rays from the position of the camera, delivers the lateral, longitudinal, and depth positions of the intersection together with a three-dimensional bounding box encompassing the resulting cluster of three-dimensional points, yielding the predictable result that each detection supplied by Choi acquires the three-dimensional position and size values that Newcombe's storage scheme then persists alongside the source images and the reconstruction. Forsblom discloses calculation device includes an image semantic and spatial structure summary unit configured to summarize data to be presented to a user (para 35, “if the density of marker objects in the particular bounding area 50 is too high, that is, where too much visual clutter would result by displaying each of the marker objects in that bounding area 50, they are not displayed and instead represented with an aggregate marker”; also, para 37, “The aggregate marker 56 may be implemented in a variety of different ways to convey representative information of its constituent marker objects. In its simplest form, the aggregate marker 56 is an icon including the number of marker objects within the particular bounding area 50”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a sequence of images, in which a spatial structure recognition unit constructs a spatial structure from a plurality of images, a fusion unit perspective projects between the image plane and the volumetric representation, and a storage device persists the source images and the constructed spatial structure, with the features of Choi's invention of convolutional-neural-network-based per-image object detection producing bounding boxes and class labels, Stauber's invention of projecting the two-dimensional bounding box of a detected object onto a three-dimensional manifold and ray casting from the tracked position and orientation of the camera to obtain the three-dimensional position and the three-dimensional bounding box of that object, and Forsblom's invention of clustering marker objects and representing each dense group with a single aggregate marker that conveys representative information about its constituent markers. A person of ordinary skill would have made this combination because Choi teaches that a convolutional neural network directly processes an input image and outputs bounding boxes of object detections, which are image-plane observations of the same form that Newcombe's projection pipeline already consumes, namely observations with image-plane locations paired to an image frame whose camera pose has been estimated; because Stauber teaches that projecting the two-dimensional bounding box onto a manifold generated from preceding frames and the concurrent tracked position and orientation of the device, and then casting rays from the position of the camera, delivers the lateral, longitudinal, and depth positions of the intersection together with a three-dimensional bounding box encompassing the resulting cluster of three-dimensional points, so that the detections supplied by Choi acquire the three-dimensional position and size values that Newcombe's storage scheme then persists; and because Forsblom teaches that displaying each of the marker objects individually results in too much visual clutter, that an aggregate marker conveys representative information of its constituent marker objects in the form of an icon including the number of marker objects within the particular bounding area, and that the aggregate marker so generated is itself displayed on the map in place of the individual markers it stands for, which under the broadest reasonable interpretation of a summary unit configured to summarize data to be presented to a user is a presentation of the summarized data to the person viewing that map, yielding the predictable result of a semantically annotated three-dimensional reconstruction in which each detected object carries a three-dimensional position and a size, a storage scheme that persists the source images, the reconstruction, and the per-detection metadata, and a condensed presentation of that per-detection data that a user can read at a glance, without modification to Newcombe's volumetric integration pipeline. Regarding claim 2, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, Choi further discloses the storage device stores information on a position and a size of the object (Choi: paragraph 29, line(s) 4-6, “a CNN directly processes an input image and outputs bounding boxes of object detections”; also, paragraph 17, line(s) 4-5, “estimate the location of such object in the image space”; also, paragraph 18, line(s) 20-21, “object category classification, subcategory classification, and bounding box regression”; also, paragraph 7, line(s) 3-4, “The system includes a memory and a processor in communication with the memory”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction with a memory that stores the reconstructed model and the data used by its processors, as modified by Stauber and Forsblom, with the features of Choi's invention of per-image object detection producing bounding-box position and size output. A person of ordinary skill would have made this combination because Choi teaches that the detection network performs bounding box regression to estimate the location of the object in the image space and that the system includes a memory and a processor in communication with the memory, and the bounding-box output so produced is the standard per-object position-and-size output of a detector, yielding the predictable result that those position and size values are persisted in the storage Newcombe already maintains, which supports downstream rendering and retrieval of the detected objects in the reconstructed scene. Regarding claim 3, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 2, wherein Newcombe further discloses the storage device stores information on a direction and a distance of the object from an imaging position of the image in which the object is imaged.(Newcombe: paragraph 27, line(s) 8-11, “each image element (i.e. pixel) comprises a depth value such as a length or distance from the camera to an object in the captured scene which gave rise to that image element”; also, paragraph 49, line(s) 5-7, “The 6DOF pose estimate indicates the location and orientation of the depth camera 302”). Regarding claim 4, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, wherein Newcombe further discloses the storage device stores a feature of the image as the source of constructing the spatial structure (Newcombe: paragraph 62, line(s) 10-12, “The averaged signed distance function values can then be stored 430 at the current voxel”; also, paragraph 63, line(s) 1-4, “two values can be stored at each voxel. A weighted sum of the signed distance function values can be calculated and stored, and also a sum of the weights calculated and stored”). Regarding claim 5, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 2, wherein Newcombe further discloses the spatial structure recognition unit estimates a three-dimensional position and a posture of the image based on at least one of information on an imaging position of the image and a value estimated in a processing step of spatial structure construction (Newcombe: paragraph 49, line(s) “receiving 402 from the mobile environment capture device 300 a depth image 314 and the 6DOF pose estimate of the depth camera 302 when capturing that depth image. The 6DOF pose estimate indicates the location and orientation of the depth camera 302”; also, paragraph 41, line(s) 1-4, “The frame alignment engine 318 of the real-time tracker is arranged to align pairs of depth image frames, or a depth image frame and an estimate of a depth image frame from the dense 3D model”), and the image semantic and spatial structure fusion unit estimates a three-dimensional position and a size of the object on the spatial structure based on a two-dimensional position of the detected object in the image, the estimated three-dimensional position and the posture of the image, and the spatial structure, and stores the estimated three-dimensional position and the size in the storage device (Newcombe: paragraph 53, line(s) 12-14, “the real-world coordinate for the current voxel associated with the thread is perspective projected through the depth camera's projection”; also, paragraph 56, line(s) 2-5, “a factor relating to the distance between the voxel and a point in the environment at the corresponding location to the voxel from the camera's perspective is determined”; also, paragraph 62, line(s) 10-12, “The averaged signed distance function values can then be stored 430 at the current voxel”). Regarding claim 7, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, wherein Forsblom further discloses the image semantic and spatial structure summary unit clusters a plurality of the objects according to similarity in information on the objects, and places a group icon indicating a region including all of the objects included in a generated cluster on the spatial structure (paragraph 35, line(s) 9-14, “if the density of marker objects in the particular bounding area 50 is too high, that is, where too much visual clutter would result by displaying each of the marker objects in that bounding area 50, they are not displayed and instead represented with an aggregate marker”; also, paragraph 37, line(s) 1-5, “The aggregate marker 56 may be implemented in a variety of different ways to convey representative information of its constituent marker objects. In its simplest form, the aggregate marker 56 is an icon including the number of marker objects within the particular bounding area 50”; also, paragraph 37, line(s) 11-13, “the aggregate markers 56 may be positioned in a central region of the bounding area 50”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction, as modified by Choi and Stauber, with the features of Forsblom's invention of selective map marker aggregation. A person of ordinary skill would have made this combination because Forsblom teaches that displaying each of the marker objects individually results in too much visual clutter, that dense groups are instead represented with a single aggregate marker conveying representative information of its constituent markers, and that the aggregate markers may be positioned in a central region of the bounding area, and the same visual-clutter problem arises once the detections supplied by Choi are located on Newcombe's reconstruction by the projection taught by Stauber, yielding the predictable result of reduced clutter and improved navigability of the reconstructed three-dimensional scene. Regarding claim 8, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, Forsblom further discloses wherein the image semantic and spatial structure summary unit displays an icon label that varies in a display mode depending on types of all objects included in a cluster (paragraph 40, line(s) 1-11, “several embodiments of the present disclosure contemplate additional information being provided in the aggregate markers 56 in the form of icons, colors, further text-based indicators, and so forth. For example, the first aggregate marker 56a may include an icon 66 that is representative of some demographic information culled from the profiles 32 of those contacts located within the northeast bounding area 50a. The second aggregate marker 56b indicates that the southeast bounding area 50d is comprised of 75% males, another demographic data point retrieved from the profiles 32.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction, as combined with Choi's invention of CNN-based per-image object detection, with the features of Forsblom's invention of type-dependent aggregate marker display, because Forsblom teaches that aggregate markers can convey representative information about the composition of the clustered objects through varying icons, colors, and text, and a person of ordinary skill in the art would recognize that when Choi's object detections (each carrying a category and subcategory label) are clustered on Newcombe's reconstructed spatial structure, displaying the cluster icon in a manner that reflects the class composition of the constituent objects provides the user with at-a-glance information about what types of objects are present in each region of the 3D scene. One of ordinary skill in the art would recognize that Forsblom's technique of varying the aggregate marker's display mode based on the types of clustered objects applies directly to Choi's classified detections, yielding the predictable result of category-aware cluster visualization on the spatial structure. Regarding claim 13, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, wherein Newcombe further discloses the storage device stores (Newcombe: para 89, “The memory 816 can also provide a data store 830, which can be used to provide storage for data used by the processors 802 when performing the 3D modeling techniques”), and Choi further discloses a reliability value representing reliability of an image recognition result for the detected object (para 34, “The subcategory cony layer 212 outputs heat map 214 for each scale, where each location in the heat map 214 indicates the confidence of an object in the corresponding location, scale, and subcategory”; also, para 52, “Finally, the detection network 300 terminates at three output layers 322. The first output layer applies a softmax function directly on the output of the "subcategory FC" layer for subcategory classification. The other two output layers operate on the RoI feature vector, and apply FC layers for object class classification and bounding box regression.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction with a memory providing a data store for the data used by its processors, as modified by Stauber and Forsblom, with the features of Choi's invention of a detection network that outputs a confidence measure and a softmax probability for the object class of each detected region of interest. A person of ordinary skill would have made this combination because Choi teaches that each location in the heat map indicates the confidence of an object in the corresponding location, scale, and subcategory and that the detection network applies a softmax function for object class classification of each detected region of interest, and a confidence value is of use to a downstream consumer only if it is retained with the detection it qualifies, yielding the predictable result that the reliability measure accompanying each detection is written to the same data store in which Newcombe already persists the reconstruction and the data used by its processors. Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) as modified by Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077) and Forsblom et al. (U.S. Pub. No. 20130044137), further in view of Shen et al. (U.S. Pub. No. 2020/0175326). Regarding claim 6, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 5, image semantic and spatial structure fusion unit estimates the three-dimensional position and the size of the object on the spatial structure, the object being detected within a predetermined range from a perpendicular line to a center position of the image. However, in a similar field of endeavor Shen discloses wherein the image semantic and spatial structure fusion unit estimates the three-dimensional position and the size of the object on the spatial structure, the object being detected within a predetermined range from a perpendicular line to a center position of the image (paragraph 13, line(s) 14-16, “a portion associated with a horizon line or vanishing line may tend to depict other vehicles on a road”; also, paragraph 18, line(s) 1-4, “a horizon line, or other field of view, may be assigned as a center third portion of an image. For example, along a vertical direction a center third of the image may be cropped”; also, paragraph 35, line(s) 5-7, “system 100 can determine the priority field of vision in a naive fashion, by taking the center of the image and classifying it as a priority field of vision”; also, paragraph 62, line(s) 5-8, “a rectangle from the image may be cropped. The rectangle may optionally extend a first threshold distance above the y-coordinate and a second threshold distance below the y-coordinate”; also, paragraph 45, line(s) 9-12, “an accuracy associated with detecting objects, assigning bounding boxes or other location information, and so on, may be greater for the crop”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction, as modified by Choi, Stauber, and Forsblom, with the features of Shen's invention of center-of-image detection filtering. A person of ordinary skill would have made this combination because Shen teaches that an accuracy associated with detecting objects and assigning bounding boxes may be greater for the crop and that the priority field of vision can be determined by taking the center of the image and classifying it as a priority field of vision, and restricting the detections admitted for fusion to those within a predetermined range from the center of the image reduces false positives arising from peripheral lens distortion, yielding the predictable result of more reliable three-dimensional position and size estimates from the fusion step without altering the projection mathematics of Newcombe or the detection architecture of Choi. Claim(s) 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) as modified by Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077) and Forsblom et al. (U.S. Pub. No. 20130044137), further in view of Hoffmann et al. (U.S. Pub. No. 2019/0272650). Regarding claim 9, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, further comprising: context utilization query generation unit configured to acquire a viewpoint position and a viewing direction of space information to be displayed and generate a search query including a three-dimensional position of a gaze point of a user as a condition based on the acquired position and posture information; and an image search unit configured to acquire an image from the storage device using the generated search query. However, in a similar field of endeavor, Hoffmann discloses a context utilization query generation unit configured to acquire a viewpoint position and a viewing direction of space information to be displayed and generate a search query including a three-dimensional position of a gaze point of a user as a condition based on the acquired position and posture information (paragraph 4, line(s) 1-4, “The present invention relates to a method and an apparatus for gaze endpoint determination, in particular for determining a gaze endpoint of a subject on a three-dimensional object in space”; also, paragraph 26, line(s) 1-7, “a calculating unit for calculating the gaze endpoint based on the gaze direction, the eye tracker position and the 3D scene structure representation, and/or for determining the object in the 3D scene the subject is gazing at based on the gaze direction, the eye tracker position and the 3D scene structure representation”; also, paragraph 29, line(s) 1-4, “The intersection of gaze direction with the 3D representation gives a geometrical approach for calculating the location where the gaze “hits” or intersects the 3D structure and therefore delivers the real gaze endpoint”); and an image search unit configured to acquire an image from the storage device using the generated search query (paragraph 36, line(s) 1-2, “a scene camera adapted to acquire one or more images of the scene from an arbitrary viewpoint”; also, paragraph 37, line(s) 1-3, “a module for mapping a 3D gaze endpoint onto the image plane of the scene image taken by the scene camera”; also, paragraph 38, line(s) 1-4, “In this way not only the 3D gaze endpoint on the 3D structure is determined, but there can be determined the corresponding location on any scene image as taken by a scene camera”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction, as modified by Choi, Stauber, and Forsblom, with the features of Hoffmann's invention of gaze-endpoint-based mapping between a user's viewpoint and corresponding images of the scene. A person of ordinary skill would have made this combination because Hoffmann teaches that the intersection of gaze direction with the three-dimensional representation delivers the real gaze endpoint and that a module maps a three-dimensional gaze endpoint onto the image plane of a scene image taken by a scene camera, and the combined system already persists a set of source images together with the three-dimensional model derived from them and the per-image camera pose that such a mapping consumes, yielding the predictable result of a principled gaze-point-based query against those stored images that returns the ones relevant to the region being viewed, without alteration of Newcombe's underlying storage scheme. Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) as modified by Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), Forsblom et al. (U.S. Pub. No. 20130044137) and Hoffmann et al. (U.S. Pub. No. 2019/0272650), further in view of Spiegel et al. (U.S. Pub. No. 2020/0401617). Regarding claim 10, Newcombe as modified by Choi, Stauber, Forsblom, and Hoffmann discloses the space visualization system according to claim 9, context utilization query generation unit calculates, using the space information viewed from the viewpoint position as the image, an image feature of the image, and generates the search query including the calculated image feature as a condition. However, in a similar field of endeavor, Spiegel discloses wherein the context utilization query generation unit calculates, using the viewing direction of space information from the viewpoint position as the image, an image feature of the image, and generates the search query including the calculated image feature as a condition (paragraph 177, line(s) 14-23, The integrated neural network 406 may include a feature extractor 407. The query image 405 may be provided to the feature extractor 407 which generates a query feature map provided to a correspondence matcher 408. The correspondence matcher 408 may be implemented within a neural network 406 in order to identify correspondence between a feature map generated by the feature extractor 407 from a query image 405 to a feature map generated from a reference image”; also, paragraph 74, line(s) -11, “feature extraction starts from an initial set of measured data and builds derived values (features) intended to be informative and non-redundant, facilitating the subsequent learning and generalization steps, and in some cases leading to better human interpretations. Feature extraction is related to dimensionality reduction. When the input data to an algorithm is too large to be processed and it is suspected to be redundant (e.g. the same measurement in both feet and meters, or the repetitiveness of images presented as pixels), then it can be transformed into a reduced set of features (also named a feature vector)”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction, as modified by Choi, Stauber, Forsblom, and Hoffmann, with the features of Spiegel's invention of query-image feature-map extraction and correspondence matching against a reference-image feature-map database. A person of ordinary skill would have made this combination because Spiegel teaches that a feature extractor generates a query feature map from a query image and that a correspondence matcher identifies correspondence between that feature map and a feature map generated from a reference image, and a query resting on gaze-point proximity alone selects source images by geometry and cannot discriminate among images captured from similar camera positions but depicting differing visual content, yielding the predictable result of an appearance-based refinement of that query which consumes the per-image content Newcombe already persists. Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), further in view of Stauber et al. (U.S. Pub. No. 20210074077). Regarding claim 11, Newcombe discloses a spatial visualization method executed by a computer, the computer including a calculation device that executes predetermined calculation processing, and a storage device accessible to the calculation device, the spatial visualization method comprising (paragraph 44, line(s) 1-3, “Reference is now made to FIG. 4, which illustrates a flowchart of a parallelizable process for generating a 3D environment model; also, paragraph 91, line(s) 1-5, “The methods described herein may be performed by software in machine readable form on a tangible storage medium e.g. in the form of a computer program comprising computer program code means adapted to perform all the steps of any of the methods described herein”): a spatial structure recognition step of the calculation device constructing a spatial structure based on a plurality of images (paragraph 22, line(s) 9-12, “images captured by the mobile depth camera 102 are used to form and build up a dense 3D model of the environment as the person moves about the room”; also, paragraph 24, line(s) 1-4, “The real-time camera tracking system 112 provides input to the dense 3D modeling system, in order to allow individual depth images to be built up into an overall 3D models.”); image semantic and spatial structure fusion step of the calculation device estimating a spatial position of the detected object on the spatial structure (paragraph 53, line(s) 12-15, “the real-world coordinate for the current voxel associated with the thread is perspective projected through the depth camera's projection, and can take into account the intrinsic parameters of the camera”; also, paragraph 62, line(s) 10-12, “The averaged signed distance function values can then be stored 430 at the current voxel”); and a step of the calculation device storing, in the storage device, information on the image as a source of constructing the spatial structure, the constructed spatial structure (paragraph 7, line(s) 2-4, “a 3D model of a real-world environment is generated in a 3D volume made up of voxels stored on a memory device.”; also, paragraph 89, line(s) 12-15, “The memory 816 can also provide a data store 830, which can be used to provide storage for data used by the processors 802 when performing the 3D modeling techniques”), image semantic recognition step of the calculation device detecting an object included in each of the plurality of images, and information on the detected object including a three-dimensional position and a size of the detected object estimated by projecting a two-dimensional position of the detected object onto the spatial structure using position and posture information of the image. However, in a similar field of endeavor, Choi discloses an image semantic recognition step of the calculation device detecting an object included in each of the plurality of images (paragraph 29, line(s) 4-6, “a CNN directly processes an input image and outputs bounding boxes of object detections”; also, paragraph 30, line(s) 2-3, “region proposal techniques are utilized to generate a set of object candidates from each image”; also, paragraph 17, line(s) 1-5, “the present invention proposes a method and system for solving visual object detection given an image. The goal is to identify the object of interest (e.g., car, pedestrian, etc.) and information on the detected object (paragraph 7, line(s) 3-10, “The system includes a memory and a processor in communication with the memory, wherein the processor is configured to generate object region proposals from an image by a region proposal network (RPN) which utilizes subcategory information, and classify and refine the object region proposals by an object detection network (ODN) that simultaneously performs object category classification, subcategory classification, and bounding box regression”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of a computer-executed method of constructing a spatial structure from a plurality of images, estimating a spatial position on that structure by perspective projection, and storing the source images and the constructed structure, with the features of Choi's invention of convolutional-neural-network-based per-image object detection producing bounding boxes and class labels. A person of ordinary skill would have made this combination because Choi teaches that a convolutional neural network directly processes an input image and outputs bounding boxes of object detections, which are image-plane observations of the same geometric form that Newcombe's projection step already consumes, yielding the predictable result of a set of semantically labeled detections that the method can project onto the volumetric representation and store with the reconstruction. Stauber discloses information on the detected object including a three-dimensional position and a size of the detected object estimated by projecting a two-dimensional position of the detected object onto the spatial structure using position and posture information of the image (para 33, “As shown in FIGS. 1A, 2, and 3, the application can: implement visual-inertial odometry, structure from motion, or similar techniques to derive depth information from a series of 2D frames recorded over a period of time and based on motion information (e.g., changes in position and orientation) of the mobile device tracked over this period of time; and then construct a 3D manifold or other virtual 3D representation of surfaces in the field of view of the 2D camera.”; also, para 34, “Upon detecting an object in a first 2D frame—in the set of consecutive frames recorded by the camera—the application can: calculate a 2D bounding box around this object in the first 2D frame; and project the 2D bounding box and the object depicted in two dimensions in the first 2D frame onto the 3D manifold generated from preceding 2D frames and concurrent motion of the mobile device and defined relative to the camera. The application can then implement ray casting techniques to: virtually project a first ray from the position of the camera into the 2D bounding box projected on the 3D manifold and determine whether the first ray intersects the projection of the object onto the 3D manifold. If so, the application can: calculate a distance from the position of the camera to a point on the 3D manifold at which the first ray intersects the projection of the object; and store the lateral, longitudinal, and depth positions of this intersection relative to the camera.”; also, para 35, “he application can then populate a 3D graph with a cluster of 3D points—defining lateral, longitudinal, and depth locations relative to the camera—wherein each point represents an intersection of a ray, virtually cast from the camera, on the object detected in the 2D frame and projected onto the 3D manifold. The application can subsequently calculate a 3D bounding box that encompasses this cluster of 3D points in the 3D graph and define this 3D bounding box relative to the camera.”; also, para 51, “The application can store additional characteristics in association with any object in the 3D constellation of objects, such as: characteristics of surfaces that define the object (e.g., the number of such surfaces); the structure of the object (e.g., represented as a 3D point cloud) or any other representation of this structure such as the total volume of the object; the orientation of the object; the dimensions of a 3D bounding box of the object; text, symbols, or visual patterns detected on the object; and/or colors present on the object; etc.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention as modified by Choi, in which the detected objects supplied by Choi are available as image-plane observations paired to image frames of estimated camera pose, with the features of Stauber's invention of projecting the two-dimensional bounding box of a detected object onto a three-dimensional manifold and ray casting from the tracked position and orientation of the camera to obtain the three-dimensional position and the three-dimensional bounding box of that object. A person of ordinary skill would have made this combination because Stauber teaches that projecting the two-dimensional bounding box onto a manifold generated from preceding frames and the concurrent tracked position and orientation of the device, and then casting rays from the position of the camera, delivers the lateral, longitudinal, and depth positions of the intersection together with a three-dimensional bounding box encompassing the resulting cluster of three-dimensional points, yielding the predictable result of per-object three-dimensional position and size values stored alongside the source images and the constructed spatial structure that Newcombe's storing step already persists. Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), and Forsblom et al. (U.S. Pub. No. 20130044137), further in view of Abeywardena (U.S. Pub. No. 20210133997). Regarding claim 12, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, image semantic and spatial structure fusion unit estimates the three-dimensional position of the detected object by extending a straight line on an optical axis from center coordinates of the detected object in the image and obtaining three-dimensional coordinates at which the straight line collides with the spatial structure. However, in a similar field of endeavor, Abeywardena discloses wherein the image semantic and spatial structure fusion unit estimates the three-dimensional position of the detected object by extending a straight line on an optical axis from center coordinates of the detected object in the image and obtaining three-dimensional coordinates at which the straight line collides with the spatial structure (para 55, “As shown, to determine a position of the structure 802 within the terrain model 808, a ray 810 is projected from the precise position of the camera 102 in the autonomous vehicle 200, along a vector that extends from the camera optical center through a center pixel of the annotated location of the structure 812 associated with the image of the structure 806. Then, a terrain model intersection point 814 is found based on where the ray 810 intersects the terrain model 808. The geographical location associated with the terrain model intersection point 814, which is a part of the terrain model 808, can then be used as the geographical location of the structure 802.”; also, para 54, “While the illustration in FIG. 8 is a two-dimensional cross-section, one will appreciate that in practice, the terrain model 808 is likely three-dimensional.”; also, para 65, “At block 512, the geographic location determination engine 318 determines geographical locations of the structures in the new aerial image based on the model-generated annotations and the pose information. In some embodiments, the determination of the geographical locations is similar to the determination made in block 416.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a plurality of images as modified by Choi, Stauber, and Forsblom with the features of Abeywardena's invention of projecting a single ray from the camera optical center through the center pixel of a detected object's annotated location in the image and taking the point at which that ray intersects the terrain model as the object's three-dimensional location. A person of ordinary skill would have made this combination because Abeywardena teaches that once the pose of the camera and the terrain model are determined, ray tracing or ray casting may be used to determine geographic locations of annotations within the terrain model, and that the single ray through the center pixel resolves to a single terrain model intersection point that can then be used as the location of the structure, yielding the predictable result of one determinate three-dimensional coordinate per detected object rather than the cluster of intersections produced by sweeping the entire detection region. Claim(s) 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), and Forsblom et al. (U.S. Pub. No. 20130044137), further in view of Nagasawa et al. (U.S. Pub. No. 20200267369). Regarding claim 14, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, display unit configured to display the spatial structure on a three-dimensional viewer and visualize the spatial structure at a viewpoint designated by a user. However, in a similar field of endeavor, Nagasawa discloses further comprising a display unit configured to display the spatial structure on a three-dimensional viewer and visualize the spatial structure at a viewpoint designated by a user (para 50, “The display control unit 24 renders the stereoscopic image in the three-dimensional space on the basis of the point cloud data received by the point cloud data receiving unit 23 and causes the display 203 to display the stereoscopic image.”; also, para 27, “The display state designating unit 21 of the receiving apparatus 200 designates the display state of the stereoscopic shape to be displayed on the display 203. The display state of the stereoscopic shape is the viewpoint and the direction of the line of sight from the viewpoint that are variable in the three-dimensional space and refers to the viewpoint designated by the user and the direction of the line of sight. Various methods can be used as a method of designating the viewpoint and the direction of the line of sight direction by the user.”; also, para 28, “For example, the user can arbitrarily designate a desired viewpoint and a direction of the line of sight by operating the operating unit 202 in a state in which the stereoscopic image based on the point cloud data is displayed on the display 203. In this case, the user can perform designation of causing the viewpoint to be moved to a desired position or causing the line of sight to a desired direction with reference to the display state of the stereoscopic image being currently displayed by operating the operating unit 202. The operating unit 202 is, for example, a mouse, a keyboard, a touch panel, a controller, a changeover switch, or the like.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a plurality of images as modified by Choi, Stauber, and Forsblom with the features of Nagasawa's invention of displaying a stereoscopic shape built from point cloud data on a display and permitting the user to designate the viewpoint and the direction of the line of sight from which that shape is rendered. A person of ordinary skill would have made this combination because Nagasawa teaches that the user can arbitrarily designate a desired viewpoint and a direction of the line of sight by operating a mouse, a keyboard, a touch panel, or a controller while the stereoscopic image is displayed, and Newcombe's rendering process already accepts a virtual camera pose of exactly the form Nagasawa's operating unit supplies, yielding the predictable result that a user can inspect any region of the reconstructed spatial structure from a viewpoint of the user's own choosing rather than only from the poses at which the source images happened to be captured. Claim(s) 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), Forsblom et al. (U.S. Pub. No. 20130044137), and Nagasawa et al. (U.S. Pub. No. 20200267369), further in view of Zhao et al. (U.S. Pub. No. 2021035810). Regarding claim 15, Newcombe as modified by Choi, Stauber, Forsblom, and Nagasawa discloses the space visualization system according to claim 14, display unit arranges an icon of the detected object on the spatial structure based on the three-dimensional position of the detected object. However, in a similar field of endeavor, Zhao discloses wherein the display unit arranges an icon of the detected object on the spatial structure based on the three-dimensional position of the detected object (para 152, “The converter 107C plots the obtained three-dimensional coordinates as the location of the deteriorated portion on the target three-dimensional model. This plotting can be achieved as association processing in the information processing (FIG. 12 described below). In addition, the converter 107C sets a deteriorated portion image for highlighting the deteriorated portion on the target three-dimensional model with a predetermined color or an icon to be shown at the time of visualization.”; also, para 151, “The perspective conversion matrix P is obtained beforehand from a separate SFM processing. The three-dimensional coordinates obtained through conversion represents the location of the deteriorated portion in an area including the target three-dimensional model.”; also, para 175, “Further, the visualizer 108 generates a screen in which the deteriorated portion is plotted on the target three-dimensional model based on conversion information and is visualized”; also, para 176, “Further, the three-dimensional coordinates (X1,Y1,Z1) of a deteriorated portion 1313 are plotted on a surface of the target three-dimensional model 1311.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a plurality of images as modified by Choi, Stauber, Forsblom, and Nagasawa with the features of Zhao's invention of plotting the three-dimensional coordinates of a detected portion onto a target three-dimensional model and setting an icon to mark that portion at the time of visualization. A person of ordinary skill would have made this combination because Zhao teaches that the converter plots the obtained three-dimensional coordinates as the location of the detected portion on the target three-dimensional model and sets an icon for highlighting that portion to be shown at the time of visualization, and the combined system already computes exactly such three-dimensional coordinates for each detected object, yielding the predictable result that a user viewing the reconstructed spatial structure sees each detected object marked at the place in the structure where it actually sits rather than having to correlate a separate list of detections against the geometry by hand. Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), Forsblom et al. (U.S. Pub. No. 20130044137), Nagasawa et al. (U.S. Pub. No. 20200267369), and Zhao et al. (U.S. Pub. No. 2021035810), further in view of Kang et al. (U.S. Pub. No. 20210295621). Regarding claim 16, Newcombe as modified by Choi, Stauber, Forsblom, Nagasawa, and Zhao discloses the space visualization system according to claim 15, user selects the icon of the detected object, the display unit acquires detailed information of an original image from which the detected object is detected from the storage device and displays the image in a pop-up window. However, in a similar field of endeavor, Knag discloses wherein, when a user selects the icon of the detected object, the display unit acquires detailed information of an original image from which the detected object is detected from the storage device and displays the image in a pop-up window (para 308, “Meanwhile, FIG. 12B illustrates a case where when the driving-related event icon 1206 is selected, specific information of the driving-related event corresponding to the selected icon is displayed as a pop-up window. According to the present invention, the user terminal device 400 may display an image of an event corresponding to the selected event driving-related event as a pop-up window. Here, the image may be a motion image generated by combining at least two images instead of one image, or may be an image configured by only one image.”; also, para 140, “In addition, the vehicle driving support function unit 180 may detect a danger of a collision a vehicle in front of the vehicle and determine a front collision warning system (FCWS) is required for the driver, based on the driving image captured by the imaging unit 110.”; also, para 148, “In addition, the event recording data may include an image and/or a recorded sound captured for a period from before a predetermined time to after a predetermined time based on the occurrence of the driving-related event”; also, para 168, “Also, the driving recording data stored in the driving recording data storage unit 322 may include the event recording data, driving time data, driving route data, and driving distance data.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a plurality of images as modified by Choi, Stauber, Forsblom, Nagasawa, and Zhao with the features of Kang's invention of responding to a user's selection of an icon by retrieving from a storage unit, and displaying in a pop-up window, the captured image in which the corresponding object was detected. A person of ordinary skill would have made this combination because Kang teaches that when the icon is selected the terminal device may display an image of the corresponding event as a pop-up window, that this image may be an image configured by only one image, and that the image so displayed is the captured image held in the recording data storage unit from which the vehicle was detected, and the combination of Newcombe, Choi, Stauber, Forsblom, Nagasawa, and Zhao already places an icon for each detected object on the spatial structure at the three-dimensional position of that object while separately persisting the source image in which the object was detected, so that Kang supplies only the response to the selection of that icon, yielding the predictable result that a user who selects an icon is shown the source image itself for confirmation instead of having to locate that image among the stored source images by hand. Claim(s) 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), and Forsblom et al. (U.S. Pub. No. 20130044137), further in view of Tysowski (U.S. Pub. No. 20120147216) and Hoffmann et al. (U.S. Pub. No. 2019/0272650). Regarding claim 17, Newcombe as modified by Choi, Stauber, and Forsblom, discloses the space visualization system according to claim 1, image search unit configured to search for objects included in a predetermined search distance range from a three-dimensional position of a gaze point of a user. However, in a similar field of endeavor, Tysowski discloses further comprising an image search unit configured to search for objects included in a predetermined search distance range (para 68, “Other options for overlaying images could indicate, for example, that the user wishes to view images that are within a set radius of a specific point. Such as the current location of the user. At this point, the user on the mobile device could choose images that are within the geographic radius of the selected location. Thus, a search routine on the mobile device or at the network server may need to look in multiple subfolders in order to determine whether or not images exist in those subfolders for the chosen radius.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction as modified by Choi, Stauber, and Forsblom, in which the source images are persisted and each detected object is located on the constructed spatial structure at a three-dimensional position, with the features of Tysowski's invention of an image search routine that returns the stored images lying within a chosen radius of a specified point. A person of ordinary skill would have made this combination because Tysowski teaches that a search routine on the mobile device or at the network server looks through the stored subfolders to determine whether images exist within the chosen radius of a selected location, a search distance range being a scalar distance measured about a point that the combination can evaluate in the same three-dimensional coordinate system in which it already holds each detected object together with the source image in which that object was detected, yielding the predictable result of a retrieval limited to the objects whose associated images fall within the chosen distance of the specified point rather than the whole reconstruction. Hoffman discloses from a three-dimensional position of a gaze point of a user (para 22, “a calculating unit for calculating the gaze endpoint based on the gaze direction, the eye tracker position and the 3D scene structure representation, and/or for determining the object in the 3D scene the subject is gazing at based on the gaze direction, the eye tracker position and the 3D scene structure representation”; also, para 29, “The intersection of gaze direction with the 3D representation gives a geometrical approach for calculating the location where the gaze "hits" or intersects the 3D structure and therefore delivers the real gaze endpoint.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention as modified by Choi, Stauber, Forsblom, and Tysowski, in which an image search routine returns the objects whose associated images fall within a chosen distance of a specified point, with the features of Hoffmann's invention of intersecting a user's gaze direction with a three-dimensional scene structure representation to deliver a three-dimensional gaze endpoint. A person of ordinary skill would have made this combination because Hoffmann teaches that the intersection of the gaze direction with the three-dimensional representation delivers the real gaze endpoint on the three-dimensional structure, and the search routine supplied by Tysowski requires a specified point about which the chosen distance is measured, so that using the user's three-dimensional gaze endpoint as that point directs the search to the region of the reconstruction the user is looking at, yielding the predictable result that the objects lying within the chosen distance of the user's three-dimensional gaze point are returned. Claim(s) 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), and Forsblom et al. (U.S. Pub. No. 20130044137), further in view of Ohtomo et al. (U.S. Pub. No. 20130135440). Regarding claim 18, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, plurality of images are acquired by an unmanned aerial vehicle including a global navigation satellite system for acquiring position information and a gyro sensor for acquiring posture information. However, in a similar field of endeavor, Ohtomo discloses wherein the plurality of images are acquired by an unmanned aerial vehicle including a global navigation satellite system for acquiring position information and a gyro sensor for acquiring posture information (para 42, “The flying object 1 is, e.g., a helicopter as a small flying object for making autonomous flight. The helicopter 1 is operated by remote control from the base control device 2, or the flight plan is set in a control device (as described later) of the helicopter 1 from the base control device 2, and autonomous flight is carried out according to the flight plan.”; also, para 44, “A GPS device 12, a control device, and at least one image pickup device 13 are installed on the helicopter body 3 of the helicopter 1. In the GPS device 12 to be used in the present embodiment, real time kinematic GPS (i.e. RTK-GPS) is used, for one example.”; also, para 44, “The GPS device 12 is arranged so as to measure a reference position, e.g. a mechanical center, of the helicopter 1. The GPS device 12 also measures absolute three-dimensional coordinates of the reference position, and the measured value indicates ground coordinate system and altitude as obtained from geocentric coordinate system (absolute coordinate system).”; also, para 49, “The image taken by the image pickup device 13 is associated with time for taking the image, with geocentric coordinates (three-dimensional coordinates) measured by the GPS device 12 and is stored in a storage unit 18 as described later.”; also, para 50, “FIG. 2 shows a control device 16 provided in the helicopter body 3. The control device 16 is mainly constituted of an arithmetic control unit 17, a storage unit 18, a communication unit 19, an image pickup controller 21, a motor controller 22, a gyro unit 23, and a power supply unit 24.”; also, para 61, “The arithmetic control unit 17 also controls the helicopter 1 in horizontal position during the flight and maintains the helicopter 1 in standstill flying (hovering) condition at the predetermined position by controlling the first motor 8, the second motor 9, the third motor 10, and the fourth motor 11 via the motor controller 22 according to the flying posture control program and based on the signals from the gyro unit 23.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a plurality of images as modified by Choi, Stauber, and Forsblom with the features of Ohtomo's invention of an unmanned aerial vehicle carrying an image pickup device, a global positioning system device that measures the absolute three-dimensional coordinates of the aircraft, and a gyro unit whose signals drive the flying posture control program. A person of ordinary skill would have made this combination because Ohtomo teaches that the image taken by the image pickup device is associated with the three-dimensional coordinates measured by the global positioning system device and stored in a storage unit, and that the posture of the aircraft is controlled based on the signals from the gyro unit, a gyro unit whose signals are the input to a flying posture control program necessarily senses the posture that the program controls, so that those signals are posture information under the broadest reasonable interpretation of that term, and the combined system requires exactly this per-image position and posture information in order to project each detection onto the spatial structure, yielding the predictable result that the plurality of images arrives already tagged with the position and posture data the projection consumes, and that a space too large or too hazardous to walk through with a hand-held capture device can be surveyed from the air. Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), and Forsblom et al. (U.S. Pub. No. 20130044137), further in view of Watanabe et al. (U.S. Pub. No. 20160217158). Regarding claim 19, Newcombe as modified by Choi, Stauber, and Forsblom discloses the space visualization system according to claim 1, storage device stores an image feature obtained by digitizing a visual feature in the image, the image feature being given by fixed-length vector data for use in similar image search. However, in a similar field of endeavor, Watanabe discloses wherein the storage device stores an image feature obtained by digitizing a visual feature in the image, the image feature being given by fixed-length vector data for use in similar image search (para 44, “The image ID field 201 holds identification numbers of each image data piece. The image data field 202 is a field in which image data is held in a binary form and is used when a user confirms a search result. The image feature value field 203 holds image feature value data. An image feature value means fixed-length numerical vector data obtained by digitizing features such as color and a shape which an image itself has.”; also, para 47, “Normalization for correcting bias of distribution for each pattern is made and the histogram is stored as fixed-length vector data of approximately several hundreds of dimensions which the system can readily handle and which is obtained by compressing dimensions of obtained vectors of several thousands of dimensions by principal component analysis and the like. As the vector data obtained as described above has close values between seemingly similar images, it can be used for the similar image search.”; also, para 35, “The image database 108 means a database that stores image data, its feature value and a tag. The feature value means data required for a search based upon apparent similarity of an image.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a plurality of images as modified by Choi, Stauber, and Forsblom with the features of Watanabe's invention of an image database holding, for each image, a fixed-length numerical vector obtained by digitizing the color and shape features of that image for use in similar image search. A person of ordinary skill would have made this combination because Watanabe teaches that the histogram is stored as fixed-length vector data of approximately several hundreds of dimensions which the system can readily handle, and that such vector data has close values between seemingly similar images and can therefore be used for the similar image search, and Newcombe already persists the source images from which the spatial structure was constructed without any means of comparing them by appearance, yielding the predictable result that the stored source images become retrievable by visual similarity at a uniform and readily handled per-image cost. Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Newcombe et al. (U.S. Pub. No. 2012/0194516) in view of Choi et al. (U.S. Pub. No. 2017/0124415 ), Stauber et al. (U.S. Pub. No. 20210074077), and Forsblom et al. (U.S. Pub. No. 20130044137), further in view of Kutliroff et al. (U.S. Pub. No. 20170243352). Regarding claim 20, Newcombe as modified by Choi, Stauber, and Forsblom, discloses the space visualization system according to claim 1, storage device stores a direction of a straight line connecting an imaging device and the detected object in the three-dimensional space. However, in a similar field of endeavor, Kutliroff discloses wherein the storage device stores a direction of a straight line connecting an imaging device and the detected object in the three-dimensional space (para 63, “In order to represent this point in the object boundary set, two 3-element vectors are stored: the 3D (x,y,z) position of the point in the global coordinate system, and the vector representing the ray extending from the camera's position to that point (which is referred to herein as the "camera ray").”; also, para 63, “For each pixel, the associated 3D position of the 2D pixel is computed, by sampling the associated depth map to obtain the associated depth pixel and projecting that depth pixel to a point in 3D space, at operation 1412. A ray is then generated which extends from the camera to the location of the projected point in 3D space at operation 1414.”; also, para 61, “The object detection circuit 134 may be configured to process the RGB image, and in some embodiments the associated depth map as well, to generate a list of any objects of interest recognized in the image. A label may be attached to each of the recognized objects and a 2D bounding box is generated which contains the object. Additionally, a 3D location of the center of the 2D bounding box is computed.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Newcombe's invention of three-dimensional environment reconstruction from a plurality of images as modified by Choi, Stauber, and Forsblom with the features of Kutliroff's invention of storing, for each point belonging to a recognized object, a vector representing the ray extending from the camera's position to that point. A person of ordinary skill would have made this combination because Kutliroff teaches that two 3-element vectors are stored for such a point, the three-dimensional position of the point in the global coordinate system and the camera ray from the camera's position to that point, and the combined system already computes the ray from the camera to each detected object in the course of projecting that object onto the spatial structure but discards its direction after taking the intersection, yielding the predictable result that the line of sight along which each object was observed is retained in storage alongside the position already persisted there. Response to Arguments Applicant’s arguments have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Applicant argues at pages 12 and 13 of the Remarks that Newcombe's projection operates on depth values from a depth camera and not on two-dimensional positions of semantically detected objects, that Choi's detection operates entirely in two-dimensional image space, and that the combination therefore does not teach or suggest "a three-dimensional position and a size of the detected object estimated by projecting a two-dimensional position of the detected object onto the spatial structure using position and posture information of the image" as recited in amended claims 1 and 11. This argument is persuasive with respect to Newcombe and Choi alone. However, upon further consideration and as necessitated by Applicant's amendment, a new ground of rejection under 35 U.S.C. 103 over Newcombe in view of Choi, Stauber, and Forsblom is made as set forth above with respect to claim 1, and a new ground of rejection under 35 U.S.C. 103 over Newcombe in view of Choi and Stauber is made as set forth above with respect to claim 11. Stauber teaches the operation Applicant identifies as absent from Newcombe and Choi. Stauber calculates "a 2D bounding box around this object in the first 2D frame" and projects "the 2D bounding box and the object depicted in two dimensions in the first 2D frame onto the 3D manifold generated from preceding 2D frames and concurrent motion of the mobile device and defined relative to the camera," then casts a ray "from the position of the camera" and stores "the lateral, longitudinal, and depth positions of this intersection relative to the camera", and calculates "a 3D bounding box that encompasses this cluster of 3D points in the 3D graph". The position and posture information of the image is the tracked "changes in position and orientation" of the device from which the manifold is generated. What Stauber projects onto the three-dimensional structure is the two-dimensional bounding box of a semantically detected object, not a depth value, which is the distinction Applicant draws. Applicant argues at page 13 of the Remarks that neither Newcombe nor Choi discloses an image semantic and spatial structure summary unit configured to summarize data to be presented to a user, as recited in amended claim 1. This argument is persuasive with respect to Newcombe and Choi alone. However, upon further consideration and as necessitated by Applicant's amendment, a new ground of rejection under 35 U.S.C. 103 over Newcombe in view of Choi, Stauber, and Forsblom is made as set forth above. Forsblom teaches that where "too much visual clutter would result by displaying each of the marker objects in that bounding area 50, they are not displayed and instead represented with an aggregate marker, and that the aggregate marker "is an icon including the number of marker objects within the particular bounding area 50". Condensing a set of individual objects into a single marker that conveys representative information about them is summarizing data to be presented to a user. Applicant argues at page 13 of the Remarks that claims 2-5 depend from claim 1 and are allowable for at least the same reasons as claim 1, and that claim 11 is allowable for the same reasons as claim 1. This argument has been considered but is moot because it does not apply to the new combinations of references being used in the current rejections. Applicant argues at pages 13 and 14 of the Remarks that Shen, Forsblom, Hoffmann, and Spiegel do not cure the deficiencies of Newcombe and Choi with respect to the amended limitations of claim 1, and that claims 6-10 are therefore allowable for at least the same reasons as claim 1. This argument has been considered but is moot because it does not apply to the new combinations of references being used in the current rejections. The limitations Applicant identifies as absent are now supplied by Stauber and Forsblom in the rejection of claim 1, from which claims 6-10 depend directly or indirectly, and Shen, Hoffmann, and Spiegel are relied upon only for the additional limitations of claims 6, 9, and 10 respectively. Applicant argues at pages 14 and 15 of the Remarks that new claims 12-20 recite additional features not taught by the cited prior art, and identifies for each claim the feature said to be absent. These arguments have been considered but are not persuasive. Claims 12-20 are rejected as set forth above over combinations that include references not previously applied. Specifically, the collision-based estimation of claim 12 is taught by Abeywardena, which projects a ray "along a vector that extends from the camera optical center through a center pixel of the annotated location of the structure 812" and teaches that "a terrain model intersection point 814 is found based on where the ray 810 intersects the terrain model 808." The reliability value of claim 13 is taught by Choi, where "each location in the heat map 214 indicates the confidence of an object in the corresponding location, scale, and subcategory." The user-designated viewpoint of claim 14 is taught by Nagasawa at paragraph [0028], where "the user can arbitrarily designate a desired viewpoint and a direction of the line of sight by operating the operating unit 202." The icon placement of claim 15 is taught by Zhao, which "plots the obtained three-dimensional coordinates as the location of the deteriorated portion on the target three-dimensional model" and "sets a deteriorated portion image for highlighting the deteriorated portion on the target three-dimensional model with a predetermined color or an icon." The pop-up display of claim 16 is taught by Kang, where "when the driving-related event icon 1206 is selected" the terminal "may display an image of an event corresponding to the selected event driving-related event as a pop-up window." The distance-range search of claim 17 is taught by Tysowski, where a search routine determines "whether or not images exist in those subfolders for the chosen radius," in combination with Hoffmann's gaze endpoint. The unmanned aerial vehicle of claim 18 is taught by Ohtomo. The fixed-length vector data of claim 19 is taught by Watanabe, where "[a]n image feature value means fixed-length numerical vector data obtained by digitizing features such as color and a shape which an image itself has." The stored line direction of claim 20 is taught by Kutliroff , which stores "the vector representing the ray extending from the camera's position to that point (which is referred to herein as the "camera ray")." Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jai Li whose telephone number is (571)272-1170. The examiner can normally be reached Mon-Thu between 06:00-16:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571)272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAI W LI/Junior Examiner, Art Unit 2613 /DAVID T WELCH/Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Oct 18, 2024
Application Filed
Apr 22, 2026
Non-Final Rejection mailed — §103
Jun 22, 2026
Response Filed
Jul 23, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month