Prosecution Insights
Last updated: October 04, 2026
Application No. 18/919,404

VOLUMETRIC SCENE RECONSTRUCTION WITH VARIABLE VOXEL RESOLUTION IN TRUNCATED SIGNED DISTANCE FUNCTION (TSDF) FUSION

Final Rejection §103
Filed
Oct 17, 2024
Examiner
LI, JAI WEI TOMMY
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Niantic, Inc.
OA Round
2 (Final)
Grant Probability
Favorable
3-4
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-62.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
32 currently pending
Career history
33
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The objections to the claims and the specifications have been withdrawn in view of the applicants amendments filed on 08/18/2026. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Meilland et al. (U.S. Pub. No 2021/0225074) in view of Godard et al. (U.S. Pub. No. 2019/0356905) and Armstrong et al. (U.S. Pub. No. 2023/235198). Regarding claim 1, Meilland et al. discloses a computer-implemented method comprising (paragraph 4, line(s) 1-6, “Various implementations disclosed herein include devices, systems, and methods that generate a mesh (e.g., vertices that form connected triangles) representing the surfaces in a physical environment using multi-resolution voxels that are generated based on detected depth information.”; also, paragraph 54, line(s) 11-13, “the method 400 is performed by processing logic, including hardware, firmware, software, or a combination thereof”): receiving image data capturing a real-world environment and captured by a camera assembly of a client device, the image data comprising a plurality of frames (paragraph 7, line(s) 4-8, “The method involves obtaining depth data of a physical environment using a sensor. For example, the depth data can include pixel depth values from a viewpoint and sensor position and orientation data.”; also, paragraph 14, line(s) 1-5, “In some implementations, the depth data is obtained using one or more depth cameras. For example, the one or more depth cameras can acquire depth based on structured light (SL), passive stereo (PS), active stereo (AS), time of flight (ToF), and the like.”; also, paragraph 79, line(s) 5-8, “Example environment 1100 is an example of acquiring image data (e.g., light intensity data and depth data) for a plurality of image frames.”); semantic segmentation model to each frame to output a segmentation mask corresponding to the frame, wherein the segmentation mask classifies pixels in the frame into one of a plurality of semantic classes (paragraph 61, line(s) 4-7, “semantic labeling uses a machine learning model, where a semantic segmentation model may be configured to identify semantic labels for pixels or voxels of image data.”; also, paragraph 61, line(s) 7-12, “if a voxel is labeled as a “wall,” regardless of how far away the object (wall) is, the system can select a bigger/sparser resolution based on the assumption or specification that a wall type object will have a consistent or insignificant texture that need not be represented using fine resolution.”); determining level hints for each frame based on the segmentation mask and the depth map corresponding to the frame (paragraph 9, line(s) 7-9, “the resolution level used for each voxel may be determined based on distance from the sensor, noise, semantics, and the like.”; also, paragraph 61, line(s) 1-3, “the resolution levels may be determined based on semantic labeling identifying object type (e.g., table, teapot, chair, vase, etc.).”), variable-resolution truncated signed distance function (TSDF) grid by fusing depth predictions from the depth maps corresponding to the plurality of frames (paragraph 8, line(s) 1-6, “The exemplary method further involves generating a first hash table storing 3D positions of a first set of voxels having a first resolution (e.g., big voxels) and signed distance values representing distances to the surfaces (e.g., to a nearest surface) of the physical environment based on the depth data.”; also, paragraph 8, line(s) 9-12, “the signed distance values include TSDF values that may be used to represent voxel distances of each voxel to a nearest surface of the surfaces of the physical environment.”), the variable-resolution TSDF grid comprising TSDF values indicating distance to a surface in the real-world environment, wherein the variable-resolution TSDF grid includes at least one portion at a first voxel resolution level and another portion at a second voxel resolution level of finer resolution than the first voxel resolution level (paragraph 5, line(s) 2-7, “Using multi-resolution voxels provides some portions of a reconstruction with smaller voxels to provide finer resolution and thus potentially higher accuracy and fidelity, and other portions of the reconstruction with larger voxels to provide coarser resolution and thus less accuracy and fidelity.”; also, paragraph 12, line(s) 13-16, “In some implementations, voxels of the first set of voxels have a first size and voxels of the second set of voxels have a second size, where the first size is larger than the second size.”), and wherein the voxel resolution at each portion of the variable- resolution TSDF grid is based on the level hints (paragraph 5, line(s) 12-16, “voxel size, e.g., which voxels are small and which voxels are large, may be determined using criteria that provides for the use of smaller voxels in areas where doing so will likely result in greater accuracy, e.g., where there is less noise in the data,”; also, paragraph 4, line(s) 12-15, “Those resolutions may be selected based upon distances of the voxels from the sensor, noise in the depth data associated with different voxels, or other factors.”); generating a polygon mesh from the variable-resolution TSDF grid digitally representing surfaces in the real-world environment captured by the image data (paragraph 6, line(s) 1-6, “Techniques disclosed herein may use the multi-resolution voxel data to generate a mesh that reconstructs the geometry of the physical environment. This may involve using a meshing algorithm that combines multi-resolution voxel information stored in multiple hash tables to generate a single mesh.”; also, paragraph 10, line(s) 11-15, “the mesh may be generated using a marching cubes meshing algorithm technique that identifies lines connecting points associated with the voxels in each hash table and interpolates to identify vertices along those lines that correspond to the surfaces.”); . Meilland does not disclose applying a depth estimation model to each frame to output a depth map corresponding to the frame, wherein the depth map comprises depth predictions for pixels in the frame, the level hints indicate a voxel resolution level for each pixel, storing the polygon mesh in a map database. However, in a similar field of endeavor, Godard et al. discloses applying a depth estimation model to each frame to output a depth map corresponding to the frame (Godard: paragraph 30, line(s) 1-3, “The depth estimation model 130 receives an input image of a scene and outputs a depth of the scene based on the input image”; also, paragraph 6, line(s) 3-5, “The system inputs the images into a depth model to extract a depth map for each image based on parameters of the depth model.”), Armstrong et al. discloses the level hints indicate a voxel resolution level for each pixel (Armstrong: paragraph 24, line(s) 30-32, “the multi-resolution voxel space may have a plurality of semantic layers in which each semantic layer comprises a plurality of voxel grids representing voxels as covariance ellipsoids at different resolutions.”; also, paragraph 26, line(s) 10-12, “the semantic labels may comprise a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree, road surface, curb, sidewalk, unknown, etc.”) and storing the polygon mesh in a map database (Armstrong: paragraph 24, line(s) 28-29, “the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing using TSDF values from depth camera data with the features of Godard et al.'s invention of a learned depth estimation model that generates depth maps from input images, further in view of Armstrong et al.'s invention of multi-resolution voxel spaces with semantic layers for environment representation and storage. Regarding element “discloses a computer-implemented method”, Meilland et al. teaches "devices, systems, and methods that generate a mesh representing the surfaces in a physical environment using multi-resolution voxels." The specification further confirms that "the method 400 is performed by processing logic, including hardware, firmware, software, or a combination thereof." This directly corresponds to the claimed computer-implemented method. Regarding element” receiving image data capturing a real-world environment and captured by a camera assembly of a client device, the image data comprising a plurality of frames”, Meilland et al. teaches "obtaining depth data of a physical environment using a sensor" where "the depth data can include pixel depth values from a viewpoint and sensor position and orientation data." Meilland et al. further teaches that the data is obtained using "one or more depth cameras" capable of acquiring depth through multiple modalities including structured light, passive stereo, active stereo, and time of flight. Under BRI, the one or more depth cameras constitute a camera assembly of a client device that captures image data (including depth data) of a real-world environment comprising a plurality of frames over time for real-time reconstruction. Regarding element “applying a depth estimation model to each frame to output a depth map corresponding to the frame, wherein the depth map comprises depth predictions for pixels in the frame”, Godard et al. teaches "the depth estimation model receives an input image of a scene and outputs a depth of the scene based on the input image" and that "the system inputs the images into a depth model to extract a depth map for each image based on parameters of the depth model." This directly teaches applying a learned depth estimation model to each frame to output a depth map comprising depth predictions for pixels in the frame. Regarding element “applying a semantic segmentation model to each frame to output a segmentation mask corresponding to the frame, wherein the segmentation mask classifies pixels in the frame into one of a plurality of semantic classes”, Meilland et al. explicitly teaches that "semantic labeling uses a machine learning model, where a semantic segmentation model may be configured to identify semantic labels for pixels or voxels of image data." Meilland et al. further teaches that these semantic labels are used to classify scene elements — for example, labeling a voxel as a "wall" to determine appropriate resolution. Under BRI, this teaches applying a semantic segmentation model to each frame to output a segmentation mask classifying pixels into a plurality of semantic classes. Regarding element “determining level hints for each frame based on the segmentation mask and the depth map corresponding to the frame”, Meilland et al. teaches that "the resolution level used for each voxel may be determined based on distance from the sensor, noise, semantics, and the like" and "the resolution levels may be determined based on semantic labeling identifying object type (e.g., table, teapot, chair, vase, etc.)." This teaches determining resolution guidance (i.e., level hints) for each frame based on both the semantic information (segmentation mask) and depth-related information (distance from sensor/depth map). The explicit inclusion of both "semantics" and "distance from the sensor" as resolution determinants maps directly to level hints based on both the segmentation mask and the depth map. Regarding element “wherein the level hints indicate a voxel resolution level for each pixel”, Armstrong et al. teaches multi-resolution voxel spaces with "a plurality of semantic layers in which each semantic layer comprises a plurality of voxel grids representing voxels as covariance ellipsoids at different resolutions" and that "the semantic labels may comprise a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree, road surface, curb, sidewalk, unknown, etc." Under BRI, the assignment of semantic labels combined with different resolution levels per semantic layer teaches that the level hints indicate a voxel resolution level, which is determinable per-pixel based on the pixel's semantic class. Regarding element “generating a variable-resolution truncated signed distance function (TSDF) grid by fusing depth predictions from the depth maps corresponding to the plurality of frames”, Meilland et al. teaches generating "a first hash table storing 3D positions of a first set of voxels having a first resolution (e.g., big voxels) and signed distance values representing distances to the surfaces (e.g., to a nearest surface) of the physical environment based on the depth data." Meilland et al. further teaches that "the signed distance values include TSDF values that may be used to represent voxel distances of each voxel to a nearest surface." This directly teaches generating a TSDF grid by fusing depth data from multiple frames, where the grid contains TSDF values indicating distance to surfaces. Regarding element “the variable-resolution TSDF grid comprising TSDF values indicating distance to a surface in the real-world environment, wherein the variable-resolution TSDF grid includes at least one portion at a first voxel resolution level and another portion at a second voxel resolution level of finer resolution than the first voxel resolution level”, Meilland et al. teaches that "multi-resolution voxels provides some portions of a reconstruction with smaller voxels to provide finer resolution and thus potentially higher accuracy and fidelity, and other portions of the reconstruction with larger voxels to provide coarser resolution" and that "voxels of the first set of voxels have a first size and voxels of the second set of voxels have a second size, where the first size is larger than the second size." This directly teaches a variable-resolution TSDF grid with at least one portion at a first resolution level and another at a second, finer resolution level. Regarding element “and wherein the voxel resolution at each portion of the variable-resolution TSDF grid is based on the level hints”, Meilland et al. teaches that "voxel size, e.g., which voxels are small and which voxels are large, may be determined using criteria" and that "resolutions may be selected based upon distances of the voxels from the sensor, noise in the depth data associated with different voxels, or other factors." Combined with Meilland et al.'s teaching in element “determining level hints for each frame based on the segmentation mask and the depth map corresponding to the frame” that resolution is determined based on "distance from the sensor, noise, semantics, and the like," this teaches that voxel resolution at each portion of the TSDF grid is based on the resolution guidance criteria (level hints). Regarding element “generating a polygon mesh from the variable-resolution TSDF grid digitally representing surfaces in the real-world environment captured by the image data”, Meilland et al. teaches using "a meshing algorithm that combines multi-resolution voxel information stored in multiple hash tables to generate a single mesh" and specifically teaches "a marching cubes meshing algorithm technique that identifies lines connecting points associated with the voxels in each hash table and interpolates to identify vertices along those lines that correspond to the surfaces." This directly teaches generating a polygon mesh from the variable-resolution TSDF grid that digitally represents surfaces in the real-world environment. Regarding element “storing the polygon mesh in a map database”, Armstrong et al. teaches that "the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces" for purposes including "localization, object tracking, and/or navigation of an autonomous vehicle with respect to the physical environment." Under BRI, storing 3D reconstruction data representing the environment in a persistent data structure for later use by autonomous vehicles constitutes storing the polygon mesh in a map database, which is a standard design choice for mapping and navigation applications. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches generating multi-resolution voxel meshes using TSDF values from depth camera data and using semantic labeling to determine voxel resolution, but relies on hardware depth sensors such as structured light and time of flight cameras for depth data. Godard et al. teaches a trained depth estimation model that "receives an input image of a scene and outputs a depth of the scene" and extracts "a depth map for each image based on parameters of the depth model," providing a learned alternative to hardware depth sensors that could directly replace or supplement Meilland et al.'s depth cameras. Armstrong et al. teaches multi-resolution voxel spaces with "a plurality of semantic layers" at "different resolutions" and storing environment data as multi-resolution voxel spaces for "localization, object tracking, and/or navigation," providing the semantic layer structure and persistent storage that Meilland et al.'s system would benefit from. A POSITA would recognize that substituting Godard et al.'s learned depth model for Meilland et al.'s hardware sensors eliminates the need for specialized depth cameras, and that Armstrong et al.'s storage and semantic layer architecture extends Meilland et al.'s reconstruction for downstream navigation use, yielding a predictable combined system. Regarding claim 2, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses, the computer-implemented method of claim 1, further comprising: variable-resolution TSDF grid is further based on the tracked one or more objects in the real-world environment (Meilland et al.: paragraph 61, line(s) 1-3, “the resolution levels may be determined based on semantic labeling identifying object type (e.g., table, teapot, chair, vase, etc.).”; also, paragraph 41, line(s) 13-15, “the detection, tracking, and representing of objects in 3D space relative to one another based on stored 3D models”). Meilland et al. as modified by Godard et al. and Armstrong et al. does not disclose applying an object detection model to each frame to identify one or more objects in the frame; and tracking one or more of the objects across frames. However, in a similar field of endeavor, Armstrong et al. further discloses applying an object detection model to each frame to identify one or more objects in the frame (Armstrong et al.: paragraph 26, line(s) 10-12, “the semantic labels may comprise a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree, road surface, curb, sidewalk, unknown, etc.”; also, paragraph 12, line(s) 27-29, “map data represented by a multi-resolution voxel space may be generated from data points representing a physical environment, such as an output of a light detection and ranging (lidar) system.”); and tracking one or more of the objects across frames (Armstrong et al.: paragraph 30, line(s) 17-18, “to assist with localization, object tracking, and/or navigation of an autonomous vehicle with respect to the physical environment.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the features of Armstrong et al.'s object detection, semantic labeling, and object tracking for environment representation. Regarding element “applying an object detection model to each frame to identify one or more objects in the frame”, Armstrong et al. teaches that semantic labels are assigned comprising "a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree, road surface, curb, sidewalk, unknown, etc." Under BRI, identifying entities by class or type from sensor data constitutes applying an object detection model to identify one or more objects in each frame. Armstrong et al. generates multi-resolution voxel spaces from sensor data with semantic classifications, which requires detecting and classifying objects in the input data. Regarding element “tracking one or more of the objects across frames”, Armstrong et al. teaches that the multi-resolution voxel space system assists with "localization, object tracking, and/or navigation of an autonomous vehicle with respect to the physical environment." This directly teaches tracking one or more objects across frames, as object tracking inherently requires maintaining object identity across sequential observations. Regarding element “generating the variable-resolution TSDF grid is further based on the tracked one or more objects in the real-world environment”, Meilland et al. teaches that "resolution levels may be determined based on semantic labeling identifying object type (e.g., table, teapot, chair, vase, etc.)" and that the system supports "detection, tracking, and representing of objects in 3D space relative to one another." This teaches that the variable-resolution TSDF generation is based on detected and tracked objects — objects identified by semantic type receive appropriate resolution treatment. One of ordinary skill in the art would have been motivated to combine because Armstrong et al. teaches assigning semantic labels comprising "a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree" and assisting with "object tracking," while Meilland et al. already teaches that "resolution levels may be determined based on semantic labeling identifying object type" and supports "detection, tracking, and representing of objects in 3D space." A POSITA working with Meilland et al.'s semantic-based resolution system would naturally incorporate Armstrong et al.'s object detection and tracking to maintain consistent object identity across frames, because knowing what objects are present and tracking them over time directly improves the resolution allocation that Meilland et al. already performs based on object type. Regarding claim 3, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the computer-implemented method of claim 2, wherein the object detection model is trained as a machine-learning model (Meilland et al.: paragraph 61, line(s) 16-18, “the machine learning model is a neural network (e.g., an artificial neural network), decision tree, support vector machine”; also, paragraph 61, line(s) 4-7, “semantic labeling uses a machine learning model, where a semantic segmentation model may be configured to identify semantic labels for pixels or voxels of image data.”) in Meilland et al. as modified by Godard et al. and Armstrong et al. does not disclose a supervised manner with training image data labeled with identified objects. However, in a similar field of endeavor, Godard et al. discloses a supervised manner with training image data labeled with identified objects (Godard et al.: paragraph 3, line(s) 5-10, “A depth estimation system may be trained using a detection and ranging system to establish a ground truth depth for objects in an environment (i.e., radio detecting and ranging (RADAR), light detection and ranging (LIDAR), etc.) paired with images taken of the same scene by a camera.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces with object tracking, with the features of supervised machine learning training for the object detection model. Regarding element “the object detection model is trained as a machine-learning model”, Meilland et al. teaches that "the machine learning model is a neural network (e.g., an artificial neural network), decision tree, support vector machine" and that "semantic labeling uses a machine learning model." This teaches that the object detection model is a machine-learning model such as a neural network. Regarding element “a supervised manner with training image data labeled with identified objects”, Godard et al. teaches that "a depth estimation system may be trained using a detection and ranging system to establish a ground truth depth for objects in an environment (i.e., radio detecting and ranging (RADAR), light detection and ranging (LIDAR), etc.) paired with images taken of the same scene by a camera." This teaches the supervised training paradigm of pairing labeled ground truth data with corresponding images. A POSITA would understand that applying this same supervised training paradigm to train Meilland et al.'s neural network-based object detection model with labeled training image data identifying objects is the standard approach for such models in computer vision. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches that "the machine learning model is a neural network" and that "semantic labeling uses a machine learning model", but does not disclose how these models are trained. Godard et al. teaches that "a depth estimation system may be trained using a detection and ranging system to establish a ground truth depth for objects in an environment... paired with images taken of the same scene by a camera", which is the supervised training paradigm of pairing labeled ground truth data with corresponding images. A POSITA implementing Meilland et al.'s neural network for object detection would naturally apply the same supervised training paradigm taught by Godard et al., because training a neural network to identify objects requires labeled training data with identified objects, and this is the standard methodology for such models in computer vision. Regarding claim 4, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the computer-implemented method of claim 1, generating the variable-resolution TSDF grid is further based on the one or more surface orientations of the one or more surfaces (Meilland: paragraph 5, line(s) 12-16, “voxel size, e.g., which voxels are small and which voxels are large, may be determined using criteria that provides for the use of smaller voxels in areas where doing so will likely result in greater accuracy, e.g., where there is less noise in the data,”; also, paragraph 4, line(s) 12-15, “Those resolutions may be selected based upon distances of the voxels from the sensor, noise in the depth data associated with different voxels, or other factors.”). Meilland et al. as modified by Godard et al. and Armstrong et al. does not disclose applying the depth estimation model to each frame further comprises applying the depth estimation model to identify surface orientation of one or more surfaces present in the frame. However, in a similar field of endeavor, Godard et al. further discloses applying the depth estimation model to each frame further comprises applying the depth estimation model to identify surface orientation of one or more surfaces present in the frame (Godard: paragraph 6, line(s) 3-5, “The system inputs the images into a depth model to extract a depth map for each image based on parameters of the depth model.”; also, paragraph 6, line(s) 20-22, “Upsampled depth features may also be used during generation of the synthetic frames which would affect the appearance matching loss calculations.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the features of surface orientation estimation from the depth estimation model for resolution guidance. Regarding element “applying the depth estimation model to each frame further comprises applying the depth estimation model to identify surface orientation of one or more surfaces present in the frame”, Godard et al. teaches that "the system inputs the images into a depth model to extract a depth map for each image" and that "upsampled depth features may also be used during generation of the synthetic frames which would affect the appearance matching loss calculations." The depth features extracted by the model inherently encode geometric information about the scene, including surface orientation. Under BRI, a depth estimation model that produces per-pixel depth maps implicitly captures surface orientation information, as surface normals can be computed from the spatial gradients of the depth map. Extending a depth estimation model to explicitly output surface orientation alongside depth is a routine design choice well-known in the art, as surface normal estimation networks share the same encoder architecture as depth estimation networks. Regarding element “generating the variable-resolution TSDF grid is further based on the one or more surface orientations of the one or more surfaces”, Meilland et al. teaches that "voxel size may be determined using criteria that provides for the use of smaller voxels in areas where doing so will likely result in greater accuracy, e.g., where there is less noise in the data" and that "resolutions may be selected based upon distances of the voxels from the sensor, noise in the depth data associated with different voxels, or other factors." The explicit mention of "or other factors" encompasses surface orientation as a resolution criterion. A POSITA would recognize that surfaces viewed at oblique angles produce noisier depth estimates, making surface orientation a geometrically relevant factor for the resolution criteria Meilland et al. already teaches. One of ordinary skill in the art would have been motivated to combine because Godard et al. teaches that "the system inputs the images into a depth model to extract a depth map for each image" and that "upsampled depth features may also be used during generation of the synthetic frames", meaning the depth model extracts rich geometric features including surface information from each frame. Meilland et al. teaches that voxel resolution is selected based on "distances of the voxels from the sensor, noise in the depth data associated with different voxels, or other factors". A POSITA would recognize that the depth features extracted by Godard et al.'s model inherently encode surface orientation information (as surface normals are computable from depth gradients), and that surface orientation is one of Meilland et al.'s "other factors" affecting depth noise. Extending Godard et al.'s depth model to also output surface orientation and feeding that into Meilland et al.'s resolution selection would predictably improve reconstruction quality for oblique surfaces. Regarding claim 5, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the computer-implemented method of claim 1, wherein depth estimation model is trained as a machine-learning model in a self-supervised manner by projecting frames from training image data onto other frames of the training image data based on depth predictions by the depth estimation model. However, in a similar field of endeavor, Godard et al. further discloses the depth estimation model is trained as a machine-learning model in a self-supervised manner by projecting frames from training image data onto other frames of the training image data based on depth predictions by the depth estimation model (Godard: paragraph 5, line(s) 1-6, “Some more novel methods train a depth estimation system utilizing monocular video data of an ever changing scene. The depth estimation system trains by projecting from one temporal image in the monocular video data to a subsequent temporal image while minimizing a photometric reconstruction error”; also, abstract, line(s) 6-8, “The method includes generating a plurality of synthetic frames based on the depth map and the pose for each image”; also, paragraph 6, line(s) 13-16, “The loss function includes a calculation of the photometric reconstruction error per pixel between a synthetic frame and an input image”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the features of Godard et al.'s self-supervised depth estimation training methodology. Regarding element “depth estimation model is trained as a machine-learning model in a self-supervised manner by projecting frames from training image data onto other frames of the training image data based on depth predictions by the depth estimation model”, Godard et al. explicitly teaches self-supervised training, disclosing that "some more novel methods train a depth estimation system utilizing monocular video data of an ever changing scene" and that "the depth estimation system trains by projecting from one temporal image in the monocular video data to a subsequent temporal image while minimizing a photometric reconstruction error." Godard et al. further teaches that "the method includes generating a plurality of synthetic frames based on the depth map and the pose for each image" and that "the loss function includes a calculation of the photometric reconstruction error per pixel between a synthetic frame and an input image." This is exactly the claimed self-supervised training methodology — the depth estimation model is trained by projecting frames from training image data (monocular video) onto other frames based on the depth map (depth predictions by the depth estimation model) and pose, and the photometric error between the synthetic frame and the actual input image serves as the self-supervised loss signal, without requiring ground truth depth labels. One of ordinary skill in the art would have been motivated to combine because Godard et al. teaches that "the depth estimation system trains by projecting from one temporal image in the monocular video data to a subsequent temporal image while minimizing a photometric reconstruction error", that "the method includes generating a plurality of synthetic frames based on the depth map and the pose for each image," and that "the loss function includes a calculation of the photometric reconstruction error per pixel between a synthetic frame and an input image." This self-supervised training approach eliminates the need for ground truth depth labels, which Godard et al. acknowledges are costly: "constantly utilizing detection and ranging systems to sense depth of many different environments can be a costly endeavor in time and resources." A POSITA seeking to train the depth estimation model used in the combined system would naturally adopt Godard et al.'s self-supervised methodology because it enables training on the same monocular video data that Meilland et al.'s pipeline already processes, without requiring any additional labeled datasets. Regarding claim 6, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the computer-implemented method of claim 1, wherein generating the variable-resolution TSDF grid is constrained by limiting neighboring voxel cells to be at most one voxel resolution level difference (Meilland: paragraph 11, line(s) 8-11, “Additional hash tables can be generated for each level of resolution desired. In an exemplary implementation, at least four hash tables are utilized for four different resolutions”; also, paragraph 10, line(s) 3-7, “The mesh may be generated by positioning vertices along a line connecting a first voxel (e.g., a position at the center of the first voxel) of the first set of voxels with a second voxel (e.g., a position at the center of the second voxel) of the second set of voxels.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the constraint of limiting neighboring voxel cells to at most one resolution level difference. Regarding element “generating the variable-resolution TSDF grid is constrained by limiting neighboring voxel cells to be at most one voxel resolution level difference”, Meilland et al. teaches using multiple hash tables for multiple resolution levels, with "at least four hash tables utilized for four different resolutions," and generates meshes by connecting voxels across resolution levels. When multiple resolution levels are present in a spatial grid, the transition between adjacent regions of different resolution must be managed to maintain geometric consistency. Constraining neighboring voxel cells to at most one voxel resolution level difference is a well-established technique known as "2:1 balancing" or "graded" refinement in adaptive mesh refinement and octree-based spatial data structures. Meilland et al.'s meshing across resolution levels inherently requires managing resolution transitions, and a POSITA working with Meilland et al.'s four-level multi-resolution system would naturally apply this standard 2:1 balancing constraint to prevent discontinuities at resolution boundaries and ensure mesh quality. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches using "at least four hash tables for four different resolutions" and generating meshes by "positioning vertices along a line connecting a first voxel of the first set of voxels with a second voxel of the second set of voxels". When connecting voxels across multiple resolution levels in this manner, a POSITA would recognize that allowing arbitrary resolution jumps between neighboring voxels (e.g., from the finest to the coarsest resolution) creates geometric discontinuities and mesh artifacts at the transition boundaries. Constraining neighboring voxel cells to at most one resolution level difference (2:1 balancing) is a well-established technique in adaptive mesh refinement that directly addresses this problem, and applying it to Meilland et al.'s four-level multi-resolution system would predictably produce smoother, higher-quality meshes at resolution transitions. Regarding claim 7, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the computer-implemented method of claim 1, wherein generating the variable- resolution TSDF grid comprises implementing a hyperparameter that sets a quantity of depth predictions fused into the TSDF value per voxel, wherein the hyperparameter is fit to an error curve for depth predictions by the depth estimation model (Meilland et al.: paragraph 8, line(s) 12-15, “The TSDF values can save storage space by including only values within a truncation band in the representation, e.g., only storing data for voxels that are within a threshold distance of a surface.”; also, paragraph 59, line(s) 15-18, “Distance may be used as an approximation of noise based on a correlation (e.g., quadratic noise with respect to distance such that farther distance means more noise).”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with a hyperparameter controlling the fusion quantity based on the depth estimation model's error characteristics. Regarding element “generating the variable- resolution TSDF grid comprises implementing a hyperparameter that sets a quantity of depth predictions fused into the TSDF value per voxel, wherein the hyperparameter is fit to an error curve for depth predictions by the depth estimation model”, Meilland et al. teaches that "TSDF values can save storage space by including only values within a truncation band" (a threshold parameter governing TSDF fusion) and that "distance may be used as an approximation of noise based on a correlation (e.g., quadratic noise with respect to distance such that farther distance means more noise)." This teaches that Meilland et al. explicitly models the error characteristics of depth measurements as a function of distance and uses this noise model to inform TSDF parameters. Under BRI, fitting a hyperparameter that controls fusion quantity to an error curve of the depth model is a logical extension of Meilland et al.'s teaching of correlating noise with distance — a POSITA would understand that the number of depth predictions needed to achieve a reliable TSDF value depends on the noise characteristics of the depth estimation model, and that calibrating (fitting) this relationship to the model's actual error curve is a routine optimization step. However, the specific implementation of fitting a fusion quantity hyperparameter to the depth model's error curve represents a more targeted optimization. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches that "TSDF values can save storage space by including only values within a truncation band" and that "distance may be used as an approximation of noise based on a correlation (e.g., quadratic noise with respect to distance)". This shows Meilland et al. already models depth error as a function of distance and uses error characteristics to set TSDF parameters. When integrating Godard et al.'s learned depth estimation model into this pipeline, a POSITA would naturally calibrate the fusion quantity hyperparameter to the depth model's actual error curve, because the number of depth predictions needed per voxel depends on how noisy the depth model's predictions are at that distance. Regions with higher depth noise need more observations to average out errors, while regions with lower noise achieve accurate reconstruction with fewer, predictably improving both TSDF accuracy and computational efficiency. Regarding claim 8, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the computer-implemented method of claim 1, wherein generating the polygon mesh from the variable-resolution TSDF grid comprises interpolating between neighboring voxels of different voxel resolution levels (Meilland et al.: paragraph 10, line(s) 3-7, “The mesh may be generated by positioning vertices along a line connecting a first voxel (e.g., a position at the center of the first voxel) of the first set of voxels with a second voxel (e.g., a position at the center of the second voxel) of the second set of voxels.”; also, paragraph 10, line(s) 22-28, “generating the mesh includes generating lines connecting points associated with the voxels in each of the multiple hash tables (e.g., between the first hash table and the second hash table, between the second hash table and the third hash table, etc.) and interpolating along the lines to identify vertices for the mesh that correspond to the surfaces.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with Meilland et al.'s interpolation technique for cross-resolution mesh generation. Regarding element “generating the polygon mesh from the variable-resolution TSDF grid comprises interpolating between neighboring voxels of different voxel resolution levels”, Meilland et al. explicitly teaches generating a mesh by "positioning vertices along a line connecting a first voxel of the first set of voxels with a second voxel of the second set of voxels" where the first and second sets are at different resolutions stored in different hash tables. Meilland et al. further teaches "interpolating along the lines to identify vertices for the mesh that correspond to the surfaces" specifically "between the first hash table and the second hash table, between the second hash table and the third hash table, etc." This directly teaches interpolating between neighboring voxels of different voxel resolution levels during polygon mesh generation from the multi-resolution TSDF grid. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches generating a mesh by "positioning vertices along a line connecting a first voxel of the first set of voxels with a second voxel of the second set of voxels" and "interpolating along the lines to identify vertices for the mesh that correspond to the surfaces" specifically "between the first hash table and the second hash table, between the second hash table and the third hash table, etc.". This interpolation between voxels of different resolution levels is already an integral part of Meilland et al.'s own multi-resolution meshing approach. Applying this same technique within the combined system's variable-resolution TSDF grid would produce a continuous, seamless polygon mesh without gaps or discontinuities at resolution boundaries. Regarding claim 9, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the computer-implemented method of claim 1, further comprising: augmenting the polygon mesh with visual patterns from the image data corresponding to one or more surfaces represented by the polygon mesh (Meilland et al.: paragraph 4, line(s) 1-6, “Various implementations disclosed herein include devices, systems, and methods that generate a mesh (e.g., vertices that form connected triangles) representing the surfaces in a physical environment using multi-resolution voxels that are generated based on detected depth information.”; also, paragraph 14, line(s) 1-5, “In some implementations, the depth data is obtained using one or more depth cameras. For example, the one or more depth cameras can acquire depth based on structured light (SL), passive stereo (PS), active stereo (AS), time of flight (ToF), and the like”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the augmentation of the polygon mesh with image-derived visual patterns (texture mapping). Regarding element “augmenting the polygon mesh with patterns from the image data corresponding to one or more surfaces represented by the polygon mesh”, Meilland et al. teaches generating "a mesh (e.g., vertices that form connected triangles) representing the surfaces in a physical environment" from depth data obtained by "one or more depth cameras" that can acquire data through multiple modalities including passive stereo and active stereo, which inherently capture RGB image data alongside depth. Under BRI, "augmenting the polygon mesh with patterns from the image data" encompasses texture mapping — projecting the captured image data onto the mesh surfaces — which is a fundamental and well-known technique in 3D reconstruction. A POSITA working with Meilland et al.'s mesh reconstruction system would readily understand that once a polygon mesh is generated and camera poses are known, the original image data can be projected onto mesh surfaces to add visual texture. Texture mapping is a standard step in virtually all photogrammetric and depth-based reconstruction pipelines, and its application to Meilland et al.'s mesh output would yield a predictable result. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches generating "a mesh (e.g., vertices that form connected triangles) representing the surfaces in a physical environment" from depth data obtained using "one or more depth cameras" that can acquire data through "passive stereo (PS), active stereo (AS)" and other modalities, which inherently capture RGB image data alongside depth. Once the polygon mesh is generated and the camera poses are known from the image data, projecting the original image data onto the mesh surfaces to add visual texture is one of the most fundamental techniques in 3D reconstruction. A POSITA would naturally augment the bare geometric mesh with the patterns from the already-available image data, producing a photorealistic 3D model far more useful for the navigation and visualization applications taught by Armstrong et al. Regarding claim 10, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the computer-implemented method of claim 1, further comprising: receiving a request from a second client device to view the polygon mesh (Meilland et al.: paragraph 31, line(s)13-15, “the controller 110 is communicatively coupled with the device 120 via one or more wired or wireless communication channels”; also, paragraph 31, line(s) 10-12, “the controller 110 is a remote server located outside of the physical environment”); polygon mesh from the map database; and transmitting the polygon mesh to the second client device for presentation on the second client device. However, in a similar field of endeavor, Armstrong et al. further discloses retrieving the polygon mesh from the map database; and transmitting the polygon mesh to the second client device for presentation on the second client device (Armstrong et al.: paragraph 24, line(s) 28-29, “the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces”; also, paragraph 30, line(s) 17-18, “to assist with localization, object tracking, and/or navigation of an autonomous vehicle with respect to the physical environment”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the features of client-server architecture for mesh retrieval and transmission. Regarding element “receiving a request from a second client device to view the polygon mesh”, Meilland et al. teaches that "the controller 110 is communicatively coupled with the device 120 via one or more wired or wireless communication channels" and that "the controller 110 is a remote server located outside of the physical environment." This teaches a client-server architecture where devices communicate with a remote server, and under BRI, receiving a request from a second client device to view data is a standard client-server interaction. Regarding element “retrieving the polygon mesh from the map database; and transmitting the polygon mesh to the second client device for presentation on the second client device”, Armstrong et al. teaches that "the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces" for "localization, object tracking, and/or navigation of an autonomous vehicle with respect to the physical environment." Under BRI, a system that stores environment representations for use by autonomous vehicles necessarily retrieves and transmits that data to requesting vehicles (second client devices) for their localization and navigation needs. This is the standard architecture for map services in autonomous vehicle applications. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches that "the controller 110 is communicatively coupled with the device 120 via one or more wired or wireless communication channels" and that "the controller 110 is a remote server located outside of the physical environment", establishing a client-server architecture. Armstrong et al. teaches storing "the scene as well as data representative of environment as multi-resolution voxel spaces" for "localization, object tracking, and/or navigation of an autonomous vehicle". A POSITA working with this client-server architecture and stored environment data would naturally extend it to serve the polygon mesh to additional client devices upon request, because the whole purpose of Armstrong et al.'s stored environment data is to assist multiple autonomous vehicles with localization and navigation, which inherently requires retrieving and transmitting the data to those vehicles. Regarding claim 11, Meilland et al. discloses a non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the computer processor to perform operations comprising (paragraph 15, line(s) 8-11, “a non-transitory computer readable storage medium has stored therein instructions, which, when executed by one or more processors”; also, paragraph 36, line(s) 11-13, “The memory 220 comprises a non-transitory computer readable storage medium”): receiving image data capturing a real-world environment and captured by a camera assembly of a client device, the image data comprising a plurality of frames (paragraph 7, line(s) 4-8, “The method involves obtaining depth data of a physical environment using a sensor. For example, the depth data can include pixel depth values from a viewpoint and sensor position and orientation data.”; also, paragraph 14, line(s) 1-5, “In some implementations, the depth data is obtained using one or more depth cameras. For example, the one or more depth cameras can acquire depth based on structured light (SL), passive stereo (PS), active stereo (AS), time of flight (ToF), and the like.”; also, paragraph 79, line(s) 5-8, “Example environment 1100 is an example of acquiring image data (e.g., light intensity data and depth data) for a plurality of image frames.”); semantic segmentation model to each frame to output a segmentation mask corresponding to the frame, wherein the segmentation mask classifies pixels in the frame into one of a plurality of semantic classes (paragraph 61, line(s) 4-7, “semantic labeling uses a machine learning model, where a semantic segmentation model may be configured to identify semantic labels for pixels or voxels of image data.”; also, paragraph 61, line(s) 7-12, “if a voxel is labeled as a “wall,” regardless of how far away the object (wall) is, the system can select a bigger/sparser resolution based on the assumption or specification that a wall type object will have a consistent or insignificant texture that need not be represented using fine resolution.”); determining level hints for each frame based on the segmentation mask and the depth map corresponding to the frame (paragraph 9, line(s) 7-9, “the resolution level used for each voxel may be determined based on distance from the sensor, noise, semantics, and the like.”; also, paragraph 61, line(s) 1-3, “the resolution levels may be determined based on semantic labeling identifying object type (e.g., table, teapot, chair, vase, etc.).”), variable-resolution truncated signed distance function (TSDF) grid by fusing depth predictions from the depth maps corresponding to the plurality of frames (paragraph 8, line(s) 1-6, “The exemplary method further involves generating a first hash table storing 3D positions of a first set of voxels having a first resolution (e.g., big voxels) and signed distance values representing distances to the surfaces (e.g., to a nearest surface) of the physical environment based on the depth data.”; also, paragraph 8, line(s) 9-12, “the signed distance values include TSDF values that may be used to represent voxel distances of each voxel to a nearest surface of the surfaces of the physical environment.”), the variable-resolution TSDF grid comprising TSDF values indicating distance to a surface in the real-world environment, wherein the variable-resolution TSDF grid includes at least one portion at a first voxel resolution level and another portion at a second voxel resolution level of finer resolution than the first voxel resolution level (paragraph 5, line(s) 2-7, “Using multi-resolution voxels provides some portions of a reconstruction with smaller voxels to provide finer resolution and thus potentially higher accuracy and fidelity, and other portions of the reconstruction with larger voxels to provide coarser resolution and thus less accuracy and fidelity.”; also, paragraph 12, line(s) 13-16, “In some implementations, voxels of the first set of voxels have a first size and voxels of the second set of voxels have a second size, where the first size is larger than the second size.”), and wherein the voxel resolution at each portion of the variable-resolution TSDF grid is based on the level hints (paragraph 5, line(s) 12-16, “voxel size, e.g., which voxels are small and which voxels are large, may be determined using criteria that provides for the use of smaller voxels in areas where doing so will likely result in greater accuracy, e.g., where there is less noise in the data,”; also, paragraph 4, line(s) 12-15, “Those resolutions may be selected based upon distances of the voxels from the sensor, noise in the depth data associated with different voxels, or other factors.”); generating a polygon mesh from the variable-resolution TSDF grid digitally representing surfaces in the real-world environment captured by the image data (paragraph 6, line(s) 1-6, “Techniques disclosed herein may use the multi-resolution voxel data to generate a mesh that reconstructs the geometry of the physical environment. This may involve using a meshing algorithm that combines multi-resolution voxel information stored in multiple hash tables to generate a single mesh.”; also, paragraph 10, line(s) 11-15, “the mesh may be generated using a marching cubes meshing algorithm technique that identifies lines connecting points associated with the voxels in each hash table and interpolates to identify vertices along those lines that correspond to the surfaces.”); depth estimation model to each frame to output a depth map corresponding to the frame, wherein the depth map comprises depth predictions for pixels in the frame, wherein the level hints indicate a voxel resolution level for each pixel, and storing the polygon mesh in a map database However, in a similar field of endeavor Godard et al. discloses applying a depth estimation model to each frame to output a depth map corresponding to the frame, wherein the depth map comprises depth predictions for pixels in the frame (Godard: paragraph 30, line(s) 1-3, “The depth estimation model 130 receives an input image of a scene and outputs a depth of the scene based on the input image”; also, paragraph 6, line(s) 3-5, “The system inputs the images into a depth model to extract a depth map for each image based on parameters of the depth model.”), Armstrong et al. discloses the level hints indicate a voxel resolution level for each pixel (Armstrong: paragraph 24, line(s) 30-32, “the multi-resolution voxel space may have a plurality of semantic layers in which each semantic layer comprises a plurality of voxel grids representing voxels as covariance ellipsoids at different resolutions.”; also, paragraph 26, line(s) 10-12, “the semantic labels may comprise a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree, road surface, curb, sidewalk, unknown, etc.”) and storing the polygon mesh in a map database (Armstrong: paragraph 24, line(s) 28-29, “the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing using TSDF values from depth camera data with the features of Godard et al.'s invention of a learned depth estimation model that generates depth maps from input images, further in view of Armstrong et al.'s invention of multi-resolution voxel spaces with semantic layers for environment representation and storage. Regarding element “a non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the computer processor to perform operations comprising”, Meilland et al. explicitly teaches "a non-transitory computer readable storage medium has stored therein instructions, which, when executed by one or more processors" and that "the memory 220 includes a non-transitory computer readable storage medium." This directly corresponds to the claimed non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the processor to perform operations. Regarding element “receiving image data capturing a real-world environment and captured by a camera assembly of a client device, the image data comprising a plurality of frames”, Meilland et al. teaches "obtaining depth data of a physical environment using a sensor" where "the depth data can include pixel depth values from a viewpoint and sensor position and orientation data" using "one or more depth cameras" capable of multiple modalities. Under BRI, the one or more depth cameras constitute a camera assembly of a client device that captures image data of a real-world environment comprising a plurality of frames. Regarding element “applying a depth estimation model to each frame to output a depth map corresponding to the frame, wherein the depth map comprises depth predictions for pixels in the frame”, Godard et al. teaches "the depth estimation model receives an input image of a scene and outputs a depth of the scene based on the input image" and "the system inputs the images into a depth model to extract a depth map for each image based on parameters of the depth model." This directly teaches applying a learned depth estimation model to each frame to output a depth map comprising per-pixel depth predictions. Regarding element “applying a semantic segmentation model to each frame to output a segmentation mask corresponding to the frame, wherein the segmentation mask classifies pixels in the frame into one of a plurality of semantic classes”, Meilland et al. teaches "semantic labeling uses a machine learning model, where a semantic segmentation model may be configured to identify semantic labels for pixels or voxels of image data" and further teaches classifying scene elements by semantic type (e.g., "wall") for resolution determination. This teaches applying a semantic segmentation model to output a segmentation mask classifying pixels into semantic classes. Regarding element “determining level hints for each frame based on the segmentation mask and the depth map corresponding to the frame”, Meilland et al. teaches "the resolution level used for each voxel may be determined based on distance from the sensor, noise, semantics, and the like" and "the resolution levels may be determined based on semantic labeling identifying object type." This teaches determining resolution guidance based on both semantic information and depth-related information, corresponding to level hints based on both the segmentation mask and the depth map. Regarding element “wherein the level hints indicate a voxel resolution level for each pixel”, Armstrong et al. teaches multi-resolution voxel spaces with "a plurality of semantic layers in which each semantic layer comprises a plurality of voxel grids representing voxels as covariance ellipsoids at different resolutions" with semantic labels including "vehicle, pedestrian, cyclist, animal, building, tree, road surface, curb, sidewalk, unknown, etc." Under BRI, different resolution levels assigned per semantic layer teaches that the level hints indicate a voxel resolution level determinable per pixel. Regarding element “generating a variable-resolution truncated signed distance function (TSDF) grid by fusing depth predictions from the depth maps corresponding to the plurality of frames”, Meilland et al. teaches generating hash tables storing "3D positions of a first set of voxels having a first resolution and signed distance values representing distances to the surfaces of the physical environment based on the depth data" where "the signed distance values include TSDF values." This teaches generating a TSDF grid by fusing depth data from multiple frames. Regarding element “the variable-resolution TSDF grid comprising TSDF values indicating distance to a surface in the real-world environment, wherein the variable-resolution TSDF grid includes at least one portion at a first voxel resolution level and another portion at a second voxel resolution level of finer resolution than the first voxel resolution level”, Meilland et al. teaches multi-resolution voxels providing "some portions of a reconstruction with smaller voxels to provide finer resolution" and "other portions with larger voxels to provide coarser resolution" with voxels "of a first size" and voxels "of a second size, where the first size is larger." This teaches a variable-resolution TSDF grid with portions at different resolution levels. Regarding element “wherein the voxel resolution at each portion of the variable-resolution TSDF grid is based on the level hints”, Meilland et al. teaches that "voxel size may be determined using criteria" and "resolutions may be selected based upon distances of the voxels from the sensor, noise in the depth data, or other factors." This teaches voxel resolution based on resolution determination criteria (level hints). Regarding element “generating a polygon mesh from the variable-resolution TSDF grid digitally representing surfaces in the real-world environment captured by the image data”, Meilland et al. teaches "a meshing algorithm that combines multi-resolution voxel information stored in multiple hash tables to generate a single mesh" using "a marching cubes meshing algorithm technique." This teaches generating a polygon mesh from the variable-resolution TSDF grid. Regarding element “and storing the polygon mesh in a map database”, Armstrong et al. teaches "the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces" for "localization, object tracking, and/or navigation." This teaches storing the reconstruction in a persistent map database. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches multi-resolution voxel meshing using TSDF values from depth camera data and semantic labeling to determine voxel resolution, but relies on hardware depth sensors. Godard et al. teaches a trained depth estimation model that "receives an input image of a scene and outputs a depth of the scene" and extracts "a depth map for each image based on parameters of the depth model", providing a learned alternative that could directly replace Meilland et al.'s hardware depth cameras. Armstrong et al. teaches multi-resolution voxel spaces with "a plurality of semantic layers" at "different resolutions" and storing environment data for "localization, object tracking, and/or navigation". A POSITA would recognize that substituting Godard et al.'s learned depth model for hardware sensors eliminates the need for specialized depth cameras, and that Armstrong et al.'s storage and semantic layer architecture extends Meilland et al.'s reconstruction for downstream navigation use. Meilland et al. further teaches that "memory 220 includes a non-transitory computer-readable storage medium", and implementing the combined system as processor-executable instructions stored on such a medium is a standard design choice. Regarding claim 12, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 11, the operations further comprising: variable-resolution TSDF grid is further based on the tracked one or more objects in the real-world environment (Meilland et al.: paragraph 61, line(s) 1-3, “the resolution levels may be determined based on semantic labeling identifying object type (e.g., table, teapot, chair, vase, etc.). Meilland et al. as modified by Godard et al. and Armstrong et al. does not discloses applying an object detection model to each frame to identify one or more objects in the frame; and tracking one or more of the objects across frames. However, in a similar field of endeavor Armstrong et al. further discloses applying an object detection model to each frame to identify one or more objects in the frame (Armstrong et al.: paragraph 26, line(s) 10-12, “the semantic labels may comprise a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree, road surface, curb, sidewalk, unknown, etc.”; also, paragraph 12, line(s) 27-29, “map data represented by a multi-resolution voxel space may be generated from data points representing a physical environment, such as an output of a light detection and ranging (lidar) system.”); and tracking one or more of the objects across frames (Armstrong et al.: paragraph 30, line(s) 17-18, “to assist with localization, object tracking, and/or navigation of an autonomous vehicle with respect to the physical environment.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the features of Armstrong et al.'s object detection, semantic labeling, and object tracking. Regarding element “applying an object detection model to each frame to identify one or more objects in the frame”, Armstrong et al. teaches assigning semantic labels comprising "a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree, road surface, curb, sidewalk, unknown, etc." from sensor data including "an output of a light detection and ranging (lidar) system." Under BRI, identifying entities by class or type from sensor data constitutes applying an object detection model to each frame to identify one or more objects. Regarding element ”tracking one or more of the objects across frames”, Armstrong et al. teaches that the system assists with "localization, object tracking, and/or navigation of an autonomous vehicle," directly teaching tracking one or more objects across frames. Regarding element “generating the variable-resolution TSDF grid is further based on the tracked one or more objects in the real-world environment”, Meilland et al. teaches that "resolution levels may be determined based on semantic labeling identifying object type" and supports "detection, tracking, and representing of objects in 3D space." This teaches that the variable-resolution TSDF generation is based on detected and tracked objects. One of ordinary skill in the art would have been motivated to combine because Armstrong et al. teaches assigning semantic labels comprising "a class or an entity type, such as vehicle, pedestrian, cyclist, animal, building, tree" and assisting with "object tracking", while Meilland et al. already teaches that "resolution levels may be determined based on semantic labeling identifying object type" and supports "detection, tracking, and representing of objects in 3D space." A POSITA working with Meilland et al.'s semantic-based resolution system would naturally incorporate Armstrong et al.'s object detection and tracking to maintain consistent object identity across frames, because knowing what objects are present and tracking them over time directly improves the resolution allocation that Meilland et al. already performs based on object type. Regarding claim 13, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 12, wherein the object detection model is trained as a machine-learning model (Meilland et al.: paragraph 61, line(s) 16-18, “the machine learning model is a neural network (e.g., an artificial neural network), decision tree, support vector machine”; also, paragraph 61, line(s) 4-7, “semantic labeling uses a machine learning model, where a semantic segmentation model may be configured to identify semantic labels for pixels or voxels of image data.”) in It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces with object tracking, with the features of supervised machine learning training. Regarding element “the object detection model is trained as a machine-learning mode”, Meilland et al. teaches that "the machine learning model is a neural network (e.g., an artificial neural network), decision tree, support vector machine" and that "semantic labeling uses a machine learning model." This teaches that the object detection model is a machine-learning model such as a neural network. Regarding element “a supervised manner with training image data labeled with identified objects”, Godard et al. teaches that "a depth estimation system may be trained using a detection and ranging system to establish a ground truth depth for objects in an environment (i.e., radio detecting and ranging (RADAR), light detection and ranging (LIDAR), etc.) paired with images taken of the same scene by a camera." This teaches the supervised training paradigm of pairing labeled ground truth data with corresponding images. A POSITA would understand that applying this same supervised training paradigm to train Meilland et al.'s neural network-based object detection model with labeled training image data identifying objects is the standard approach for such models. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches that "the machine learning model is a neural network" and that "semantic labeling uses a machine learning model", but does not disclose how these models are trained. Godard et al. teaches that "a depth estimation system may be trained using a detection and ranging system to establish a ground truth depth for objects in an environment... paired with images taken of the same scene by a camera", which is the supervised training paradigm of pairing labeled ground truth data with corresponding images. A POSITA implementing Meilland et al.'s neural network for object detection would naturally apply the same supervised training paradigm taught by Godard et al., because training a neural network to identify objects requires labeled training data with identified objects, and this is the standard methodology for such models in computer vision. Regarding claim 14, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 1, wherein variable-resolution TSDF grid is further based on the one or more surface orientations of the one or more surfaces (Meilland et al.: paragraph 5, line(s) 12-16, “voxel size, e.g., which voxels are small and which voxels are large, may be determined using criteria that provides for the use of smaller voxels in areas where doing so will likely result in greater accuracy, e.g., where there is less noise in the data,”; also, paragraph 4, line(s) 12-15, “Those resolutions may be selected based upon distances of the voxels from the sensor, noise in the depth data associated with different voxels, or other factors.”). Meilland et al. as modified by Godard et al. and Armstrong et al. does not discloses applying the depth estimation model to each frame further comprises applying the depth estimation model to identify surface orientation of one or more surfaces present in the frame. However, in a similar field of endeavor, Godard et al. further discloses applying the depth estimation model to each frame further comprises applying the depth estimation model to identify surface orientation of one or more surfaces present in the frame (Godard: paragraph 6, line(s) 3-5, “The system inputs the images into a depth model to extract a depth map for each image based on parameters of the depth model.”; also, paragraph 6, line(s) 20-22, “Upsampled depth features may also be used during generation of the synthetic frames which would affect the appearance matching loss calculations.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the features of surface orientation estimation for resolution guidance. Regarding element “applying the depth estimation model to each frame further comprises applying the depth estimation model to identify surface orientation of one or more surfaces present in the frame”, Godard et al. teaches "the system inputs the images into a depth model to extract a depth map for each image" and "upsampled depth features may also be used during generation of the synthetic frames which would affect the appearance matching loss calculations." The depth features extracted by the model inherently encode geometric information including surface orientation. Under BRI, a depth estimation model that produces per-pixel depth maps implicitly captures surface orientation, as surface normals are computable from spatial gradients of the depth map. Extending such a model to explicitly output surface orientation is a routine design choice. Regarding element “generating the variable-resolution TSDF grid is further based on the one or more surface orientations of the one or more surfaces”, Meilland et al. teaches that "voxel size may be determined using criteria that provides for the use of smaller voxels in areas where doing so will likely result in greater accuracy" and "resolutions may be selected based upon distances of the voxels from the sensor, noise in the depth data, or other factors." Surface orientation is a geometrically relevant factor for resolution determination — surfaces at oblique angles produce noisier depth estimates — and falls within Meilland et al.'s "other factors" category. One of ordinary skill in the art would have been motivated to combine because Godard et al. teaches that "the system inputs the images into a depth model to extract a depth map for each image" and that "upsampled depth features may also be used during generation of the synthetic frames", meaning the depth model extracts rich geometric features including surface information. Meilland et al. teaches that voxel resolution is selected based on "distances of the voxels from the sensor, noise in the depth data associated with different voxels, or other factors". A POSITA would recognize that Godard et al.'s depth features inherently encode surface orientation information (as surface normals are computable from depth gradients), and that surface orientation is one of Meilland et al.'s "other factors" affecting depth noise. Extending Godard et al.'s depth model to also output surface orientation and feeding that into Meilland et al.'s resolution selection would predictably improve reconstruction quality for oblique surfaces. Regarding claim 15, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 1, wherein depth estimation model is trained as a machine-learning model in a self-supervised manner by projecting frames from training image data onto other frames of the training image data based on depth predictions by the depth estimation model. However, in a similar field of endeavor, Godard et al. further discloses the depth estimation model is trained as a machine-learning model in a self-supervised manner by projecting frames from training image data onto other frames of the training image data based on depth predictions by the depth estimation model (Godard: paragraph 5, line(s) 5-10, “Some more novel methods train a depth estimation system utilizing monocular video data of an ever changing scene. The depth estimation system trains by projecting from one temporal image in the monocular video data to a subsequent temporal image while minimizing a photometric reconstruction error.”; also, abstract, line(s) 6-8, “The method includes generating a plurality of synthetic frames based on the depth map and the pose for each image”; also, paragraph 6, line(s) 13-16, “The loss function includes a calculation of the photometric reconstruction error per pixel between a synthetic frame and an input image”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the features of Godard et al.'s self-supervised depth estimation training methodology. Regarding element “the depth estimation model is trained as a machine-learning model in a self-supervised manner by projecting frames from training image data onto other frames of the training image data based on depth predictions by the depth estimation model”, Godard et al. explicitly teaches self-supervised training, disclosing that "some more novel methods train a depth estimation system utilizing monocular video data of an ever changing scene" and that "the depth estimation system trains by projecting from one temporal image in the monocular video data to a subsequent temporal image while minimizing a photometric reconstruction error." Godard et al. further teaches that "the method includes generating a plurality of synthetic frames based on the depth map and the pose for each image" and that "the loss function includes a calculation of the photometric reconstruction error per pixel between a synthetic frame and an input image." This directly teaches the claimed self-supervised training methodology — the depth estimation model is trained by projecting frames from training image data (monocular video) onto other frames based on the depth map (depth predictions) and pose, with the photometric error serving as the self-supervised loss signal without ground truth depth labels. One of ordinary skill in the art would have been motivated to combine because Godard et al. teaches that "the depth estimation system trains by projecting from one temporal image in the monocular video data to a subsequent temporal image while minimizing a photometric reconstruction error", that "the method includes generating a plurality of synthetic frames based on the depth map and the pose for each image," and that "the loss function includes a calculation of the photometric reconstruction error per pixel between a synthetic frame and an input image." This self-supervised training approach eliminates the need for ground truth depth labels. A POSITA seeking to train the depth estimation model used in the combined system would naturally adopt Godard et al.'s self-supervised methodology because it enables training on the same monocular video data that Meilland et al.'s pipeline already processes, without requiring any additional labeled datasets. Regarding claim 16, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 11, wherein generating the variable-resolution TSDF grid is constrained by limiting neighboring voxel cells to be at most one voxel resolution level difference (Meilland: paragraph 11, line(s) 8-11, “Additional hash tables can be generated for each level of resolution desired. In an exemplary implementation, at least four hash tables are utilized for four different resolutions”; also, paragraph 10, line(s) 3-7, “The mesh may be generated by positioning vertices along a line connecting a first voxel (e.g., a position at the center of the first voxel) of the first set of voxels with a second voxel (e.g., a position at the center of the second voxel) of the second set of voxels.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the constraint of limiting neighboring voxel cells to at most one resolution level difference. Regarding element “generating the variable-resolution TSDF grid is constrained by limiting neighboring voxel cells to be at most one voxel resolution level difference”, Meilland et al. teaches using "at least four hash tables for four different resolutions" and generates meshes by connecting voxels across resolution levels. Constraining neighboring voxels to at most one resolution level difference (2:1 balancing) is a well-established technique in adaptive mesh refinement. A POSITA working with Meilland et al.'s multi-resolution system would naturally apply this standard constraint to ensure mesh quality at resolution transitions. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches using "at least four hash tables for four different resolutions" and generating meshes by "positioning vertices along a line connecting a first voxel of the first set of voxels with a second voxel of the second set of voxels". When connecting voxels across multiple resolution levels, a POSITA would recognize that allowing arbitrary resolution jumps between neighboring voxels creates geometric discontinuities and mesh artifacts at transition boundaries. Constraining neighboring voxel cells to at most one resolution level difference (2:1 balancing) is a well-established technique in adaptive mesh refinement that directly addresses this problem, and applying it to Meilland et al.'s four-level system would predictably produce smoother, higher-quality meshes at resolution transitions. Regarding claim 17, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 11, wherein generating the variable-resolution TSDF grid comprises implementing a hyperparameter that sets a quantity of depth predictions fused into the TSDF value per voxel, wherein the hyperparameter is fit to an error curve for depth predictions by the depth estimation model (Meilland et al.: paragraph 8, line(s) 12-15, “The TSDF values can save storage space by including only values within a truncation band in the representation, e.g., only storing data for voxels that are within a threshold distance of a surface.”; also, paragraph 59, line(s) 15-18, “Distance may be used as an approximation of noise based on a correlation (e.g., quadratic noise with respect to distance such that farther distance means more noise).”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with a hyperparameter controlling fusion quantity based on depth model error characteristics. Regarding element “generating the variable-resolution TSDF grid comprises implementing a hyperparameter that sets a quantity of depth predictions fused into the TSDF value per voxel, wherein the hyperparameter is fit to an error curve for depth predictions by the depth estimation model”, Meilland et al. teaches TSDF threshold parameters ("values within a truncation band") and that "distance may be used as an approximation of noise based on a correlation (e.g., quadratic noise with respect to distance such that farther distance means more noise)." This teaches modeling error characteristics and using them to inform TSDF parameters. Under BRI, fitting a fusion quantity hyperparameter to the depth model's error curve is an extension of Meilland et al.'s noise-distance correlation teaching, though the specific implementation represents a more targeted optimization. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches that "TSDF values can save storage space by including only values within a truncation band" and that "distance may be used as an approximation of noise based on a correlation (e.g., quadratic noise with respect to distance)". This shows Meilland et al. already models depth error as a function of distance and uses error characteristics to set TSDF parameters. When integrating Godard et al.'s learned depth estimation model into this pipeline, a POSITA would naturally calibrate the fusion quantity hyperparameter to the depth model's actual error curve, because the number of depth predictions needed per voxel depends on how noisy the model's predictions are at that distance. Regarding claim 18, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 11, wherein generating the polygon mesh from the variable-resolution TSDF grid comprises interpolating between neighboring voxels of different voxel resolution levels (Meilland et al.: paragraph 10, line(s) 3-7, “The mesh may be generated by positioning vertices along a line connecting a first voxel (e.g., a position at the center of the first voxel) of the first set of voxels with a second voxel (e.g., a position at the center of the second voxel) of the second set of voxels.”; also, paragraph 10, line(s) 22-28, “generating the mesh includes generating lines connecting points associated with the voxels in each of the multiple hash tables (e.g., between the first hash table and the second hash table, between the second hash table and the third hash table, etc.) and interpolating along the lines to identify vertices for the mesh that correspond to the surfaces.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with Meilland et al.'s interpolation technique for cross-resolution mesh generation. Regarding element “wherein generating the polygon mesh from the variable-resolution TSDF grid comprises interpolating between neighboring voxels of different voxel resolution levels”, Meilland et al. explicitly teaches generating a mesh by "positioning vertices along a line connecting a first voxel of the first set of voxels with a second voxel of the second set of voxels" across different resolutions, and "interpolating along the lines to identify vertices for the mesh that correspond to the surfaces" between hash tables of different resolution levels. This directly teaches interpolating between neighboring voxels of different voxel resolution levels. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches generating a mesh by "positioning vertices along a line connecting a first voxel of the first set of voxels with a second voxel of the second set of voxels" and "interpolating along the lines to identify vertices for the mesh that correspond to the surfaces" specifically "between the first hash table and the second hash table, between the second hash table and the third hash table, etc.". This interpolation between voxels of different resolution levels is already an integral part of Meilland et al.'s own multi-resolution meshing approach. Applying this technique within the combined system's variable-resolution TSDF grid would produce a continuous, seamless polygon mesh without gaps or discontinuities at resolution boundaries. Regarding claim 19, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 1, the operations further comprising: augmenting the polygon mesh with visual patterns from the image data corresponding to one or more surfaces represented by the polygon mesh (Meilland et al.: paragraph 4, line(s) 1-6, “Various implementations disclosed herein include devices, systems, and methods that generate a mesh (e.g., vertices that form connected triangles) representing the surfaces in a physical environment using multi-resolution voxels that are generated based on detected depth information.”; also, paragraph 14, line(s) 1-5, “In some implementations, the depth data is obtained using one or more depth cameras. For example, the one or more depth cameras can acquire depth based on structured light (SL), passive stereo (PS), active stereo (AS), time of flight (ToF), and the like”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with the augmentation of the polygon mesh with image-derived visual patterns. Regarding element “augmenting the polygon mesh with patterns from the image data corresponding to one or more surfaces represented by the polygon mesh”, Meilland et al. teaches generating "a mesh (e.g., vertices that form connected triangles) representing the surfaces in a physical environment" from depth data obtained by "one or more depth cameras" that can acquire data through modalities including passive stereo and active stereo, which inherently capture RGB image data alongside depth. Under BRI, augmenting the polygon mesh with patterns from the image data encompasses texture mapping, which is a fundamental and well-known technique in 3D reconstruction. A POSITA would readily apply texture mapping to Meilland et al.'s mesh output using the captured image data once camera poses are known. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches generating "a mesh (e.g., vertices that form connected triangles) representing the surfaces in a physical environment" from depth data obtained using depth cameras that also capture image data. Once the polygon mesh is generated and camera poses are known from the image data, projecting the original image data onto the mesh surfaces to add visual texture is one of the most fundamental techniques in 3D reconstruction. A POSITA would naturally augment the bare geometric mesh with patterns from the already-available image data, producing a photorealistic 3D model far more useful for the navigation and visualization applications taught by Armstrong et al. Regarding claim 20, Meilland et al. as modified by Godard et al. and Armstrong et al. discloses the non-transitory computer-readable storage medium of claim 11, the operations further comprising: receiving a request from a second client device to view the polygon mesh (Meilland et al.: paragraph 31, line(s)13-15, “the controller 110 is communicatively coupled with the device 120 via one or more wired or wireless communication channels”; also, paragraph 31, line(s) 10-12, “the controller 110 is a remote server located outside of the physical environment”); polygon mesh from the map database; and transmitting the polygon mesh to the second client device for presentation on the second client device. However, in a similar field of endeavor, Armstrong et al. further discloses the retrieving the polygon mesh from the map database; and transmitting the polygon mesh to the second client device for presentation on the second client device (Armstrong et al.: paragraph 24, line(s) 28-29, “the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces”; also, paragraph 30, line(s) 17-18, “to assist with localization, object tracking, and/or navigation of an autonomous vehicle with respect to the physical environment”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Meilland et al.'s invention of multi-resolution voxel meshing, as combined with Godard et al.'s depth estimation model and Armstrong et al.'s multi-resolution voxel spaces, with client-server architecture for mesh retrieval and transmission. Regarding element “receiving a request from a second client device to view the polygon mesh”, Meilland et al. teaches "the controller 110 is communicatively coupled with the device 120 via one or more wired or wireless communication channels" and "the controller 110 is a remote server located outside of the physical environment." This teaches a client-server architecture where devices communicate with a remote server, and under BRI, receiving a request from a second client device is a standard client-server interaction. Regarding element “retrieving the polygon mesh from the map database; and transmitting the polygon mesh to the second client device for presentation on the second client device”, Armstrong et al. teaches "the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces" for "localization, object tracking, and/or navigation of an autonomous vehicle." A system that stores environment data for autonomous vehicles necessarily retrieves and transmits that data to requesting vehicles for navigation. This is the standard architecture for map services. One of ordinary skill in the art would have been motivated to combine because Meilland et al. teaches that "the controller 110 is communicatively coupled with the device 120 via one or more wired or wireless communication channels" and that "the controller 110 is a remote server located outside of the physical environment", establishing a client-server architecture. Armstrong et al. teaches storing "the scene as well as data representative of environment as multi-resolution voxel spaces" for "localization, object tracking, and/or navigation of an autonomous vehicle". A POSITA working with this client-server architecture and stored environment data would naturally extend it to serve the polygon mesh to additional client devices upon request, because the whole purpose of Armstrong et al.'s stored environment data is to assist multiple autonomous vehicles with localization and navigation, which inherently requires retrieving and transmitting the data to those vehicles. Response to Arguments Applicant's arguments filed 08/18/2026 have been fully considered but they are not persuasive. Applicant argues that Meilland does not teach "determining level hints for each frame based on the segmentation mask and the depth map corresponding to the frame," because the cited portion of Meilland describes determining a voxel resolution level directly from semantics or distance from the sensor, whereas claim 1 requires a two step process in which the level hints are first determined as an intermediate value derived from the segmentation mask and the depth map and are only subsequently used to generate the variable-resolution TSDF grid. The argument does not address the rejection as made. Applicant states that the Office Action relies on Meilland at paragraph 9. The rejection of claim 1 cited Meilland paragraph 9 and Meilland paragraph 67 for the level hints limitation, and Applicant has not addressed the citation to paragraph 67. Meilland teaches that "the resolution levels may be determined based on semantic labeling identifying object type (e.g., table, teapot, chair, vase, etc.)", and further teaches that "semantic labeling uses a machine learning model, where a semantic segmentation model may be configured to identify semantic labels for pixels or voxels of image data”. Meilland therefore determines resolution levels from semantic labels that a semantic segmentation model identifies for pixels of the image data. Applicant's characterization of the recited level hints is also not commensurate in scope with the claim language. Claim 1 recites "wherein the level hints indicate a voxel resolution level for each pixel." The claim states what a level hint indicates, and what it indicates is a voxel resolution level. Applicant's assertions that the level hints "are not themselves the assignment of voxel resolution" and that they are "independent of" that assignment do not appear in the claim. Claim 1 further recites that "the voxel resolution at each portion of the variable-resolution TSDF grid is based on the level hints," which requires the assignment of voxel resolution to depend on the level hints rather than to be independent of them. Limitations appearing in the specification but not recited in the claim are not read into the claim. Meilland further performs the sequence that claim 1 recites. Meilland determines the resolution level, teaching that "the resolution level used for each voxel may be determined based on distance from the sensor, noise, semantics, and the like", and thereafter generates the voxel representation at the determined resolutions, teaching that "The exemplary method further involves generating a first hash table storing 3D positions of a first set of voxels having a first resolution (e.g., big voxels) and signed distance values representing distances to the surfaces (e.g., to a nearest surface) of the physical environment based on the depth data". The determination of resolution level in Meilland is an act separate from, and antecedent to, the generation of the voxel representation that is built at those resolutions. Applicant's premise that Meilland performs a single step is therefore not supported by Meilland. Applicant argues that Armstrong at paragraphs 24 and 26 does not teach "wherein the level hints indicate a voxel resolution level for each pixel," because those paragraphs describe an architecture for encoding semantic information across multiple resolution layers of the same space rather than a level hint calculated for a pixel, and because Armstrong is silent as to determining a hint value based on a segmentation mask and a depth map (Remarks, page 12). This argument is not persuasive. The argument is directed to Armstrong individually. The limitation is taught by Meilland. As set forth above, Meilland teaches that a "semantic segmentation model may be configured to identify semantic labels for pixels or voxels of image data" and that "the resolution levels may be determined based on semantic labeling identifying object type (e.g., table, teapot, chair, vase, etc.)". A resolution level that is determined from a semantic label identified for a pixel is a voxel resolution level indicated for that pixel, which is what claim 1 requires the level hints to indicate. Armstrong's disclosure is cumulative with respect to this clause. Armstrong remains relied upon for "storing the polygon mesh in a map database," for which Armstrong teaches that "the system may be configured to store the scene as well as data representative of environment as multi-resolution voxel spaces". Applicant argues that there is no adequate articulated reason to combine, because Meilland's multi-resolution voxel space is deployed to generate a single-resolution TSDF value at each voxel for mesh reconstruction while Armstrong's multi-resolution voxel space is deployed to encode semantic class information across semantic layers for localization, object tracking, or navigation, and because importing Armstrong's semantic-layer encoding structure into Meilland would not yield the claimed level hints (Remarks, pages 12-13). This argument is not persuasive. The argument rests on the premise that neither reference teaches the recited level hints. That premise is addressed above. The level hints are taught by Meilland, and the combination is not required to produce them by importing Armstrong's semantic layer structure into Meilland. Armstrong is relied upon for storing the environment representation as a map database for downstream use, and the reason to add that teaching to Meilland's reconstruction does not depend on Armstrong supplying the level hints. Applicant's characterization of Meilland as generating "a single-resolution TSDF value at each voxel" is also not consistent with what Meilland teaches. Meilland teaches that "Using multi-resolution voxels provides some portions of a reconstruction with smaller voxels to provide finer resolution and thus potentially higher accuracy and fidelity, and other portions of the reconstruction with larger voxels to provide coarser resolution and thus less accuracy and fidelity". Meilland's reconstruction is itself of differing resolution across its portions, which is the arrangement claim 1 recites for the variable-resolution TSDF grid. Applicant argues that independent claim 11 recites commensurate limitations and is allowable for the same reasons as claim 1, and that claims 2-10 and 12-20 are allowable at least by virtue of their dependency from claims 1 and 11 (Remarks, page 13). These arguments are not persuasive. Claim 11 recites the level hints limitation in the same terms as claim 1 and is rejected on the same teachings, for the reasons given above. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jai Li whose telephone number is (571)272-1170. The examiner can normally be reached Mon-Thu between 06:00-16:00 EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571)272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAI W LI/Junior Patent Examiner, Art Unit 2613 /XIAO M WU/Supervisory Patent Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Oct 17, 2024
Application Filed
Apr 20, 2026
Non-Final Rejection mailed — §103
Aug 18, 2026
Response Filed
Aug 31, 2026
Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
Grant Probability
Moderate
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month