Prosecution Insights
Last updated: October 02, 2026
Application No. 19/192,277

THREE-DIMENSIONAL ASSET RECONSTRUCTION

Non-Final OA §103§DOUBLEPATENT
Filed
Apr 28, 2025
Priority
Sep 12, 2022 — continuation of 12/333,650
Examiner
CHIN, MICHELLE
Art Unit
Tech Center
Assignee
Snap Inc.
OA Round
1 (Non-Final)
85%
Grant Probability
Favorable
1-2
OA Rounds
9m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
560 granted / 656 resolved
+25.4% vs TC avg
Moderate +12% lift
Without
With
+11.6%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
25 currently pending
Career history
677
Total Applications
across all art units

Statute-Specific Performance

§101
9.3%
-30.7% vs TC avg
§103
71.0%
+31.0% vs TC avg
§102
5.5%
-34.5% vs TC avg
§112
1.7%
-38.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 656 resolved cases

Office Action

§103 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement 2. The information disclosure statement (IDS) submitted on 07/24/2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Double Patenting 3. The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory obviousness-type double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b). 4. Claims 1-20 are rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-20 of Patent No. 12,333,650. Although the conflicting claims are not identical, they are not patentably distinct from each other because the instant application claims are broader in every aspect than the patent claims and are therefore an obvious variant thereof. 5. Regarding claim 1, the application claim discloses An asset reconstruction system comprising: a plurality of statically positioned cameras, the cameras configured and positioned to capture images of an object from a plurality of viewpoints surrounding the object; at least one light source configured to illuminate the object, each light source having a respective location; one or more computer processors; and a memory storing instructions that, when executed by the one or more computer processors, cause the asset reconstruction system to perform operations for generating a three-dimensional (3D) asset representing the object, the operations comprising: capturing a set of images of the object from the plurality of viewpoints surrounding the object with the plurality of statically positioned cameras, estimating camera poses for each respective image of the set of images, constructing a 3D surface mesh comprising a plurality of surfaces using the set of images and the camera poses estimated for each respective image, and optimizing texture properties of the plurality of surfaces of the 3D surface mesh to generate the 3D asset. Claim 1 of Patent No. 12,333,650 discloses An asset reconstruction system comprising: a plurality of cameras statically positioned within a darkroom, the cameras configured and positioned to capture images of an object positioned within the darkroom from a plurality of viewpoints surrounding the object; at least one light source configured to illuminate the object, each light source having a respective location; one or more computer processors; and a memory storing instructions that, when executed by the one or more computer processors, cause the asset reconstruction system to perform operations for generating a three-dimensional (3D) asset representing the object, the operations comprising: capturing a set of images of the object from the plurality of viewpoints surrounding the object with the plurality of cameras statically positioned within the darkroom; estimating camera poses for each respective image of the set of images; constructing a 3D surface mesh comprising a plurality of surfaces using the set of images and the camera poses estimated for each respective image; and optimizing texture properties of the plurality of surfaces of the 3D surface mesh to generate the 3D asset. Regarding claim 1, the only difference is that claim 1 of the instant application does not recite “positioned within the darkroom” while claim 1 of Patent No. 12,333,650 recites. Regarding claims 10 and 19, the analyses are similar to that of claim 1; the rationale of claim 1 rejection is applied in rejecting claims 10 and 19. Therefore, the claims in the present application recite a broader scope than the claims in the Patent No. 12,333,650. 6. The following table shows the claims of the current application being examined and the conflicting claims of Patent No. 12,333,650. Current Application No. 19/192,277 Patent No. 12,333,650 1 1 2-9 2-9 10 10 11-18 11-18 19 19 20 20 The following table shows an example of the corresponding conflicting claims of the current application and Patent No. 12,333,650. Current Application No. 19/192,277 Claim 1 Patent No. 12,333,650 Claim 1 An asset reconstruction system comprising: An asset reconstruction system comprising: a plurality of statically positioned cameras, the cameras configured and positioned to capture images of an object from a plurality of viewpoints surrounding the object; a plurality of cameras statically positioned within a darkroom, the cameras configured and positioned to capture images of an object positioned within the darkroom from a plurality of viewpoints surrounding the object; at least one light source configured to illuminate the object, each light source having a respective location; one or more computer processors; and a memory storing instructions that, when executed by the one or more computer processors, cause the asset reconstruction system to perform operations for generating a three-dimensional (3D) asset representing the object, the operations comprising: at least one light source configured to illuminate the object, each light source having a respective location; one or more computer processors; and a memory storing instructions that, when executed by the one or more computer processors, cause the asset reconstruction system to perform operations for generating a three-dimensional (3D) asset representing the object, the operations comprising: capturing a set of images of the object from the plurality of viewpoints surrounding the object with the plurality of statically positioned cameras, capturing a set of images of the object from the plurality of viewpoints surrounding the object with the plurality of cameras statically positioned within the darkroom; estimating camera poses for each respective image of the set of images, estimating camera poses for each respective image of the set of images; constructing a 3D surface mesh comprising a plurality of surfaces using the set of images and the camera poses estimated for each respective image, and optimizing texture properties of the plurality of surfaces of the 3D surface mesh to generate the 3D asset. constructing a 3D surface mesh comprising a plurality of surfaces using the set of images and the camera poses estimated for each respective image; and optimizing texture properties of the plurality of surfaces of the 3D surface mesh to generate the 3D asset. Claim Rejections - 35 USC § 103 7. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 8. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 9. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 10. Claim(s) 1, 5, 6, 10, 14, 15 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over von Cramon (US 2022/0092849 A1) in view of Mullins (US 2017/0193686 A1). 11. With reference to claim 1, von Cramon teaches An asset reconstruction system comprising: a plurality of statically positioned cameras, the cameras configured and positioned to capture images of an object (“In generating photorealistic models of complex real world environments the image capture system is charged with providing a post-processing workflow digital assets containing data that is both useful to a photogrammetry engine employed to solve for geometry in scene reconstruction and provide usable texture data allowing a lighting and rendering engine to realistically simulate the diffuse color and specular reflectance of surface materials. “ [0019] “The phrase “image-processing system” as used herein means an apparatus that receives images however illuminated of substantially the same scene and uses one or more repetitive and/or iterative data manipulations to generate a modified image or another data asset that is stored for use in generating a representation of a scene from a model in conjunction with information from the images” [0051] “an image-capture system could be arranged with paired cameras. In such an arrangement a single camera orientation would apply to the image pairs and would provide optimal inputs for a difference blend operation to isolate specular reflections from diffuse color. A single emitter could be used in conjunction with a film polarizer to illuminate a subject-of-interest with polarized light. A first camera may receive the reflected light after it is further redirected by a beam splitter. A second or “through-path” camera is provided after the beam splitter.” [0087] “a series of images captured from a fixed location and orientation while a light traverses the field of view such that with each exposure the light source is in a new position along a path provides an additional stream of image information that can be coupled with the paired images for training CNNs,” [0097]) von Cramon also teaches at least one light source configured to illuminate the object, each light source having a respective location; (“The phrase “virtual light source” as used herein means information related to one or more of relative location from an origin, luminous flux, and frequency that changes one or more characteristics of the subject matter in the virtual environment.” [0077] “When a light source moves in the real world an observer sees shadows and specular reflections change accordingly. Similarly, when an observer moves in the real world, specular reflections and in the case of partially translucent materials, subsurface scatter changes from the perspective of the observer. Accordingly, it is a benefit when moving a virtual light in a virtual environment for an observer to see shadows and specular reflections shift in accordance with changes in the location and orientation of the virtual light source. Likewise, when the perspective of the virtual observer is changing it is further beneficial for specular reflections, and in the case of translucent materials, for the behaviors of subsurface scatter to change in the virtual representation.” [0080] “The surface or surfaces of interest at a location to be modeled are illuminated by a light source that provides sufficient light under different polarization states to adequately expose photosensitive elements in an image sensor.” [0082]) von Cramon further teaches one or more computer processors; and a memory storing instructions that, when executed by the one or more computer processors, cause the asset reconstruction system to perform operations for generating a three-dimensional (3D) asset representing the object, (“a surface construction or reconstruction process has been performed. Such a surface reconstruction uses the locations defined by the points of the point cloud 210 to define a four-sided polygon or quadrilateral. Alternative surface reconstruction algorithms may use three points from the point cloud or other collections of points greater in number to represent surfaces of features in a real-world scene 200. However, surfaces represented by triangles and quadrilaterals are generally preferred. The closed areas of sub-portions of a polygonal mesh 215 are often associated with a two-dimensional unfolded version of the corresponding surface geometry. These two-dimensional representations are commonly called UV maps. The letters “U” and “V” denote axes of a two-dimensional texture. When matched or projected with appropriate color and relatively finer texture information in proper registration with the surface geometry over the entirety of the surfaces in the polygonal mesh 215 a three-dimensional color model of the real-world scene 200 is created.” [0143] “FIG. 4A is a schematic diagram of an embodiment of the image-capture device 400 of FIG. 3. As illustrated, the image-capture device 400 is an assembly of subsystems including an illumination source 410, illumination controller 420, an optional scanner subsystem 425, optical subsystem 430, shutter 440, processor 450 and memory 460.” [0156]) von Cramon teaches the operations comprising: capturing a set of images of the object with the plurality of statically positioned cameras, (“an image-capture system could be arranged with paired cameras. In such an arrangement a single camera orientation would apply to the image pairs and would provide optimal inputs for a difference blend operation to isolate specular reflections from diffuse color. A single emitter could be used in conjunction with a film polarizer to illuminate a subject-of-interest with polarized light. A first camera may receive the reflected light after it is further redirected by a beam splitter. A second or “through-path” camera is provided after the beam splitter.” [0087] “the present image-capturing techniques and processing methods can be applied in conjunction with structured light, sonar (sound navigation and ranging), LiDAR (a portmanteau of “light” and “radar”), light field camera technology, and other scanning methods to leverage camera projection mapping to produce information models to support the creation of more realistic virtual environments that adapt to changes in point of view, changes in position of a virtual or CG light source and for some environments changes in position of the sun.” [0092] “a series of images captured from a fixed location and orientation while a light traverses the field of view such that with each exposure the light source is in a new position along a path provides an additional stream of image information that can be coupled with the paired images for training CNNs,” [0097]) von Cramon also teaches estimating camera poses for each respective image of the set of images, constructing a 3D surface mesh comprising a plurality of surfaces using the set of images and the camera poses estimated for each respective image, (“The present image-capturing techniques can be used to forward a set of diffuse albedo surface textures to a photogrammetry engine to generate a dense surface mesh, which after post-processing delivers a render mesh. The render mesh includes a three-dimensional model of the geometry of the subject matter captured in the images and a set of UV maps. The render mesh is used with camera orientation information and the surface textures to create corresponding diffuse albedo projection maps. The render mesh and the diffuse albedo projection maps are inputs that can be used by an image processor to create a three-dimensional color representation of the subject matter captured in the images.” [0091] “To relate datasets between imagery exposed using a repositioned light source as captured from a fixed image sensor with imagery separately captured using the above-described co-polarized and cross-polarized exposures, a most direct correlation implies a shared camera orientation. That is, while a workflow that collects the first and second data sets is characterized by handheld photography, or photographic collection from any free-standing chassis, comparing changes in per pixel values within a sequence of photos of substantially the same scene illuminated with a moving light source separate from the image sensor implies the need for a fixed perspective, such as from a tripod. By placing the previously described image capture system on a tripod it is possible to more effectively train CNNs when co-polarized and cross-polarized exposures of a scene provide nearly identical rasters using the same camera to record reflections from surfaces or objects of interest present in the scene. In a given environment, to ensure anything between believability and comprehensive coverage, a technician would first navigate the space taking an inventory of all materials present. From a first location a given perspective may optimize for taking in as broad a sampling of materials present, these seen from an angle to optimize for recovering reflectance properties. While Li et al. features camera angles perpendicular to largely flat surfaces, such as ceramic tiles or the side of figurines, wider views of more complex scenes may require material acquisition from multiple views of the same set of materials present to provide critical mass in terms of the quality of texture information used to train CNNs, as defined by ideal camera angles relative to surfaces.” [0098]) PNG media_image1.png 399 540 media_image1.png Greyscale von Cramon does not explicitly teach capture images of an object from a plurality of viewpoints surrounding the object; This is what Mullins teaches (“Reconstruction of a three-dimensional model of an animated physical object typically requires multiple cameras that are statically positioned at predefined locations around the physical object.” [0023]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Mullins into von Cramon, in order to facilitate accurate photogrammetric reconstruction and texture mapping. 12. With reference to claim 5, von Cramon teaches the texture properties include one or more of reflectance, color, or roughness. (“The phrase “specular-surface texture” as used herein means a two-dimensional data set that includes specular color from one or more co-polarized exposures or non-polarized exposures.” [0072] “The present methods combine substantially shadow-free lighting with photography to capture surface textures that isolate diffuse color data from specular color data. A set of diffuse albedo surface textures are used with conventional photogrammetry techniques to generate a model of a real-world location or scene. Matched images or images of substantially the same subject matter exposed under different lighting conditions are used to generate a modified image. This modified image or specular roughness surface texture is used as a separate input when rendering a virtual environment from the model. Accordingly, a set of exposures captured at a location are temporarily stored as image files and processed using an image-processing technique to generate the modified image.” [0081]) 13. With reference to claim 6, von Cramon teaches the operations further comprise: storing each of the texture properties as a separate texture parameter. (“The present methods combine substantially shadow-free lighting with photography to capture surface textures that isolate diffuse color data from specular color data. A set of diffuse albedo surface textures are used with conventional photogrammetry techniques to generate a model of a real-world location or scene. Matched images or images of substantially the same subject matter exposed under different lighting conditions are used to generate a modified image. This modified image or specular roughness surface texture is used as a separate input when rendering a virtual environment from the model. Accordingly, a set of exposures captured at a location are temporarily stored as image files and processed using an image-processing technique to generate the modified image.” [0081] “The result of the difference blend operation, a modified image or specular roughness surface texture, can be stored and applied as an input to an image processor to generate specular roughness projection maps using the three-dimensional model and UV maps generated from the diffuse albedo surface textures, structured light techniques or alternate scanning methods.” [0093]) 14. Claim 10 is similar in scope to claim 1, and thus is rejected under similar rationale. 15. Claims 14 is similar in scope to claim 5, and thus is rejected under similar rationale. 16. Claim 15 is similar in scope to claim 6, and thus is rejected under similar rationale. 17. Claim 19 is similar in scope to claim 1, and thus is rejected under similar rationale. von Cramon additionally teaches A non-transitory computer-readable storage medium that stores instructions that when executed by at least one processor (“The image-processing system includes an image processor 500 that communicates via a data interface 510 with the image-capture system 400, an image store 590, an image editor 592 and a display apparatus 570. The data interface 510 communicates directly or indirectly with each of the image store 590, image editor 592 and display apparatus 570. … the data interface 510 may include one or more readers or ports for receiving data from a portable data storage medium such as a secure digital memory card or SD card (not shown).” [0185] “the image processor 500 may be used to render, manipulate and store a video product. Such a video product may be distributed to theaters, network access providers or other multi-media distributors via a wired or wireless network or a data storage medium.” [0193]) 18. Claim(s) 2-4, 7-9, 11-13, 16-18 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over von Cramon (US 2022/0092849 A1) and Mullins (US 2017/0193686 A1), as applied to claims 1, 10 and 19 above, and further in view of Pizer et al. (US 2020/0219272 A1). 19. With reference to claim 2, the combination of von Cramon and Mullins does not explicitly teach estimating the camera pose for each respective image of the set of images comprises: extracting features from a respective image; matching the features extracted from the respective image to features extracted from at least one other image from the set of images to obtain a set of matching features providing image correspondence; and estimating a camera pose of the respective image based on the set of matching features and geometric relationships of the set of matching features between the respective image and the at least one other image. This is what Pizer teaches (“SfM and SfS. An example SfMS framework or method may be based on two classical methods: SfM and SfS. SfM [26], [13], [12] is the simultaneous estimation of camera motion and 3D scene structure from multiple images taken at different viewpoints. Typical SfM methods produce a sparse scene representation by first detecting and matching local features in a series of input images, which are the individual frames of the endoscope video in the application described herein. Then, starting from an initial two-view reconstruction, these methods incrementally estimate both camera poses (rotation and position for each image) and scene structure. The scene structure is parameterized by a set of 3D points projecting to corresponding 2D image features. One point of interest to the generality of the SfMS framework is that sparse non-rigid reconstruction in medical settings is an unsolved problem [2]. However, the approach described herein can handle any sparse data as input, so the approach could easily be integrated with non-rigid SfM formulations that produce time-dependent sparse 3D geometry.” [0073] “in order to produce more robust correspondence matching results, the temporal coherence constraints can be leveraged by using a KLT tracker. … This section presents a method that solves this problem and augments the tracking-based correspondence matching. …In order to solve the short-term loop closure problem, Algorithm 2 may be improved by using a frame-skipping strategy. … Once the network is trained, video frames can be fed to it sequentially and the DenseSLAMNet will output the dense depth map and relative camera pose for each input frame.” [0111-0117] “Then for each frame, a depth map that corresponds to its camera position is extracted from S.sub.fused for reflectance model estimation. In such a way, all the single frame reconstructions may be using the same reference surface as used for reflectance model estimation, so more coherent results are generated.” [0129] “After a complete endoscopogram is generated using 3D reconstruction and group-wise geometry fusion algorithms described herein, the endoscopogram may be registered to CT for achieving the fusion between endoscopic video and CT. To allow a good initialization of the registration, first the tissue-gas surface from the CT may be identified or derived (e.g., extracted) and then a surface-to-surface registration between the endoscopogram and the surface derived from the CT may be performed.” [0144]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Pizer into the combination of von Cramon and Mullins, in order to provide precise 3D spatial information. 20. With reference to claim 3, the combination of von Cramon and Mullins does not explicitly teach constructing the 3D surface mesh comprises: applying a structure from motion (SfM) algorithm based on the set of images and the camera pose estimated for each respective image in the set of images. This is what Pizer teaches (“given an input endoscopic video sequence, a throat surface is reconstructed as a textured 3D mesh (see FIG. 1), also referred to herein as an endoscopogram. The endoscopogram is generated by first reconstructing a textured 3D partial surface for each frame. Then these multiple partial surfaces are fused into an endoscopogram using a group-wise surface registration algorithm and a seamless texture fusion from the partial surfaces.” [0032] “a 3D reconstruction of a throat surface may be performed using the preprocessed images. In some embodiments, the 3D reconstruction may utilize sparse, multi-view data obtained via Structure-from-Motion (SfM) to guide Shape-from-Shading (SfS) reconstruction of the throat surface in individual frames. … the 3D reconstruction may utilize a recurrent neural network (RNN) based depth estimation method that implicitly models the complex tissue reflectance property and performs depth estimation and camera pose estimation in real-time.” [0041-0042]) “SfM and SfS. An example SfMS framework or method may be based on two classical methods: SfM and SfS. SfM [26], [13], [12] is the simultaneous estimation of camera motion and 3D scene structure from multiple images taken at different viewpoints. Typical SfM methods produce a sparse scene representation by first detecting and matching local features in a series of input images, which are the individual frames of the endoscope video in the application described herein. Then, starting from an initial two-view reconstruction, these methods incrementally estimate both camera poses (rotation and position for each image) and scene structure. The scene structure is parameterized by a set of 3D points projecting to corresponding 2D image features.” [0073]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Pizer into the combination of von Cramon and Mullins, in order to provide precise 3D spatial information. 21. With reference to claim 4, the combination of von Cramon and Mullins does not explicitly teach the operations further comprise: receiving manual input instructions for the 3D surface mesh; and correcting the 3D surface mesh responsive to the manual input instructions. This is what Pizer teaches (“given an input endoscopic video sequence, a throat surface is reconstructed as a textured 3D mesh (see FIG. 1), also referred to herein as an endoscopogram. The endoscopogram is generated by first reconstructing a textured 3D partial surface for each frame. Then these multiple partial surfaces are fused into an endoscopogram using a group-wise surface registration algori “training data for the DispNet architecture may be generated such that by using some of its functions the specular points in endoscopic images can be manually removed. For example, 256 manually generated frames may be used as training data for training the DispNet architecture to perform specularity removal.thm and a seamless texture fusion from the partial surfaces.” [0032] “training data for the DispNet architecture may be generated such that by using some of its functions the specular points in endoscopic images can be manually removed. For example, 256 manually generated frames may be used as training data for training the DispNet architecture to perform specularity removal.” [0061] “SfM [26], [13], [12] is the simultaneous estimation of camera motion and 3D scene structure from multiple images taken at different viewpoints. Typical SfM methods produce a sparse scene representation by first detecting and matching local features in a series of input images, which are the individual frames of the endoscope video in the application described herein. Then, starting from an initial two-view reconstruction, these methods incrementally estimate both camera poses (rotation and position for each image) and scene structure. The scene structure is parameterized by a set of 3D points projecting to corresponding 2D image features.” [0073] “In the SfMS reconstruction method introduced in section IV-A, there are no temporal constraints between successive frame-by-frame reconstructions. This fact and the method's reliance on reflectance model initialization can lead to inconsistent reconstructions and may even result in failure to reconstruct some frames. As such, manual intervention may be needed for selecting partial surface reconstructions for fusion.” [0127]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Pizer into the combination of von Cramon and Mullins, in order to provide precise 3D spatial information. 22. With reference to claim 7, the combination of von Cramon and Mullins does not explicitly teach optimizing texture properties of the plurality of surfaces of the 3D surface mesh comprises: modeling a scene using the 3D surface mesh, the camera pose estimated for each respective image of the set of images, and the at least one light source; initializing the texture properties; and optimizing the texture properties using inverse rendering and parameters of the at least one light source until convergence. This is what Pizer teaches (“given an input endoscopic video sequence, a colon surface is reconstructed as a textured 3D mesh (see FIG. 3), also referred to herein as an endoscopogram. The endoscopogram is generated by first reconstructing a textured 3D partial surface for each frame. Then these multiple partial surfaces are fused into an endoscopogram using a group-wise surface registration algorithm and a seamless texture fusion from the partial surfaces.” [0033] “SfM [26], [13], [12] is the simultaneous estimation of camera motion and 3D scene structure from multiple images taken at different viewpoints. Typical SfM methods produce a sparse scene representation by first detecting and matching local features in a series of input images, which are the individual frames of the endoscope video in the application described herein. Then, starting from an initial two-view reconstruction, these methods incrementally estimate both camera poses (rotation and position for each image) and scene structure. “ [0073] “The reflectance model described herein is based on the set of BRDF basis functions introduced by Koenderink et al. [33]. These functions form a complete, orthonormal basis on the half-sphere derived via a mapping from the Zernike polynomials, which are defined on the unit disk. The BRDF basis of Koenderink et al. is adapted to produce a multi-lobe reflectance model for camera-centric SfS. First, taking the light source to be at the camera center, let θ.sub.i=θ.sub.r and Δϕ.sub.ir=0; where α.sub.k and β.sub.k are coefficients that specify the BRDF.” [0082-0084] “FIG. 9 shows a visual comparison of surfaces generated by an example textured 3D reconstruction approach described herein for an image from a ground truth dataset. In FIG. 9, the top row depicts visualizations of a surface without texture from an original image and the bottom row depicts visualizations of the surface with texture from the original image. Columns from left to right: (1) using a Lambertian BRDF, (2) using the proposed BRDF described herein (K=2) without image-weighted derivatives, (3) using the proposed BRDF described herein (K=2) with image-weighted derivatives, and (4) the ground-truth surface. Note the oversmoothing along occlusion boundaries in column (2) versus column (3). Reflectance Model Estimation. From this warped surface, reflectance model parameters Θ may be optimized for the specified BRDF (where the parameters depend on the chosen BRDF). In some embodiments, this optimization may be performed by minimizing the least-squares error, where I.sub.est(x,y; Θ) is the estimated image intensity (see Equation (3)) as determined by S.sub.warp.sup.n and the estimated BRDF.” [0104-0106]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Pizer into the combination of von Cramon and Mullins, in order to provide precise 3D spatial information. 23. With reference to claim 8, the combination of von Cramon and Mullins does not explicitly teach initializing the texture properties comprises: setting all pixels of a respective surface of the 3D surface mesh equal to predefined or pseudo-random values. This is what Pizer teaches (“given an input endoscopic video sequence, a colon surface is reconstructed as a textured 3D mesh (see FIG. 3), also referred to herein as an endoscopogram. The endoscopogram is generated by first reconstructing a textured 3D partial surface for each frame. Then these multiple partial surfaces are fused into an endoscopogram using a group-wise surface registration algorithm and a seamless texture fusion from the partial surfaces.” [0033] “SfM [26], [13], [12] is the simultaneous estimation of camera motion and 3D scene structure from multiple images taken at different viewpoints. Typical SfM methods produce a sparse scene representation by first detecting and matching local features in a series of input images, which are the individual frames of the endoscope video in the application described herein. Then, starting from an initial two-view reconstruction, these methods incrementally estimate both camera poses (rotation and position for each image) and scene structure. “ [0073] “The reflectance model described herein is based on the set of BRDF basis functions introduced by Koenderink et al. [33]. These functions form a complete, orthonormal basis on the half-sphere derived via a mapping from the Zernike polynomials, which are defined on the unit disk. The BRDF basis of Koenderink et al. is adapted to produce a multi-lobe reflectance model for camera-centric SfS. First, taking the light source to be at the camera center, let θ.sub.i=θ.sub.r and Δϕ.sub.ir=0; where α.sub.k and β.sub.k are coefficients that specify the BRDF.” [0082-0084] “FIG. 9 shows a visual comparison of surfaces generated by an example textured 3D reconstruction approach described herein for an image from a ground truth dataset. In FIG. 9, the top row depicts visualizations of a surface without texture from an original image and the bottom row depicts visualizations of the surface with texture from the original image. Columns from left to right: (1) using a Lambertian BRDF, (2) using the proposed BRDF described herein (K=2) without image-weighted derivatives, (3) using the proposed BRDF described herein (K=2) with image-weighted derivatives, and (4) the ground-truth surface. Note the oversmoothing along occlusion boundaries in column (2) versus column (3). Reflectance Model Estimation. From this warped surface, reflectance model parameters Θ may be optimized for the specified BRDF (where the parameters depend on the chosen BRDF). In some embodiments, this optimization may be performed by minimizing the least-squares error, where I.sub.est(x,y; Θ) is the estimated image intensity (see Equation (3)) as determined by S.sub.warp.sup.n and the estimated BRDF.” [0104-0106] “In the first stage, an initial texture is created: for each voxel on the endoscopogram surface, the image is selected whose reconstruction has the closest distance to that voxel to color it. A Markov Random Field (MRF) based regularization is used to make the pixel selection more spatially consistent, resulting in a texture map that has multiple patches with clear seams, as shown in FIG. 14.” [0133]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Pizer into the combination of von Cramon and Mullins, in order to provide precise 3D spatial information. 24. With reference to claim 9, the combination of von Cramon and Mullins does not explicitly teach optimizing the texture properties using inverse rendering and parameters of the at least one light source until convergence comprises: generating a portion of the scene from a known camera position; computing Mean Squared Error (MSE) between the portion of the scene and a corresponding portion of an image from the set of images; and determining that a difference between the portion of the scene and the corresponding portion of the image is less than a threshold difference. This is what Pizer teaches (“given an input endoscopic video sequence, a colon surface is reconstructed as a textured 3D mesh (see FIG. 3), also referred to herein as an endoscopogram. The endoscopogram is generated by first reconstructing a textured 3D partial surface for each frame. Then these multiple partial surfaces are fused into an endoscopogram using a group-wise surface registration algorithm and a seamless texture fusion from the partial surfaces.” [0033] “SfM [26], [13], [12] is the simultaneous estimation of camera motion and 3D scene structure from multiple images taken at different viewpoints. Typical SfM methods produce a sparse scene representation by first detecting and matching local features in a series of input images, which are the individual frames of the endoscope video in the application described herein. Then, starting from an initial two-view reconstruction, these methods incrementally estimate both camera poses (rotation and position for each image) and scene structure. “ [0073] “The reflectance model described herein is based on the set of BRDF basis functions introduced by Koenderink et al. [33]. These functions form a complete, orthonormal basis on the half-sphere derived via a mapping from the Zernike polynomials, which are defined on the unit disk. The BRDF basis of Koenderink et al. is adapted to produce a multi-lobe reflectance model for camera-centric SfS. First, taking the light source to be at the camera center, let θ.sub.i=θ.sub.r and Δϕ.sub.ir=0; where α.sub.k and β.sub.k are coefficients that specify the BRDF.” [0082-0084] “FIG. 9 shows a visual comparison of surfaces generated by an example textured 3D reconstruction approach described herein for an image from a ground truth dataset. In FIG. 9, the top row depicts visualizations of a surface without texture from an original image and the bottom row depicts visualizations of the surface with texture from the original image. Columns from left to right: (1) using a Lambertian BRDF, (2) using the proposed BRDF described herein (K=2) without image-weighted derivatives, (3) using the proposed BRDF described herein (K=2) with image-weighted derivatives, and (4) the ground-truth surface. Note the oversmoothing along occlusion boundaries in column (2) versus column (3). Reflectance Model Estimation. From this warped surface, reflectance model parameters Θ may be optimized for the specified BRDF (where the parameters depend on the chosen BRDF). In some embodiments, this optimization may be performed by minimizing the least-squares error, where I.sub.est(x,y; Θ) is the estimated image intensity (see Equation (3)) as determined by S.sub.warp.sup.n and the estimated BRDF.” [0104-0106] “In the first stage, an initial texture is created: for each voxel on the endoscopogram surface, the image is selected whose reconstruction has the closest distance to that voxel to color it. A Markov Random Field (MRF) based regularization is used to make the pixel selection more spatially consistent, resulting in a texture map that has multiple patches with clear seams, as shown in FIG. 14. Then in the second stage, to generate a seamless texture, within-patch intensity gradient magnitude differences and inter-patch-boundary color differences are minimized.” [0133]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Pizer into the combination of von Cramon and Mullins, in order to provide precise 3D spatial information. 25. Claims 11-13 are similar in scope to claims 2-4, and they are rejected under similar rationale. 26. Claims 16-18 are similar in scope to claims 7-9, and they are rejected under similar rationale. 27. Claim 20 is similar in scope to claim 7, and thus is rejected under similar rationale. Conclusion 28. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michelle Chin whose telephone number is (571)270-3697. The examiner can normally be reached on Monday-Friday 8:00 AM-4:30 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http:/Awww.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Kent Chang can be reached on (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is (571)273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https:/Awww.uspto.gov/patents/apply/patent- center for more information about Patent Center and https:/Awww.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHELLE CHIN/ Primary Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

Apr 28, 2025
Application Filed
Sep 15, 2026
Non-Final Rejection mailed — §103, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749140
HANDLING PIPELINE SUBMISSIONS ACROSS MANY COMPUTE UNITS
2y 1m to grant Granted Sep 29, 2026
Patent 12740832
SYSTEMS AND METHODS FOR PREOPERATIVE PLANNING AND POSTOPERATIVE ANALYSIS OF SURGICAL PROCEDURES
1y 11m to grant Granted Sep 22, 2026
Patent 12737983
CONSTRAINT-DRIVEN IMPRINT-BASED MESH GENERATION FOR COMPUTER-AIDED DESIGN (CAD) OBJECTS
2y 4m to grant Granted Sep 15, 2026
Patent 12725358
METHODS FOR MODELLING AND MANAUFACTURING A DEVICE
2y 4m to grant Granted Sep 01, 2026
Patent 12720004
AUTOMATICALLY GENERATING COLORS FOR OVERLAID CONTENT OF VIDEOS
3y 2m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
85%
Grant Probability
97%
With Interview (+11.6%)
2y 2m (~9m remaining)
Median Time to Grant
Low
PTA Risk
Based on 656 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month