Prosecution Insights
Last updated: August 18, 2026
Application No. 19/341,119

GENERATING THREE-DIMENSIONAL (3D) IMAGES FROM IMAGES USING MACHINE LEARNING MODELS

Non-Final OA §102§103
Filed
Sep 26, 2025
Priority
Sep 27, 2024 — provisional 63/700,024
Examiner
HE, YINGCHUN
Art Unit
2613
Tech Center
2600 — Communications
Assignee
Stability AI Ltd.
OA Round
3 (Non-Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
1y 5m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
539 granted / 657 resolved
+20.0% vs TC avg
Moderate +15% lift
Without
With
+14.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 4m
Avg Prosecution
19 currently pending
Career history
679
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
57.0%
+17.0% vs TC avg
§102
6.4%
-33.6% vs TC avg
§112
17.4%
-22.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 657 resolved cases

Office Action

§102 §103
DETAILED ACTION *Note in the following document: 1. Texts in italic bold format are limitations quoted either directly or conceptually from claims/descriptions disclosed in the instant application. 2. Texts in regular italic format are quoted directly from cited reference or Applicant’s arguments. 3. Texts with underlining are added by the Examiner for emphasis. 4. Texts with 5. Acronym “PHOSITA” stands for “Person Having Ordinary Skill In The Art”. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 4 May 2026 has been entered. Status of Claims This is in response to applicant’s amendment/response file on 4 May 2026, which has been entered and made of record. Claims 1, 3-4, 8-10, 12-13, 16 and 18 have been amended. Claims 21-22 have been added or cancelled. Claim 19-20 has been added or cancelled. Claims 1-18 and 21-22 are pending in the application. Response to Arguments Applicant’s arguments, see p.7-11, filed on 4 May 2026, with respect to the rejection(s) of Claim(s) 1/13/18 and their dependent claims under 35 USC §103 have been fully considered but are moot because the arguments do not apply to any of the references being used in the current rejection. The newly amended Claim(s) 1/13/18 is/are now rejected under 35 USC §103 as being unpatentable over Joachim (US 2024/0378832 A1) in view of Li et al. (Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF from a Single Image, 2020). See detailed rejections below. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-3, 5, 8, 13-16, 18 and 21-22 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Joachim (US 2024/0378832 A1) in view of Li et al. (Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF from a Single Image, 2020). Regarding Claim 1, Joachim discloses a system ([0002]: One or more embodiments described herein provide benefits and/or solve one or more problems in the art with systems, methods, and non-transitory computer-readable media that implement artificial intelligence models to facilitate flexible and efficient scene-based image editing) comprising: one or more storage media storing instructions (Fig.94: Memory 9404; Storage 9406); and one or more processors (Fig.94: Processor 9402) configured to execute the instructions to cause the system to: receive an image depicting a first object under a first illumination ([0578]: As illustrated in FIG. 45, the scene-based image editing system 106 determines a two-dimensional image 4500. For example, the two-dimensional image 4500 includes a plurality of objects—e.g., one or more objects in a foreground region and/or one or more objects in a background region. Any image is under illumination otherwise the image would be pure black); PNG media_image1.png 347 726 media_image1.png Greyscale generate a three-dimensional mesh object based at least in part on the first object, wherein the three-dimensional mesh object represents the first object (Fig.45: In one or more embodiments, the scene-based image editing system 106 utilizes the semantic map 4502 to generate a segmented three-dimensional mesh 4504. [0579]: In one or more embodiments, the scene-based image editing system 106 utilizes the semantic map 4502 to generate a segmented three-dimensional mesh 4504. Specifically, the scene-based image editing system 106 utilizes the object classifications of the pixels in the two-dimensional image 4500 to determine portions of a three-dimensional mesh that correspond to the objects in the two-dimensional image 4500. For example, the scene-based image editing system 106 utilizes a mapping between the two-dimensional image 4500 and the three-dimensional mesh representing the two-dimensional image 4500 to determine object classifications of portions of the three-dimensional mesh. To illustrate, the scene-based image editing system 106 determines specific vertices of the three-dimensional mesh that correspond to a specific object (e.g., a foreground object) detected in the two-dimensional image 4500 based on the mapping between the two-dimensional image 4500 and the two-dimensional image 4500); generate a representation of one or more material features of the first object, wherein the representation indicates at least one of a roughness feature or a metallic feature of the first object ([0350]: By utilizing both low-level feature maps and high-level feature maps, the scene-based image editing system 106 accurately predicts attributes across the wide range of semantic levels. For instance, the scene-based image editing system 106 utilizes low-level feature maps to accurately predict attributes such as, but not limited to, colors (e.g., red, blue, multicolored), patterns (e.g., striped, dotted, striped), geometry (e.g., shape, size, posture), texture (e.g., rough, smooth, jagged), or material (e.g., wooden, metallic, glossy, matte) of a portrayed object. Note Joachim teaches generating low-level feature 1710 and high-level feature map 1708 based on an input image through embedding neural network as shown in Fig.17. Also see [0347]: As shown in FIG. 17, the scene-based image editing system 106 utilizes an embedding neural network within the multi-attribute contrastive classification neural network. In particular, as illustrated in FIG. 17, the scene-based image editing system 106 utilizes a low-level embedding layer 1704 (e.g., embedding NNI) (e.g., of the embedding neural network 1604 of FIG. 16) to generate a low-level attribute feature map 1710 from a digital image 1702. Furthermore, as shown in FIG. 17, the scene-based image editing system 106 utilizes a high-level embedding layer 1706 (e.g., embedding NNh) (e.g., of the embedding neural network 1604 of FIG. 16) to generate a high-level attribute feature map 1708 from the digital image 1702. PNG media_image2.png 480 767 media_image2.png Greyscale ); and generate a texture for the three-dimensional mesh object based at least in part on the three-dimensional mesh object ([0605]: Additionally, the scene-based image editing system 106 generates a texture to apply to the portion of the first three-dimensional mesh 4700 utilizing an additional inpainting model. [0722]: In some embodiments, in connection with modifying the two-dimensional image 6200 based on the modified three-dimensional human model 6204, the scene-based image editing system 106 also extracts a texture map 6208 corresponding to the three-dimensional human model 6202. Specifically, the scene-based image editing system 106 extracts the texture map 6208 from pixel values of the two-dimensional image 6200 in connection with the three-dimensional human model 6202. For instance, the scene-based image editing system 106 utilizes a neural network to generate the texture map 6208 including a UV mapping from the image space to the three-dimensional human model 6202). Joachim discloses generate a texture for the three-dimensional mesh object based at least in part on the three-dimensional mesh object as explained above. But Joachim fails to explicitly disclose generate a texture for the three-dimensional mesh object based at least in part on the three-dimensional mesh object and the representation of one or more material features of the first object, wherein the texture reflects the first object absent the first illumination and is usable to apply a second illumination that is different from the first illumination. However inverse rendering which is a process of taking 2D image(s) to deduce the physical properties of a scene had been known to a PHOSITA before the effective filing date of the claimed invention. Li, in the same field of endeavor, addresses a long-standing challenge in inverse rendering to reconstruct geometry, spatially-varying complex reflectance and spatially-varying lighting from a single RGB image of an arbitrary indoor scene captured under uncontrolled conditions (p.2475 left column Section 1. Introduction lines 1-5). Li discloses Given a single image of an indoor scene (a), we recover its diffuse albedo (b), normals (c), specular roughness (d), depth (e) and spatially-varying lighting (f). We build a large-scale high-quality synthetic training dataset rendered with photorealistic SVBRDF (p.2475 Fig.1). PNG media_image3.png 574 574 media_image3.png Greyscale The specular roughness feature is related to material features of an object. Li teaches the texture is based on SVBRDF texture (p.2477 right column last four lines: we first search for an optimal crop from our SVBRDF texture by minimizing gradients for diffuse albedo, normals and roughness perpendicular to the patch boundaries). Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Li into that of Joachim and to include the limitation of generate a texture for the three-dimensional mesh object based at least in part on the three-dimensional mesh object and the representation of one or more material features of the first object, wherein the texture reflects the first object absent the first illumination and is usable to apply a second illumination that is different from the first illumination in order to provide an advanced tool for applications like relighting and material editing as suggested by Li (p.2479 right column lines after equation 3). Regarding Claim 2, Joachim further teaches or suggests wherein the first object is presented from a first camera view; and wherein the system further comprises processors configured to execute the instructions to further cause the system to present the three-dimensional mesh object from a second camera view that is different from the first camera view ([0618]: In one or more additional embodiments, the scene-based image editing system 106 also provides tools for editing three-dimensional characteristics of the object 4902. In particular, FIG. 49 illustrates a set of visual indicators 4906 for rotating the object 4902 in the three-dimensional space. More specifically, in response to detecting one or more interactions with the set of visual indicators 4906, the scene-based image editing system 106 modifies an orientation of the three-dimensional mesh corresponding to the object 4902 in the three-dimensional space and updates the two-dimensional depiction of the object 4902 accordingly). PNG media_image4.png 504 707 media_image4.png Greyscale Regarding Claim 3, Joachim modified by Li further teaches or suggests wherein the texture for the three-dimensional mesh object reflects the first object absent the first illumination, and wherein the texture is represented using a texture space applicable to the three-dimensional mesh object (Li p.2482 Section 6 lines 1-4: We have presented the first holistic inverse rendering framework that estimates disentangled shape, SVBRDF and spatially-varying lighting, from a single image of an indoor scene. Also see Fig.1 notes regarding recovering diffuse albedo from a single image. Also note an albedo map is a texture map that defines the base color and pure reflectance of a surface). The same reason to combine as that of Claim 1 is applied. Regarding Claim 5, Joachim discloses in one or more embodiments, the scene-based image editing system utilizes one or more machine learning models to process a digital image in anticipation of user interactions for modifying the digital image ([0101]). Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Joachim and to include the limitation of wherein the three-dimensional mesh object and the texture are used for generating an asset for digital entertainment since users interactions include digital entertainment such as video game or digital entertainment applications involving users feedbacks. Regarding Claim 8, Joachim discloses the scene-based image editing system 106 utilizes lighting data, shading data, and three-dimensional characteristics of content in the two-dimensional image to generate the modified two-dimensional image 4400 ([0570]) as shown in Fig.44. PNG media_image5.png 426 715 media_image5.png Greyscale Joachim further discloses the editing system could generating a modified two-dimensional image based on lighting data ([0570]: Additionally, FIG. 44 illustrates that the scene-based image editing system 106 utilizes lighting data, shading data, and three-dimensional characteristics of content in the two-dimensional image to generate the modified two-dimensional image 4400). Li discloses using Inverse Rendering result from a single image for relighting and material editing (p.2479 right column lines 12-23: PNG media_image6.png 49 472 media_image6.png Greyscale ). A skilled person would have known that relighting editing is to apply a different light setting. Therefore Joachim modified by Li further teaches or suggests wherein the execution of the instructions further causes the system to: apply the second illumination to the three-dimensional mesh object; and output a representation of the three-dimensional mesh object and the second illumination as a three-dimensional object file. The same reason to combine as that of Claim 1 is applied. Regarding Claim 13, Claim 13 is similar to Claim 1 except in the format of method. Therefore the same reason(s) for rejection is applied to Claim 1 is also applied to Claim 13. Regarding Claim 14, Joachim further teaches or suggests wherein the first object is presented from a first camera view; and wherein the method further comprises: presenting the three-dimensional mesh object from a second camera view that is different from the first camera view absent the texture (See Fig.49 above. Also see [0618]: In one or more additional embodiments, the scene-based image editing system 106 also provides tools for editing three-dimensional characteristics of the object 4902. In particular, FIG. 49 illustrates a set of visual indicators 4906 for rotating the object 4902 in the three-dimensional space. More specifically, in response to detecting one or more interactions with the set of visual indicators 4906, the scene-based image editing system 106 modifies an orientation of the three-dimensional mesh corresponding to the object 4902 in the three-dimensional space and updates the two-dimensional depiction of the object 4902 accordingly) absent the texture ( Joachim teaches applying textures to a 3D mesh ([0605]: Additionally, the scene-based image editing system 106 generates a texture to apply to the portion of the first three-dimensional mesh 4700 utilizing an additional inpainting model). According broadest interpretation, a mesh is a geometry portion of a 3D model which is the underlying 3D structural shape (polygons, vertices, edges) of an object, while a texture is a color portion of a 3D model which is an image pattern stretched over the mesh. A skilled person would have known before applying the texture, the 3D mesh is a mesh absent the texture). Regarding Claim 15, Joachim modified by Li further teaches or suggests wherein generating the three-dimensional mesh object is further based at least in part on the representation of one or more material features of the first object (Li p.2475 right column lines 1-5: Driven by the success of deep learning methods on similar scene inference tasks (geometric reconstruction [16], lighting estimation [17], material recognition [9]), we propose training a deep convolutional neural network to regress these scene parameters from an input image). The same reason to combine as that of Claim 1 is applied. Regarding Claim 16, Claim 16 is/are similar to Claim 8 except in the format of method. Therefore the same reason(s) for rejection is/are applied to Claim 8 is/are also applied to Claim 16. Regarding Claim 18, Claim 18 is similar to Claim 1 except in the format of non-transitory computer-readable storage media. Therefore the same reason(s) for rejection is/are applied to Claim 1 is also applied to Claim 18. Regarding Claim 21, Li teaches or suggests decomposing, from the image, an intrinsic color representation of the first object (Fig.1: diffuse albedo. Albedo color is true or intrinsic color of an object) and a shading component representation (Fig.1(c): specular roughness) of the first illumination, wherein the texture is generated further based at least in part on the intrinsic color representation (Li p.2477 right column line 3 from bottom: SVBRDF texture). The same reason to combine as that of Claim 1 is applied. Regarding Claim 22, Li further teaches or suggests determining, based at least in part on the shading component (Fig.1(d): specular roughness), an illumination distribution (Fig.1 (f): spatially-varying lighting) that is applied to enforce consistency between a luminance of the image and a luminance of a reconstruction of the first object having a uniform intrinsic color (Fig.1(b): diffuse albedo). The same reason to combine as that of Claim 1 is applied. PNG media_image3.png 574 574 media_image3.png Greyscale Claims 4, 7, 9 , 12 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Joachim (US 2024/0378832 A1) in view of Li et al. (Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF from a Single Image, 2020) as applied to Claims 1 and 13 above, and further in view of Hong et al. (“LRM: LARGE RECONSTRUCTION MODEL FOR SINGLE IMAGE TO 3D” arXiv, 9 March 2024) and Kang et al. (KR 2022/0049689 A). Regarding Claim 4, Joachim modified by Li fails to explicitly disclose wherein the execution of the instructions further causes the system to generate a triplane embedding. However Hong discloses wherein the execution of the instructions further causes the system to generate a triplane embedding (p.8 second last paragraph, see below). PNG media_image7.png 75 503 media_image7.png Greyscale Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Hong into that of Joachim as modified and to include the limitation of wherein the execution of the instructions further causes the system to generate a triplane embedding in order to use the available technology to create a 3D Shape from a single image of an arbitrary object as suggested by Hong (p.2 Section 1 Introduction line 1). Joachim modified by Li and Hong fails to disclose wherein the three-dimensional mesh object is generated based at least in part on inputting the triplane embedding to an offset feature extractor that determines at least one offset feature associated with the first object and the three-dimensional mesh object. However Kang teaches in order to configure multiple expression sets (330), the reference model setting unit (200) can extract an offset feature vector between the reference expression set of the reference model and the expression set for each facial part of the reference model ([0045]). Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Kang into that of Joachim as modified and to include the limitation of wherein the three-dimensional mesh object is generated based at least in part on inputting the triplane embedding to an offset feature extractor that determines at least one offset feature associated with the first object and the three-dimensional mesh object in order to properly follow the facial expression movements of the user as suggested by Kang ([0007]) when the object is a face of user. Regarding Claim 7, Joachim modified by Li fails to explicitly recite wherein the three-dimensional mesh object absent the texture and the texture are generated in less than 5 seconds. However Hong discloses We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image (Abstract). Hong further discloses PNG media_image8.png 249 893 media_image8.png Greyscale PNG media_image9.png 119 777 media_image9.png Greyscale Hong teaches the whole Large Reconstruction Model generation takes less than 5 seconds. Hong discloses the LRM can use DINO to reconstruct the geometry and color in 3D space. A skilled person would have known that a mesh is a geometry portion of a 3D model which is the underlying 3D structural shape (polygons, vertices, edges) of an object, while a texture is a color portion of a 3D model which is an image pattern stretched over the mesh. Note Joachim teaches applying texture to mesh ([0605]: Additionally, the scene-based image editing system 106 generates a texture to apply to the portion of the first three-dimensional mesh 4700 utilizing an additional inpainting model). Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Hong into that of Joachim as modified and to include the limitation of wherein the three-dimensional mesh object absent the texture and the texture are generated in less than 5 seconds in order to use the available technology to create a 3D Shape from a single image of an arbitrary object as suggested by Hong (p.2 Section 1 Introduction line 1). Regarding Claim 9, Joachim modified by Li fails to explicitly recite generate, by encoding the image, a triplane embedding that includes a resolution of at least 300 pixels x 300 pixels; and generate the three-dimensional mesh object and the texture based at least in part on the triplane embedding. However Hong discloses a Large Reconstruction Model (LRM) that predicts a 3D model of an object from a single input image (Abstract) and the LRM generates both texture and mesh (p.9 lines 1-3 and p.8 second last paragraph). According broadest interpretation, a mesh is a geometry portion of a 3D model which is the underlying 3D structural shape (polygons, vertices, edges) of an object, while a texture is a color portion of a 3D model which is an image pattern stretched over the mesh. Hong discloses We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image (Abstract) Our method takes an image as input and regresses a NeRF in the form of a triplane representation (Chan et al., 2022). Specifically, LRM utilizes the pre-trained visual transformer DINO (Caron et al., 2021) as the image encoder to generate the image features, and learns an image-to-triplane transformer decoder to project the 2D image features onto the 3D triplane via cross-attention and model the relations among the spatially-structured triplane tokens via self-attention (p.2 first paragraph). Hong further discloses the query resolution would be 384x384x384 (see p.8 last paragraph PNG media_image7.png 75 503 media_image7.png Greyscale ). Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Hong into that of Joachim as modified and to include the limitation of wherein the execution of the instructions further cause the system to: generate, by encoding the image, a triplane embedding that includes a resolution of at least 300 pixels x 300 pixels; and generate the three-dimensional mesh object and the texture based at least in part on the triplane embedding in order to use the available technology to create a 3D Shape from a single image of an arbitrary object as suggested by Hong (p.2 Section 1 Introduction line 1). Regarding Claim 12, Joachim modified by Li teaches or suggests wherein the execution of the instructions further causes the system to generate the three-dimensional mesh object based at least in part on using at least an albedo feature extractor, a lighting feature extractor, (Li Fig.1: Given a single image of an indoor scene (a), we recover its diffuse albedo (b), normals (c), specular roughness (d),depth (e) and spatially-varying lighting (f)). But Joachim modified by Li fails to explicitly disclose wherein the execution of the instructions further cause the system to generate the three-dimensional mesh object based at least in part on using a density feature extractor. However Hong teaches the generation of the 3D mesh object is further based at least in part on using a density feature extractor (p.3 Section 3 Method: Lastly, the 3D point features are passed to a multi-layer perception to predict RGB and density for volumetric rendering). Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Hong into that of Joachim as modified and to include the limitation of wherein the execution of the instructions further cause the system to generate the three-dimensional mesh object based at least in part on using a density feature extractor in order to use the available technology to create a 3D Shape from a single image of an arbitrary object as suggested by Hong (p.2 Section 1 Introduction line 1). Regarding Claim 17, Joachim modified by Li fails to explicitly recite generating, based at least in part on inputting the image and a camera view embedding into a transformer-based neural network, a triplane embedding that includes a resolution of at least 300 pixels x 300 pixels; and generating the three-dimensional mesh object and the texture based at least in part on the triplane embedding. However Hong discloses generating, based at least in part on inputting the image and a camera view embedding (p.4 Section 3.2: We implement a transformer decoder to project image and camera features onto learnable spatial-positional embeddings and translate them to triplane representations. This decoder can be considered as a prior network that is trained with large-scale data to provide necessary geometric and appearance information to compensate for the ambiguities of single-image reconstruction) into a transformer-based neural network (p.4 Section 3.2: We implement a transformer decoder to project image and camera features onto learnable spatial-positional embeddings and translate them to triplane representations. This decoder can be considered as a prior network that is trained with large-scale data to provide necessary geometric and appearance information to compensate for the ambiguities of single-image reconstruction), a triplane embedding that includes a resolution of at least 300 pixels x 300 pixels; and generating the three-dimensional mesh object and the texture based at least in part on the triplane embedding (Abstract: We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. P.2 first paragraph: Our method takes an image as input and regresses a NeRF in the form of a triplane representation (Chan et al., 2022). Specifically, LRM utilizes the pre-trained visual transformer DINO (Caron et al., 2021) as the image encoder to generate the image features, and learns an image-to-triplane transformer decoder to project the 2D image features onto the 3D triplane via cross-attention and model the relations among the spatially-structured triplane tokens via self-attention. Hong further discloses the query resolution would be 384x384x384 (see p.8 last paragraph PNG media_image7.png 75 503 media_image7.png Greyscale ). Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Hong into that of Joachim as modified and to include the limitation of wherein the execution of the instructions further cause the system to: generate, by encoding the image, a triplane embedding that includes a resolution of at least 300 pixels x 300 pixels; and generate the three-dimensional mesh object and the texture based at least in part on the triplane embedding in order to use the available technology to create a 3D Shape from a single image of an arbitrary object as suggested by Hong (p.2 Section 1 Introduction line 1). Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Joachim (US 2024/0378832 A1) in view of Li et al. (Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF from a Single Image, 2020) as applied to Claim 1 above, and further in view of Xu et al. (“InstantMesh Efficient 3D mesh generation from a single image with sparse-view large reconstruction models” arXiv, 14 Apr. 2024). Regarding Claim 6, Joachim modified by Li fails to explicitly recite wherein the three-dimensional mesh object is generated in less than 10 seconds. However Xu discloses InstantMesh framework, Given an input image, we first utilize a multi-view diffusion model to synthesize 6 novel views at fixed camera poses. Then we feed the generated multi-view images into a transformer-based sparse-view large reconstruction model to reconstruct a high-quality 3D mesh. The whole image-to-3D generation process takes only around 10 seconds (p.3 Fig.2). PNG media_image10.png 227 635 media_image10.png Greyscale Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Xu into that of Joachim as modified and to include the limitation of wherein the three-dimensional mesh object is generated in less than 10 seconds in order to use the available tool to improve 3D generation as suggested by Xu (p.2 Section 1 Introduction: a Crafting 3D assets from single-view images can facilitate a broad range of applications, eg, virtual reality, industrial design, gaming and animation. We have witnessed a revolution on image and video generation with the emergence of large-scale diffusion models [37, 38] trained on billionscale data, which is able to generate vivid and imaginative contents from open-domain prompts). Claims 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Joachim (US 2024/0378832 A1) in view of Li et al. (Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF from a Single Image, 2020) and Hong et al. (“LRM: LARGE RECONSTRUCTION MODEL FOR SINGLE IMAGE TO 3D” arXiv, 9 March 2024) as applied to Claim 9 above, and further in view of Barsky (US 11,935,209 B1). Regarding Claim 10, Joachim modified by Hong etc. further teaches or suggests wherein the execution of the instructions further causes the system to: generate, by inputting the triplane embedding to a feature extractor, an illumination amplitude (Joachim [0574]: In one or more embodiments, the scene-based image editing system 106 applies the merged shadow maps when shading individual objects in a scene according to the object type/category of each object. In particular, as mentioned, the scene-based image editing system 106 shades inserted objects using a physically based rendering shader by sampling the object shadow map to calculate a light intensity. In one or more embodiments, the scene-based image editing system 106 does not render proxy objects in the final color output (e.g., the proxy objects are hidden from view in the modified two-dimensional image 4400). Furthermore, the scene-based image editing system 106 generates background and foreground object colors by: COLOR(x)=(SHADOW_FACTOR(x))*TEXTURE(x)+ (1−SHADOW_FACTOR(x))*SHADOW_COLOR. SHADOW_FACTOR is a value that the scene-based image editing system 106 generates by sampling the appropriate shadow map, with a larger sampling radius producing softer shadows. Additionally, SHADOW_COLOR represents the ambient light, which the scene-based image editing system 106 determines via a shadow estimation model or based on user input. TEXTURE represents a texture applied to the three-dimensional meshes 4404 according to corresponding pixel values in the two-dimensional image. Hong p.2 first paragraph: Our method takes an image as input and regresses a NeRF in the form of a triplane representation (Chan et al., 2022). Specifically, LRM utilizes the pre-trained visual transformer DINO (Caron et al., 2021) as the image encoder to generate the image features, and learns an image-to-triplane transformer decoder to project the 2D image features onto the 3D triplane via cross-attention and model the relations among the spatially-structured triplane tokens via self-attention). But Joachim as modified fails to disclose generate the texture based at least in part on the illumination amplitude. However Barsky teaches or suggests generate the texture based at least in part on the illumination amplitude (col.6 lines 8-14: The generated (at 210) neural network model accurately represents and/or recreates the textures, contours, patterns, shapes, and/or other commonality across the different surfaces of the object or scene. Specifically, the radiance field or representative function predicts the light intensity or radiance at any point in the 2D images in order to generate points in a 3D space that reconstruct the textures, contours, patterns, shapes, and/or other commonality for the different surfaces captured in the 2D images). Therefore it would have been obvious to a PHOSITA before the effective filing date to incorporate the teaching of Barsky into that of Joachim as modified and to include the limitation of generate the texture based at least in part on the illumination amplitude in order to apply machine learning model to generate texture and to perform the dynamic backfiling as suggested by Barsky (col.4 lines 1-2). Regarding Claim 11, Joachim as modified further teaches or suggests wherein the illumination amplitude is determined using at least one spherical gaussian illumination map and the triplane embedding (Joachim [0574]: In one or more embodiments, the scene-based image editing system 106 applies the merged shadow maps when shading individual objects in a scene according to the object type/category of each object. In particular, as mentioned, the scene-based image editing system 106 shades inserted objects using a physically based rendering shader by sampling the object shadow map to calculate a light intensity. Li p.2479 left column Section “Spatially Varying Lighting Prediction: We model the lighting as a spherical function L(η) approximated by the sum of spherical Gaussian lobes. Hong p.8 second last paragraph: During inference, LRM takes an arbitrary image as input (squared and background removed) and assumes the unknown camera parameters to be the normalized cameras that we applied to train the Objaverse data. We query a resolution of 384x384x384 points from the reconstructed triplane-NeRF and extract the mesh using Marching Cubes…. Barsky col.4 lines 10-13: NeRFs and/or other neural networks may be used to generate entire point clouds or 3D mesh models from multiple images that capture the same object or scene from different positions, angles, or viewpoints). The same reason to combine as that of Claim 10 is applied. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Bi et al. (US 2025/0259389 A1) teaches An Albedo map may refer to a base color input that defines a diffuse color or a reflectively or a surface. An Albedo map may be associated with a pure color of an object ([0069]). Lopes et al. (Material Palette: Extraction of Materials from a Single Image, 16-22 June 2024 ) discloses extraction of materials from a single image. The extracted Spatially Varying BRDF (SVBRDFs) encode material intrinsics (Albedo\Normal\Roughness). These can be reused for realistic material editing of 3D scenes. Chen et al. (“3DTOPIA-XL: SCALING HIGH-QUALITY 3D ASSET GENERATION VIA PRIMITIVE DIFFUSION, 16 Sept 2024) discloses single-view conditional generative model (p.10 Section 4.3 lines 1-2) and sampling the corresponding albedo colors and material values from UV space. PNG media_image11.png 111 937 media_image11.png Greyscale PNG media_image12.png 607 922 media_image12.png Greyscale Any inquiry concerning this communication or earlier communications from the examiner should be directed to YINGCHUN HE whose telephone number is (571)270-7218. The examiner can normally be reached M-F 8:00-5:00 MT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao M Wu can be reached at 571-272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YINGCHUN HE/Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Show 1 earlier event
Nov 28, 2025
Non-Final Rejection mailed — §102, §103
Dec 23, 2025
Examiner Interview Summary
Dec 23, 2025
Applicant Interview (Telephonic)
Jan 15, 2026
Response Filed
Feb 02, 2026
Final Rejection mailed — §102, §103
May 04, 2026
Request for Continued Examination
May 06, 2026
Response after Non-Final Action
Jun 17, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12700163
METHOD FOR GENERATING TEXTURE MAP OF PHYSICAL OBJECT, ELECTRONIC DEVICE, AND COMPUTER READABLE STORAGE MEDIUM
2y 3m to grant Granted Aug 04, 2026
Patent 12700335
HEAD-UP DISPLAY WITH REDUCED MASKING
2y 5m to grant Granted Aug 04, 2026
Patent 12694603
METHOD AND APPARATUS FOR SIGNALING A USER'S SAFE ZONE FOR 5G AUGMENTED/MIXED REALITY APPLICATIONS
2y 5m to grant Granted Jul 28, 2026
Patent 12673251
METHOD OF CONTROLLING OR AUGMENTING A VIRTUAL ENVIRONMENT OR LIVE VIDEO BROADCAST VIA SENSED MOTION
4y 1m to grant Granted Jul 07, 2026
Patent 12664959
METHOD TO SAVE POWER ON PIXEL LIT DISPLAYS
2y 6m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
97%
With Interview (+14.7%)
2y 4m (~1y 5m remaining)
Median Time to Grant
High
PTA Risk
Based on 657 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month