Prosecution Insights
Last updated: August 17, 2026
Application No. 18/916,747

SYSTEM AND METHOD FOR CONTROLLABLE TEXT-TO-3D ROOM MESH GENERATION WITH LAYOUT CONSTRAINTS

Non-Final OA §103§112
Filed
Oct 16, 2024
Priority
Oct 24, 2023 — provisional 63/592,910
Examiner
SONNERS, SCOTT E
Art Unit
2613
Tech Center
2600 — Communications
Assignee
The Hong Kong University of Science and Technology
OA Round
1 (Non-Final)
69%
Grant Probability
Favorable
1-2
OA Rounds
1y 5m
Est. Remaining
81%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
268 granted / 386 resolved
+7.4% vs TC avg
Moderate +12% lift
Without
With
+11.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
11 currently pending
Career history
406
Total Applications
across all art units

Statute-Specific Performance

§101
9.3%
-30.7% vs TC avg
§103
38.3%
-1.7% vs TC avg
§102
26.9%
-13.1% vs TC avg
§112
16.4%
-23.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 386 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Applicant’s claim for the benefit of a prior-filed application under 35 U.S.C. 119(e) or under 35 U.S.C. 120, 121, 365(c), or 386(c) is acknowledged. Applicant has not complied with one or more conditions for receiving the benefit of an earlier filing date under 35 U.S.C. 119(e) as follows: The later-filed application must be an application for a patent for an invention which is also disclosed in the prior application (the parent or original nonprovisional application or provisional application). The disclosure of the invention in the parent application and in the later-filed application must be sufficient to comply with the requirements of 35 U.S.C. 112(a) or the first paragraph of pre-AIA 35 U.S.C. 112, except for the best mode requirement. See Transco Products, Inc. v. Performance Contracting, Inc., 38 F.3d 551, 32 USPQ2d 1077 (Fed. Cir. 1994). The disclosure of the prior-filed application, Application No. 63,592910 fails to provide adequate support or enablement in the manner provided by 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph for one or more claims of this application. Regarding claim 1 and analogous independent claims, the claims recite, inter alia, a “neural radiance field (NeRF) module” and “a panoptic-enhanced radiance field (PeRF) module”. A review of the documentation providing support from Application No. 63,592910 finds that there is no support or enablement in the manner provided by 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph for the claims in the application. While there is support for the user interface module, text processing module, scene code generator, layout generation module, and appearance generation module communicating with the layout generation module, there is no support for a NeRF and PeRF module as recited communicating with the appearance generation module. Rather the single panoramic image of the room is used to generate a 3D model which is textured, but there is no application of any NeRF or PeRF module. Thus the claims are not afforded the benefit of a prior-filed application under 35 U.S.C. 119(e) as Application No. 63,592910 fails to provide adequate support or enablement in the manner provided by 35 U.S.C. 112(a) or pre-AIA 35 U.S.C. 112, first paragraph for one or more claims of this application. The Examiner notes that Application No. 63,592910 would provide support for a different but broader version of the claim with layout and generation modules that are coextensive with the disclosure of the provisional documentation, and do not include a NeRF and PeRF module. However the Examiner notes that the Fang prior art below still would likely qualify as prior art and could be used in a rejection, unless an exception to 35 U.S.C. 102(a)(1) were to apply, for example. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “user interface configured to receive user input…and convert it…” and “text processing module…configured to take the user input and process it” and “scene code generator… configured to translate…” and “layout generation module… configured to generate…” and “appearance generation module… configured to transform… and to convert…and to generate…” and “neural radiance field (NeRF) module… configured to construct” and “panoptic-enhanced radiance field (PeRF) module… configured to refine…to generate…” in claim 1 and “layout modification module… configured to allow the user to modify…provides an interface…to adjust…” in claim 6 and “panoramic update module… configured to update the panorama or the fully refined 3D room model dynamically…” in claim 10. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. Here the various modules are interpreted as computer implemented means-plus-function limitations where paragraph 0044 explains that “functional units and modules of the system and methods in accordance with the embodiments disclosed herein may be embodied in hardware or software. That is, the claimed system may be implemented entirely as machine instructions or as a combination of machine instructions and hardware elements” and “Hardware elements include, but are not limited to, computing devices, computer processors, or electronic circuitries including but not limited to application specific integrated circuits (ASIC), field programmable gate arrays (FPGA), microcontrollers, and other programmable logic devices configured or programmed according to the teachings of the present disclosure. Computer instructions or software codes running in the computing devices, computer processors, or programmable logic devices can readily be prepared by practitioners skilled in the software or electronic art based on the teachings of the present disclosure”. Thus here the limitations are interpreted to correspond to the disclosed hardware types which are specially programmed to perform the functions of the various modules disclosed. Thus the various modules and their specific manner of functioning to achieve the claimed functions is limited to those disclosed in the Specification and their equivalents where figure 1 illustrates the various modules which are further explained in paragraphs 0024-0036, where figure 2 provides further details as to how the modules perform their recited functions as explained in paragraphs 0037-0040. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Regarding claim 1, the instant claim recites a “user interface configured to receive user input regarding a room from a user in a form of text input and convert it into a language or code that is recognized by the system for processing” and “a text processing module communicating with the user interface and configured to take the user input and process it into a scene description.” Each limitation is definite when considered in isolation, however, when considered as an ordered combination and in combination with the other claim limitations, the limitations render the claim indefinite. This is because the user interface appears to receive the user input and “convert it into a language or code” but then the “text processing module” is recited to function to “take the user input” which was “in a form of text input” and “process it”, thus appearing to possibly bypass or ignore the “language or code” that the user interface converted the user input into already. Furthermore the “language or code” is not referred to again in the later limitations, but rather they appear to utilize the “user input” in some manner at least. Thus it is not clear whether when later modules utilize the “user input” they are converting such user input independently as part of their functioning or whether they are using “language or code” that the user input in the form of text has been converted into. The same reasoning applies to claim 11, which is rejected for the same reasons and will be interpreted in the same manner as claim 1 as explained below. Furthermore the dependent claims are rejected for carrying through this deficiency of their parent claims without rendering the claim definite. In the interest of compact prosecution, the Examiner will interpret the limitation as if the text processing module takes the user input in the form of the language or code from the UI and processes it into a scene description. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Fang et al1 (Fang) in view of Wang et al2 (Wang). Regarding claim 1, Fang teaches a system for computer-based 3D indoor scene assessment generation (note that 3D indoor scene assessment generation” is only given patentable weight insofar as any system which functions to generate such a 3D room model as below may be considered a computer-based 3D indoor scene assessment generation system as well as such 3D room model with objects laid out may be assessed as an indoor scene; also note that as explained above the “system” of the preamble while “computer-based” does not impart any structure nor is it interpreted as any specific structure itself, though in light of the claim interpretation explained above, such system must comprise the structural components functioning as interpreted below; see below for teaching of the system and of such generation), comprising: a user interface configured to receive user input regarding a room from a user in a form of text input and convert it into a language or code that is recognized by the system for processing (see Fang, section 3 “generate the room layout from the input text and then generate the room appearance according to the layout, followed by panoramic reconstruction to generate the final 3D textured mesh” and as in figure 2 “we synthesize a scene code from the text input and convert it” where as can be seen in figure 2 and as would be understood by one of ordinary skill in the art, as the user inputs text this means that the system includes a user interface that takes such text input and converts it into a form readable by the computer at the basic level of converting such input keystrokes corresponding to the text into ASCII or Unicode); a text processing module communicating with the user interface and configured to take the user input and process it into a scene description (see Fang, section 3 teaching to “generate the room layout from the input text and then generate the room appearance according to the layout, followed by panoramic reconstruction to generate the final 3D textured mesh” and as in figure 2 “we synthesize a scene code from the text input and convert it” where pages 1-2 of the “Appendix” section explain the conversion of the text input into a scene description in the form of a “text prompt for 3D room generation” where it is disclosed “We follow the SceneFormer ? to generate text prompts describing partial scene configurations. Each text prompt contains two to four sentences. The first sentence describes how many walls are in the room, then the second sentence describes two or three existing furniture in the room” and “we can get some relation-describing sentences to depict the partial scene. Finally, we randomly sampled zero to two relation-describing sentences to form the text prompt for 3D room generation”); a scene code generator communicating with the text processing module and configured to translate the scene description from the text processing module into a set of scene codes using a scene code diffusion model (see Fang, figure 2 as explained above teaching “we synthesize a scene code from the text input and convert it” as further explained in sections 3-3.1 teaching “use a holistic scene code to parametrize the indoor scene and design a diffusion model to learn its distribution” where “the holistic scene code is generated from text” and “given a 3D scene S with m walls and n furniture items, we represent the scene layout as a holistic scene code” and “encode each object oj as a node with various attributes” and “With the scene code definition, we build a diffusion model to learn its distribution” and “denoising network ϵθ takes the scene code xt, text prompt y, and timestep t as input, and denoises them iteratively to get a clean scene code ˆx0. Then we represent ˆx0 as a set of orientated bounding boxes of various semantic types to facilitate interactive editing” such as seen in figure 2 for example the text prompt is fed to a scene code diffusion module which outputs a scene code leading to a 3D layout of bounding boxes based on the scene codes); a layout generation module communicating with the scene code generator and configured to generate a 3D layout of the room using oriented bounding boxes based on the scene codes, wherein the 3D layout of the room preserves spatial integrity and relationships between objects as specified by the scene codes (see Fang, figure 2 teaching “the Layout Generation Stage, we synthesize a scene code from the text input and convert it to a 3D bounding box representation” and as in sections 3-3.1, “generate the room layout from the input text and then generate the room appearance according to the layout,” and “the Layout Generation Stage, we use a holistic scene code to parametrize the indoor scene and design a diffusion model to learn its distribution. Once the holistic scene code is generated from text, we recover the room as a set of orientated bounding boxes of walls and objects” and “denoising network ϵθ takes the scene code xt, text prompt y, and timestep t as input, and denoises them iteratively to get a clean scene code ˆx0. Then we represent ˆx0 as a set of orientated bounding boxes of various semantic types to facilitate interactive editing” such as seen in figure 2 for example the text prompt is fed to a scene code diffusion module which outputs a scene code leading to a generation of 3D layout of bounding boxes based on the scene codes); an appearance generation module communicating with the layout generation module and configured to transform the 3D layout of the room from the layout generation module into a visual representation of the room, wherein the appearance generation module is further configured to use equirectangular projection to convert the 3D layout of the room into a semantic layout and to generate a single panoramic image of the room based on the semantic layout (see Fang, sections 3-3.1 and figure 2 where as in figure 2 it can be seen that the 3D layout of the room in form of the 3D bounding box representation is subject to an “Equirectangular projection” to obtain a “semantic layout” that is then used to generate a “panorama” which is an image of the room based on the semantic layout where “In the Appearance Generation Stage, we project the bounding boxes into a semantic segmentation map to guide the panorama synthesis” and this is done by the “appearance generation module” where “the Appearance Generation Stage, we obtain an RGB panorama through a pre-trained latent diffusion model to represent the room texture. Specifically, we project the generated layout bounding boxes into a semantic segmentation map representing the layout” and “generate an RGB panorama from the input layout panorama” and as in section 3.2 “Given the layout of an indoor scene, we seek to obtain a proper panorama image to represent its appearance” and “To condition ControlNet on the scene layout, we convert the bounding box representation into a 2D semantic layout panorama through equirectangular projection. In this way, we get a pair of RGB and semantic layout panoramic images for each scene”); a neural radiance field (NeRF) module communicating with the appearance generation module and configured to construct a base 3D room model based on the panoramic image, producing a representation of the room by capturing spatial depth (see section 3 and figure 2 teaching that using the panoramic image, a base 3D room model that captures spatial depth through depth encoding the panoramic image is constructed where “the textured 3D mesh is obtained by estimating the depth map of the generated panorama” and as in figure 2 “panorama is then reconstructed into a textured 3D mesh model” as explained in section 3.2 teaching “After getting the scene panorama, we recover the depth map…to reconstruct a textured mesh through Possion reconstruction…and MVS-texture”); and a panoptic-enhanced radiance field (PeRF) module communicating with the NeRF module and configured to refine the base 3D room model by enhancing visual coherence, so as to generate a fully refined 3D room model (see Fang, section 3 and figure 2 as explained above where the refined base 3D room and fully refined 3D room model correspond to the final output textured mesh where MVS-texture has been applied where as explained above, as in figure 2 “panorama is then reconstructed into a textured 3D mesh model” and as explained in section 3.2 teaching “After getting the scene panorama, we recover the depth map…to reconstruct a textured mesh through Possion reconstruction…and MVS-texture”). Fang teaches all of the above, but as can be seen in the above mapping, the appearance generation module does not specifically communicate with a NeRF and PeRF module that construct a base 3D room model as rather Fang teaches to provide a representation of the room capturing spatial depth and a refined model of the room through obtaining a depth map and surface reconstruction which captures spatial depth and a refining to generate a fully refined 3D room model through a texturing processing. Thus Fang stands as a base system upon which the claimed invention can be seen as an improvement through utilizing such NeRF and PeRF modules to construct the base 3D room model and fully refined 3D modules generated from the panoramic images as this could overcome issues with stretched textures and occluded objects such as explained in section 1.7 of the “Appendix” disclosing “the generated 3D room still contains incomplete structures in invisible area” and “obviously stretched texture because of the occlusion and poor performance of the panoramic depth estimator” and “problem is mainly caused by the inaccurate depth estimation and the next mesh reconstruction and mesh texturing process”. In the same field of endeavor relating to generations of 3D models from a panorama image, Wang teaches that it is known to convert a single panoramic image into a 3D volumetric representation of the scene and objects depicted in an image using a NeRF and PeRF module wherein the NeRF module communicates with an appearance generation module that provides a panoramic image and is configured to construct a base 3D room model based on the panoramic image, producing a representation of the room by capturing spatial depth (see Wang, Abstract, teaching “to lift up a 360-degree 2D scene to a 3D scene” and “first predict a panoramic depth map as initialization given a single panorama and reconstruct visible 3D regions with volume rendering” and “introduce a collaborative RGBD inpainting approach into a NeRF for completing RGB images and depth maps from random views, which is derived from an RGB Stable Diffusion model and a monocular depth estimator” which is further explained in section 3.2 and figure 1 teaching “given a single RGB panorama, we predict its depth map with a depth estimation model [45] and train a NeRF with both the RGB panorama and the depth map as initialization” such that this functions as a base 3D room model generated from the panoramic image which captures spatial depth as further explained in section 3.3 teaching “given a single panorama Irgb, we predict its depth map as Id by the depth estimator fd. We integrate Id into the training of a NeRF” and “We optimize L′ nerf to train a NeRF to be aware of the 3D shape of the visible regions. Since only a single panorama is available, the trained NeRF fnerf only works for visible regions. We then complete the invisible regions of the scene as described below” such that this is a base 3D room model and this reconstructing of the visible 3D regions using volume rendering and a depth map corresponds to such construction of a base 3D room model from a panoramic image and can be considered base as it will be completed into a fully refined 3D room model as explained below) and the PeRF module refines the base 3D model by enhancing visual coherence so as to generate a fully refined 3D room model (see Fang as explained below teaching generation of the base model where as in section 3.3 it is taught “given a single panorama Irgb, we predict its depth map as Id by the depth estimator fd. We integrate Id into the training of a NeRF” and “We optimize L′ nerf to train a NeRF to be aware of the 3D shape of the visible regions. Since only a single panorama is available, the trained NeRF fnerf only works for visible regions. We then complete the invisible regions of the scene as described below” such that the completion of such regions and additional processing explained below correspond to enhancing visual coherence where such processing in relation to the panoramic images converted to neural radiance fields corresponds to the system acting as a PeRF module as claimed where panoptic-enhanced is interpreted to mean that the radiance field module utilizes panoramic image (panoptic image data) in connection with a neural radiance field to provide such refinement as in section 3.2 teaching “note that the newly completed geometry of the invisible region may have conflicts with the observations of the already-learned reference views. To this end, we propose a progressive inpainting-and-erasing strategy to compute a mask of the conflicted regions and eliminate these regions from the NeRF training. We then fine-tune the panoramic NeRF with the reference panorama views and the new panorama generated from the random viewpoint. After fine-tuning, the new panorama is also set as a reference panorama for the upcoming NeRF learning of the random viewpoint. We progressively enlarge the range of viewpoints of the NeRF by randomly sampling a camera pose and training the NeRF until convergence” where this corresponds to the final stage of the pipeline as in figure 1 and section 3.5 teaching “proposed collaborative RGBD inpainting does not guarantee consistent geometry between the reference panorama and the new panorama from a random novel view” and “occlusion makes a part of visible regions of reference view become invisible regions, leading to geometry conflict across different views” such that there is not visual coherence and this is solved by the PeRF module refining the base 3D room model where “We iterate inpainting invisible regions and eliminating conflicted regions from sampled novel views until the algorithm converges to a completed 3D geometry”). Thus Wang teaches known techniques applicable to the base system which is ready for improvement. Therefore it would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Fang such that instead of generating a 3D model through depth map prediction and poisson surface reconstruction from a panoramic image, a NeRF and PeRF module as taught by Wang above would be utilized to take advantage of NeRF technology for 3D visualization and the PeRF module for refining of such NeRF models in relation to panoramic images and the predictable result of such use of Wang’s technique would predictably be a 3D base model generated from a panoramic image such as generated in Fang into a 3D NeRF based model which is then refined by a PeRF module which generates a fully complete 3D room model. The results would be predictable as Wang already teaches that all that is needed is a single panorama image to generated the NeRF models for refinement and Fang already teaches to generate a panoramic image to supply to a visualization component for 3D model generation. The combination would result in an improved system as it would deal with the incomplete structures due to occlusion and the poor performance of the depth estimator for the panoramic images as identified in relation to Fang, and would improve the visualization through the progressive inpainting and erasing strategy within the NeRF and PeRF modules which deals with such problems. Regarding claim 2, Fang as modified teaches all that is required as applied to claim 1 above and further teaches wherein the generation of the 3D layout of the room by the appearance generation module and the generation of the base 3D room model by the NeRF module are distinct stages performed sequentially (see Fang as modified where Fang in combination already teaches generation of the 3D layout of the room by the appearance generation module as distinct from generation of the base 3D room model as explained above and as can be seen in figure 2 of Fang where the appearance generation module acts with the layout generation module to transform the 3D layout of the room using the equirectangular projection to get a semantic layout which is converted to a single panoramic image which is distinct from that image then being fed to the reconstruction stages and visualization stages, and as already modified by Wang these stages would remain distinct as Wang takes in a panoramic image distinct from however that panoramic image was generated to depict or capture some room scene ). Regarding claim 3, Fang as modified teaches all that is required as applied to claim 1 above and further teaches wherein the scene code generator receives the scene description from the text processing module and processes it through multiple layers of a QKV (Query, Key, Value) mechanism via the scene code diffusion model, which is configured to gradually refine the scene description into a structured representation (see Fang as modified where Fang teaches this through the diffusion model where such processing through multiple layers using QKV values simply describes a basic diffusion model and as can be seen in figure 2 of Wang this is described in the “scene code diffusion” portion and as in sections 3-3.1 teaching “use a holistic scene code to parametrize the indoor scene and design a diffusion model to learn its distribution” where “the holistic scene code is generated from text” and “given a 3D scene S with m walls and n furniture items, we represent the scene layout as a holistic scene code” and “encode each object oj as a node with various attributes” and “With the scene code definition, we build a diffusion model to learn its distribution” and “denoising network ϵθ takes the scene code xt, text prompt y, and timestep t as input, and denoises them iteratively to get a clean scene code ˆx0. Then we represent ˆx0 as a set of orientated bounding boxes of various semantic types to facilitate interactive editing” such that that refines it into a structure representation of the bounding boxes). Regarding claim 4, Fang as modified teaches all that is required as applied to claim 3 above and further teaches wherein, during a translation by the scene code diffusion model, scene code noise is embedded by the scene code diffusion model to introduce variation and flexibility (see Fang as modified where Fang in combination already teaches this with Fang teaching the layout generation and scene code generator as explained above where the scene code diffusion module in Fang utilizes scene code noise as recited where “With the scene code definition, we build a diffusion model to learn its distribution” and “Given a clean scene code X0, the diffusion process gradually adds Gaussian noise to x0, until the resulting distribution is Gaussian” and “neural network is trained to reverse that process” and uses “the noise estimator which aims to find the noise ϵ added into the input x0. Here, y is the text embedding of the input text prompts. The denoising network is a 1D UNet (Ronneberger et al., 2015), with multiple self-attention and cross-attention layers designed for input text prompts. The denoising network ϵθ takes the scene code xt, text prompt y, and timestep t as input, and denoises them iteratively to get a clean scene code ˆx0. Then we represent ˆx0 as a set of orientated bounding boxes of various semantic types to facilitate interactive editing” where this introduces variation and flexibility). Regarding claim 5, Fang as modified teaches all that is required as applied to claim 1 above and further teaches wherein the layout generation module uses the oriented bounding boxes which represent key objects or key factors in a scene of the room for providing a modular way to define geometry and arrangement of the room (see Fang as modified, with Fang teaching the layout generation module as explained above and teaches that the module uses oriented bounding boxes that represent key objects or key factors in a scene of the room such as furniture and walls and doors and windows as in section 3 and figure 2 teaching “the holistic scene code is generated from text, we recover the room as a set of orientated bounding boxes of walls and objects” and “we synthesize a scene code from the text input and convert it to a 3D bounding box representation to facilitate editing” and “we consider not only furniture but also walls, doors, and windows to define the room layout” and “given a 3D scene S with m walls and n furniture items, we represent the scene layout as a holistic scene code” and “encode each object oj as a node with various attributes” and “With the scene code definition, we build a diffusion model to learn its distribution” and “denoising network ϵθ takes the scene code xt, text prompt y, and timestep t as input, and denoises them iteratively to get a clean scene code ˆx0. Then we represent ˆx0 as a set of orientated bounding boxes of various semantic types to facilitate interactive editing”). Regarding claim 6, Fang as modified teaches all that is required as applied to claim 1 above and further teaches a layout modification module communicating with a layout generation module and configured to allow the user to modify the 3D layout of the room interactively (see Fang as modified where Fang teaches the layout generation module and a layout modification module in communication allowing the user to modify the 3D layout of the room interactively as in section 3 and figure 2 teaching “users can edit these bounding boxes by dragging objects to adjust their semantic types, positions, or scales, enabling the customization of 3D scene results according to the user’s preferences” and “we synthesize a scene code from the text input and convert it to a 3D bounding box representation to facilitate editing” such that here the user is able to change the 3D layout of the room interactively), wherein the layout modification module is further configured to provides an interface such that the user is permitted to adjust at least one of the scene codes based on the displayed 3D layout of the room via the interface and that the applied scene codes are directly changed (see Fang as modified where Fang teaches such a layout modification module which adjusts scene codes based on the displayed 3D layout of the room via the interface as in section 3 and figure 2 teaching “interactive editing” of the “layout bounding box” where as in section 3.1 the scene codes correspond to an “object” with “various attributes” encoded corresponding to those as in section 3.3 teaching “user can modify the generated 3D room by changing the position, semantic class, and size of object bounding boxes” and “will update the panorama according to the user’s input” such that this adjusts the scene codes to get a changed semantic layout). Regarding claim 7, Fang as modified teaches all that is required as applied to claim 6 above and further teaches wherein the layout modification module is further configured to update the applied scene codes to reflect modification by the user, maintaining consistency between the visual representation and scene data for the room (see Fang as modified as explained above where Fang teaches such a layout modification module which adjusts scene codes based on the displayed 3D layout of the room via the interface as in section 3 and figure 2 teaching “interactive editing” of the “layout bounding box” where as in section 3.1 the scene codes correspond to an “object” with “various attributes” encoded corresponding to those as in section 3.3 teaching “user can modify the generated 3D room by changing the position, semantic class, and size of object bounding boxes” and “will update the panorama according to the user’s input” such that this adjusts the scene codes to get a changed semantic layout and then the semantic layout from such changed scene codes is applied as in the editing process which maintains consistency between the visual representation and scene data for the room). Regarding claim 8, Fang as modified teaches all that is required as applied to claim 1 above and further teaches wherein the semantic layout captures spatial relationships and object placements within the room, thereby generating the visual representation in response to the user input (see Fang as modified where Fang teaches such semantic layout and that it captures spatial relationships and object placements within the room, thereby generating the visual representation in response to the user input as can be seen as in figure 2 for example with the semantic layout capturing the spatial relationships and generating a visual representation such as the semantic layout and then generated panoramic image). Regarding claim 9, Fang as modified teaches all that is required as applied to claim 1 above and further teaches wherein the appearance generation module generate the single panoramic image of the room via employs loop-consistent sampling using a ControlNet model (see Fang as modified where Fang teaches such use of such sampling to generate the single panoramic image as in figure 2 and section 3 teaching “the Appearance Generation Stage, we obtain an RGB panorama through a pre-trained latent diffusion model to represent the room texture. Specifically, we project the generated layout bounding boxes into a semantic segmentation map representing the layout. We then fine-tune a pre-trained ControlNet (Zhang & Agrawala, 2023) model to generate an RGB panorama from the input layout panorama. To ensure loop consistency, we propose a novel loop-consistent sampling during the inference process” and section 3.2 teaching “To condition ControlNet on the scene layout, we convert the bounding box representation into a 2D semantic layout panorama through equirectangular projection. In this way, we get a pair of RGB and semantic layout panoramic images for each scene” and “panorama should be loop-consistent. In other words, its left and right should be seamlessly connected. Although the panoramic horizontal rotation in data augmentation may improve the model’s implicit understanding of the expected loop consistency, it lacks explicit constraints and might still produce inconsistent results. Therefore, we propose an explicit loop consistent sampling mechanism in the denoising process of the latent diffusion model”). Regarding claim 10, Fang as modified teaches all that is required as applied to claim 1 above and further teaches a panoramic update module communicating with the PeRF module and configured to update the panorama or the fully refined 3D room model dynamically based on any modifications made by the user, when the user reviews the fully refined 3D room model presented by the PeRF module and makes modifications (see Fang as modified where Fang as modified teaches the PeRF module and Fang already teaches modifications that can be made by a user as in section 3.3 which allows to generate a modified panorama such that Fang as modified already teaches that changing of the 3D room layout can result in a changed panorama which would then be constructed and refined to generate the final model of the layout as already taught by Wang as combined). Regarding claims 11-20, the instant claims recite the same limitations as claims 1-10, respectively, but embodied as a method using a system for performing the same limitations, though not limited by the interpretation under 35 U.S.C. 112(f). Note that while the claims recite a “system for computer-based 3D indoor scene assessment” as well as a series of “modules” and functional elements with no explicit structure, the claim is considered directed to “The method using” such a system and thus is a process and is not considered directed toward software, per se. Note that the limitations required by the method are interpreted under the broadest reasonable interpretation to require use of a specially programmed computer in order to carry out such complex steps including use of “a scene code diffusion model” which takes in computer-based data and provides outputs to other modules which require a computer to perform their functions as would be recognized by one having ordinary skill in the art. Regardless, the claims represent a broader embodiment of claims 1-10 and the limitations of claims 11-20 correspond to the limitations of claims 1-10, respectively. As claims 1-10 are rejected as obvious in view of Fang and Wang, then the broader claims 11-20, respectively; are also rejected on the same grounds. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Feng et al (Feng W, Zhu W, Fu TJ, Jampani V, Akula A, He X, Basu S, Wang XE, Wang WY. LayoutGPT: Compositional Visual Planning and Generation with Large Language Models. arXiv preprint arXiv:2305.15393. 2023 May 24.) – see Abstract and figure 1 and 2 and section 3.1-3.4 teaching using a text prompt to generate a 3D layout including 3D bounding boxes based on the text prompt being passed through an LLM to provide a 3D layout upon which a visualization of objects in a room can be generated through retrieving objects specified by the layout. Thus text is used to generate a 3D layout but the technique differs from that claimed as it does not use a diffusion model to generate the scene codes using a similar type of scene code generator and also does not teach the connected processing functions of using the 3D layout to generate a semantic layout upon which a panoramic image is generated which is then the basis for a 3D model to be generated. Rather the 3D model is generated from the objects specified directly by the 3D layout. Zhang et al (Zhang J, Li X, Wan Z, Wang C, Liao J. Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields. arXiv preprint arXiv:2305.11588. 2023 May 19.) – see Abstract and figure 2 and section III teaching using a text prompt to generate a 3D scene where a diffusion model is used to generate an image and a depth image is derived from that and a NeRF is constructed to provide a visualization of a 3D scene from the text prompt. However there is no similar layout generation stage and the NeRF is not based on a panoramic image generated from a semantic layout which is an equirectangular projection of a 3D layout generated from text as in the claimed invention. Hollein et al (Höllein, Lukas et al. “Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models.” March 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (2023): 7875-7886.) – see Abstract and figures 1 and 2 and section 3 teaching to create a textured mesh of a complete scene including of a room from text input. Again an image is generated from text, but the 3D scene is generated iteratively in a manner that differs from the claimed invention as again there is no similar layout generation stage and the NeRF is not based on a panoramic image generated from a semantic layout which is an equirectangular projection of a 3D layout generated from text as in the claimed invention. Wen et al (Wen B, Xie H, Chen Z, Hong F, Liu Z. 3d scene generation: A survey. arXiv preprint arXiv:2505.05474. 2025 May 8.) – see entire document, note that Wen is a survey of 3D scene generation which does not qualify as prior art but is relevant as it contains descriptions of analogous prior art. The Examiner has reviewed the references that would qualify as prior art cited therein which include the above references for example. Ocal et al (Öcal BM, Tatarchenko M, Karaoglu S, Gevers T. SceneTeller: Language-to-3D Scene Generation. arXiv preprint arXiv:2407.20727. 2024 Jul 30..) - see figures 1 and 2 and Abstract and Section 3 teaching “Given a textual prompt in natural language delineating the desired spatial positions and orientations of objects within the scene, a 3D scene layout is generated using in-context learning. An initial 3D scene is assembled for the predicted layout, which is then fitted with a 3D Gaussian Splatting representation. This representation is then used to stylize the scene according to the user-provided text prompt, and subsequently render the final images of the scene” such that as can be seen in figure 2, there is use of a text prompt to generate a 3D layout using a layout generation module, but again the manner in which the 3D model is generated differs from the claimed technique. Rather than using the text codes to generate the 3D layout of bounding boxes with semantic categories, which is then projected to generate a semantic image, upon which a panoramic image is then used as the basis for the 3D model as claimed, Ocal generates the 3D layout but uses the layout and semantic information to retrieve 3D models of matching objects, and then these are rendered using Gaussian splatting. Schult et al (Schult J, Tsai S, Höllein L, Wu B, Wang J, Ma CY, Li K, Wang X, Wimbauer F, He Z, Zhang P. ControlRoom3D: Room Generation using Semantic Proxy Rooms. arXiv preprint arXiv:2312.05208. 2023 Dec 8.) – see Figures 1 and 2, Abstract and section 3, teaching a user defined 3D room layout based on bounding boxes corresponding to semantic classes and use of a text prompt to guide the generation of a 3D model which includes generation of a panorama image and a depth mapping related to that panorama image, which is then used to determine a 3D mesh, which is then textured using a depth alignment from an image rendered based on the text prompt. Thus a 3D layout with semantic bounding boxes is used to guide a reconstruction, but the use of the text prompt differs, and importantly there is not equirectangular projection of a 3D semantic layout which was generated from the text prompt which is used to generate a panorama image as claimed. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SCOTT E SONNERS whose telephone number is (571)270-7504. The examiner can normally be reached Mon-Friday 9-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SCOTT E SONNERS/Examiner, Art Unit 2613 /XIAO M WU/Supervisory Patent Examiner, Art Unit 2613 1 Fang C, Dong Y, Luo K, Hu X, Shrestha R, Tan P. Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints. arXiv e-prints. 2023 Oct:arXiv-2310. 2 Wang G, Wang P, Chen Z, Wang W, Loy CC, Liu Z. PERF: Panoramic Neural Radiance Field from a Single Panorama. arXiv preprint arXiv:2310.16831. 2023 Oct 25.
Read full office action

Prosecution Timeline

Oct 16, 2024
Application Filed
Jul 13, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705829
METHODS, STORAGE MEDIA, AND SYSTEMS FOR GENERATING A THREE-DIMENSIONAL LINE SEGMENT
2y 3m to grant Granted Aug 11, 2026
Patent 12700118
SYSTEMS AND METHODS FOR PROCESSING CAPTURED IMAGES
2y 12m to grant Granted Aug 04, 2026
Patent 12700114
METHOD FOR CHARACTERISING A ZONE OF LAND INTENDED FOR THE INSTALLATION OF PHOTOVOLTAIC PANELS
2y 7m to grant Granted Aug 04, 2026
Patent 12657818
DISTORTION CORRECTION FOR ENVIRONMENT VISUALIZATIONS WITH WIDE ANGLE VIEWS
2y 11m to grant Granted Jun 16, 2026
Patent 12657769
AUGMENTING 3D PATIENT SCANS WITH CAD MODEL OF IMPLANTED OBJECT
2y 6m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
69%
Grant Probability
81%
With Interview (+11.9%)
3y 3m (~1y 5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 386 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month