Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d).
The certified copy of CN 202410339067.1 has been received on 04/21/2025.
Claim Objections
Claim 14 is objected to because of the following informalities:
In Claim 14, "device according to claim 6" is suggested to read "device according to claim 13”. For examination purposes, Examiner will interpret the language in Claim 14 to read as “device according to Claim 13”
Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a learnable module” in Claims 3, 10, and 17.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
For the sake of further prosecution, Examiner will treat “the learnable module configured to convert the two-dimensional image into the panoramic image” in Claims 3, 10, and 17 all as hardware or software configured to perform their respective recited functions/operations.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 8, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (US 12567208 B2), hereinafter referenced as Chen, in view of Fang et al. (“CTRL-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints”, 10/2023), hereinafter referenced as Fang, and in further view of Hannuksela et al. (US 2019/0268599 A1), hereinafter referenced as Hannuksela.
Regarding Claim 1, Chen discloses a method for three-dimensional scene generation (Chen: [Abs], discloses a method for constructing three-dimensional models of buildings <scene generation>), comprising:
obtaining multi-view information in a plurality of (Chen: [Col 4, ln 5], discloses performing multi-view projection correction and employing a rectangular projection spherical panorama model <multi-view projection correction requires intrinsic and extrinsic camera parameters, interpreted as multi-view information>; [Col 2, ln 23-26], discloses generating sequential multi-view images by correcting the multiple view projections of 360-degree <plurality of views> panoramic images <including a multi-view image and a panoramic image>);
performing depth estimation on the panoramic image, to determine a sparse point cloud corresponding to the panoramic image (Chen: [Col 6, ln 59 -Col 7, ln 38], discloses performing multi-view projection correction followed by dense matching, and structure from motion <depth estimation techniques> on the 360-degree panoramic data, to determine a sparse point cloud);
and generating, based on the multi-view image, the multi-view information, and the sparse point cloud, a three-dimensional scene model (Chen: [Col 2, lns 15-34], discloses generating, based on a sequence of multi-view images, the multi-view information, and the sparse point cloud, a three-dimensional scene model).
Chen fails to explicitly disclose
obtaining a target text, and generating a panoramic image described by the target text;
and generating, based on the multi-view image, the multi-view information, and the sparse point cloud, a three-dimensional scene model described by the target text
obtaining multi-view information in a plurality of preset views
However, Fang discloses
obtaining a target text, and generating a panoramic image described by the target text (Fang: [Fig. 2, pg. 5], as shown below, illustrates a method of receiving a text prompt <target text>, and generating a panorama <panoramic image> described by the target text);
PNG
media_image1.png
350
912
media_image1.png
Greyscale
and generating, based on a panorama, a semantic layout, and a layout bounding box, a three-dimensional scene model described by the target text (Fang: [Fig. 2, pg. 5], illustrates generating, based on a panorama, a semantic layout, and a layout bounding box, a textured mesh <three-dimensional model> described by the text prompt <target text>).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method disclosed by Chen by using a panoramic image generated as described by text to create the three-dimensional scene model as taught by Fang. One of ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this modification to automatically generate a customized three-dimensional scene model.
The combination of Chen and Fang fail to disclose
obtaining multi-view information in a plurality of preset views
However, Hannuksela discloses
obtaining multi-view information in a plurality of preset views (Hannuksela: [0003-0004], discloses obtaining a cubemap <multi-view information in a plurality of preset views> of panoramic data)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method disclosed by the combination of Chen and Fang by having multi-view information in a plurality of preset views as taught by Hannuksela. One of ordinary skill of the art before the effective filing date of the claimed invention would have been motivated to make this modification to simulate immersive background environments at a lower computation cost.
Regarding Claim 8, it recites limitations similar to Claim 1 but as an electronic device. As shown in the rejection, the combination of Chen, Fang, and Hannuksela disclose the method of Claim 1. The combination of Chen, Fang, and Hannuksela further disclose
An electronic device, comprising: one or more processors; and a storage apparatus having one or more programs stored thereon, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to (Hannuksela: [0327], discloses a mobile device comprising a processor and a memory, where the memory stores software <one or more programs stored thereon, when executed by a processor cause a processor to perform a method>): …
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method as disclosed by the combination of Chen, Fang, and Hannuksela by applying the method to an electronic device as further taught by Hannuksela. One of ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this modification to consolidate the performed method to one device with real-time data processing and instruction execution.
Regarding Claim 15, it recites limitations similar to Claims 1 and 8 but as a non-transitory computer-readable medium. As shown in the rejection, the combination of Chen, Fang, and Hannuksela disclose the method and device of Claims 1 and 8 respectively. The combination of Chen, Fang, and Hannuksela further disclose
A non-transitory computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, causing the processor to perform (Chen: [Claim 10], discloses a non-transitory computer-readable medium storing executable instructions, wherein, when executed by a processor, performs a method): …
Claims 2-3, 9-10, and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Chen, Fang, and Hannuksela in view of Wang et al. (“Customizing 360-Degree Panoramas through Text-to-Image Diffusion Models”, 2023), hereinafter referenced as Wang.
Regarding Claims 2, 9, and 16, the combination of Chen, Fang, and Hannuksela disclose the method, device, and medium of Claims 1, 8, and 15 respectively. The combination of Chen, Fang, and Hannuksela teach a diffusion model representing a correspondence between a text and a panoramic image (Fang: [Fig. 2]), they fail to explicitly disclose the limitations of Claims 2, 9, and 16, however, Wang discloses wherein the generating a panoramic image described by the target text comprises:
generating, using a pre-trained target diffusion model, the panoramic image described by the target text, wherein the target diffusion model is used to represent a correspondence between a text and a panoramic image (Wang: [3. Methodology, pgs. 3-4], discloses generating, using a customized pre-trained T2I diffusion model <pre-trained target diffusion model>, a panoramic image described by text, wherein the diffusion model is used to generate the output panorama of input text <represent a correspondence between a text input and a panoramic output>).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method, electronic device, or non-transitory computer-readable medium as disclosed by the combination of Chen, Fang, and Hannuksela by using a pretrained diffusion model as taught by Wang. One of ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this modification to eliminate massive computational costs and development time.
Regarding Claims 3, 10, and 17, the combination of Chen, Fang, Hannuksela, and Wang disclose the method, device, and medium of Claims 2, 9, and 16 respectively. The combination of Chen, Fang, Hannuksela, and Wang further disclose wherein
the target diffusion model is a model obtained by performing a target operation on an original diffusion model, wherein the original diffusion model is used to represent a correspondence between a text and a two-dimensional image, (Wang: [3.2. Customizing Models with LoRA, pg. 4], discloses customizing <target operation> a pre-trained T2I diffusion model <where the pre-trained T2I diffusion model is interpreted as the original diffusion model and the customized diffusion model is interpreted as the target diffusion model>; [Abs, 1. Introduction, pg. 1], discloses wherein the T2I diffusion model is used to represent a correspondence between text input and an image output <T2I stands for text-to-image>)
and the target operation comprises: freezing a parameter of the original diffusion model, and inserting a learnable module into the original diffusion model, wherein the learnable module is configured to convert the two-dimensional image into the panoramic image (Wang: [1. Introduction, pg. 1], discloses customizing the T2I Diffusion model by employing low-rank adaption technology to fine-tune the pretrained model; [3.2. Customizing Models with LoRA, pg. 4], discloses inserting a LoRA technology <interpreted as the learning module> into the diffusion model to fine tune the model for generating 360-degree panoramic images, for LoRA to fine-tune a model, the pretrained model weights <parameters> are frozen).
PNG
media_image2.png
284
634
media_image2.png
Greyscale
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method, electronic device, or non-transitory computer-readable medium disclosed by the combination of Chen, Fang, Hannuksela, and Wang by customizing a diffusion model with LoRA as further taught by Wang. One of ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this modification for efficient fine-tuning and modular stacking.
Claims 4, 11, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Chen, Fang, Hannuksela, and Wang in view of Xing et al. (CN 117973476 A), hereinafter referenced as Xing.
Regarding Claims 4, 11, and 18, the combination of Chen, Fang, Hannuksela, and Wang disclose the method, device, and medium of Claims 3, 10, and 17 respectively. The combination of Chen, Fang, Hannuksela, and Wang further disclose wherein the learnable module comprises
a low-rank matrix obtained by (Wang: [3.2. Customizing Models with LoRA, pg. 4], discloses rank decomposition matrices <fundamentally low-rank matrices> using low-rank adaption technology).
original diffusion model (Wang: [3.2. Customizing Models with LoRA, pg. 4], discloses an pre-trained diffusion model <original diffusion model>)
The combination of Chen, Fang, Hannuksela, and Wang fail to explicitly disclose
a low-rank matrix obtained by decomposing a parameter matrix of the original diffusion model
However, Xing discloses
a low-rank matrix obtained by decomposing a parameter matrix of the original (Xing: [0013], discloses decomposing the parameter matrix of an original model <therefore obtaining a low-rank matrix>)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method, electronic device, or non-transitory computer-readable medium disclosed by the combination of Cheng, Fang, Hannuksela, and Wang, by decomposing parameter matrices of the original diffusion model as taught by Xing. One of ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this modification to reduce memory storage and computational requirements, accelerating the model’s inference and training speeds.
Claims 5, 12, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Chen, Fang, and Hannuksela in view of Parker et al. (US 2017/0290721 A1), hereinafter referenced as Parker.
Regarding Claims 5, 12, and 19, the combination of Chen, Fang, and Hannuksela disclose the method, device, and medium of Claims 1, 8, and 15 respectively. The combination of Chen, Fang, and Hannuksela fail to disclose the limitations of Claims 5, 12, and 19, however, Parker discloses
determining a current view, and outputting scene information in the current view based on the current view and the three-dimensional scene model (Parker: [Abs], discloses determining a pose <current view> and displaying <outputting scene information> a display frame of the pose based on the determined pose and the 3D world <scene model>).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method, electronic device, or non-transitory computer-readable medium as disclosed by the combination of Chen, Fang, and Hannuksela by rendering a 3D scene’s current view as taught by Parker. One of ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this modification for an optimized and immersive visual experience.
Claims 6-7, 13-14, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Chen, Fang, and Hannuksela in view of Chng et al. (“Gaussian Activated Neural Radiance Fields for High Fidelity Reconstruction and Pose Estimation”, 2022), hereinafter referenced as Chng.
Regarding Claims 6, 13, and 20, the combination of Chen, Fang, and Hannuksela disclose the method, device, and medium of Claims 1, 8, and 15 respectively. The combination of Chen, Fang, and Hannuksela fail to disclose the limitations of Claims 6, 13, and 20, however, Chng discloses wherein the three-dimensional scene model comprises
a three-dimensional Gaussian radiance field (Chng: [1. Introduction, pg. 1], discloses a scene representation of Gaussian Activated Neural Radiance Fields)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method, electronic device, or non-transitory computer-readable medium as disclosed by the combination of Chen, Fang, and Hannuksela by representing a 3D scene with 3D Gaussians as taught by Chng. One of ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this modification for real-time photorealistic rendering.
Regarding Claims 7 and 14, the combination of Chen, Fang, Hannuksela, and Chng disclose the method and device of Claims 6 and 13 respectively. The combination of Chen, Fang, Hannuksela, and Chng further disclose
for each of the plurality of views, projecting the three-dimensional Gaussian radiance field to the view, comparing a projected image in the view with a multi-view image corresponding to the view, to obtain a loss value, and optimizing a parameter of the three-dimensional Gaussian radiance field with the loss value (Chng: [4.3. Implementation Details, pg. 12], discloses for the optimized poses <plurality of views>, aligning <projecting> the optimized pose to the ground truth <comparing> to obtain the translation error <loss value> for the pose; [4.1 2D Planar Image Alignment, pgs. 8-11], discloses warping parameters for optimizing a neural image representation, the 3D Gaussian Activated Neural Radiance Field representation of the paper, optimizing the weight updates using backpropagation and gradient descent <based on loss values>).
PNG
media_image3.png
292
1124
media_image3.png
Greyscale
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply and/or modify the method, electronic device, or non-transitory computer-readable medium as disclosed by the combination of Chen, Fang, Hannuksela, and Chng by backpropagation as further taught by Chng. One of ordinary skill in the art before the effective filing date of the claimed invention would have been motivated to make this modification for better alignment between model outputs and desired outcomes.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Tang et al. (“MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion”, 2023) discloses … .
Höllein et al. (“Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image models”, 2023) discloses depth estimation of image generated by text to build a 3-dimensional model.
Sun et al. (US 2019/0378330 A1) discloses a method of obtaining multi-view images from panoramic data, then reconstructing them into a sparse point cloud, then creating a three-dimensional model.
Pardeshi et al. (US 2021/0192684 A1) discloses generating panoramas then generating spherical panoramic images.
Hu et al. (“LoRA: Low-Rank Adaption of Large Language Models”, 2021) introduces low-rank adaption, which freeze the pre-trained model weights and injects trainable rank decomposition matrices into each layer of a transformer’s architecture.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ISABELLA OCHSNER whose telephone number is (571)272-9322. The examiner can normally be reached 9:30 - 6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona Faulk can be reached at (571) 272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/I.O./Examiner, Art Unit 2618
/DEVONA E FAULK/Supervisory Patent Examiner, Art Unit 2618