DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 8/10/2026 has been entered.
Response to Amendment
Claims 1, 15-16 are amended. Now claims 1-16 are pending in the instant application.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-16 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Thu et al. (Nguyen-Phuoc, Thu H., et al. "Rendernet: A deep convolutional network for differentiable rendering from 3d shapes." Advances in neural information processing systems 31 (2018).; hereinafter Thu) in view of BROOK et al. (US 20200349351 A1, hereinafter BROOK).
Regarding claim 1, Thu discloses a method (Title, Abstract; computer-implemented feed-forward CNN rendering architecture. Fig. 1, §3), comprising:
at a device (ibid, computer-implemented feed-forward CNN rendering architecture. Fig. 1, §3):
processing an input using a feed-forward neural network to generate a three- dimensional (3D) representation of a scene, wherein the input includes a plurality of labeled voxels describing the scene in 3D (receives a 3D voxel grid as input. Fig. 1: 3D shape, 64×64×64×1. §3.2 defines voxel grid
V
having height, width and depth. RenderNet expressly states its architecture renders from a “3D voxel grid input.”
By framing the rendering process as a feed-forward CNN, RenderNet has the ability to learn to express different shaders with the same network architecture, page 2, ¶3.
Architecture includes 3D convolutions → projection unit → 2D convolutions. Fig. 1, §3
Fig. 1 processes the voxel input through 3D CONVOLUTION, producing learned 3D feature tensors, including 32×32×32×16, before the projection unit. Fig. 1, §§3–3.2
); and
generating a two-dimensional (2D) image of the scene from a given viewpoint, using the 3D representation of the scene (The 3D features enter the projection unit, followed by 2D CONVOLUTION, producing FINAL OUTPUT 512×512×3. RenderNet expressly states that the CNN with projection unit produces a rendered 2D image. Fig. 1, §3.2.
Fig. 1: Camera pose. §3.1: rigid-body transformation converts the voxel grid into camera coordinates. Viewpoint is parameterized by azimuth, elevation, and distance
R
.
The learned 32×32×32×16 3D feature representation is passed to the projection unit, which converts the 3D features into 32×32×512 2D features, followed by 2D convolutions generating the final image. Fig. 1, §3.2.).
Thus, RenderNet teaches every limitation except expressly requiring its input voxels to be “labeled.”
Brook, however, cures precisely that deficiency. It expressly teaches semantically-labeled 3D voxels describing a scene, including semantic/class information associated with individual voxels. For example, the disclosed classes include structural and non-structural scene objects such as wall, floor, ceiling, chair, desk, and table (see Abstract, ¶0013-0015, figs. 1f-1h, ¶0032-0033, ¶0048, step 406, fig. 4, step 506-510, fig. 5 … etc.).
Furthermore, in ¶0033, Brook discloses, the semantic labelling algorithm can be manifest as a neural network, such as a convolutional neural network (CNN). The CNN can receive scene information, such as images, depth information, and/or surface normal information. The CNN can analyze the scene information on a pixel-by-pixel or groups of pixels basis. The CNN can output a class and confidence for each pixel or group of pixels.
Thu thus teaches processing a 3D voxel-grid input using a feed-forward CNN to generate a learned 3D representation and generating a 2D image therefrom according to a given camera viewpoint (Fig. 1; §§3–3.2), but does not expressly disclose that the input voxels are labeled. Microsoft '351 teaches a semantically-labeled 3D voxel representation of a scene, wherein semantic/class information is associated with voxels (Fig. 3, elements 308, 312, 314, 316).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (AIA ) to apply Microsoft's semantic labeling to the voxels of Thu’s RenderNet's 3D voxel-grid input so that the voxels additionally carry semantic/object information, thereby enabling scene elements represented by the voxels to be identified or distinguished. Such a combination would result in Thu processing an input comprising a plurality of labeled voxels describing a scene in 3D, while retaining Thu’s RenderNet's feed-forward 3D processing and viewpoint-dependent 2D rendering.
Regarding claim 3, Thu in view Brook discloses the method of claim 1, wherein each of the labeled voxels has a semantic meaning (Brook: Abstract, ¶0028, ¶0035, 0036).
Regarding claim 4, Thu in view Brook discloses the method of claim 1, wherein each of the labeled voxels is a voxel labeled with a descriptor of an object represented by the voxel (Brook: elements 308, 312, 314, 316, semantic labeling supplies class and instance data concerning objects represented In the 3D scene, fig. 3, ¶0034-0036).
Regarding claim 5, Thu in view Brook discloses the method of claim 1, wherein the feed-forward neural network further processes an input style code to generate the 3D representation of the scene (Thu: A novel convolutional neural network architecture that learns to render in different styles from a 3D voxel grid input. – page 2, first bullet in last ¶. Also see fig. 3, and 5, where Render style type is input to be transferred to the output).
Regarding claim 6, Thu in view Brook discloses the method of claim 1, wherein the 3D representation of the scene is a 3D feature map (Thu: 3D convolutions produce learned 32x32x16 features before projection, fig. 1).
Regarding claim 7, Thu in view Brook discloses the method of claim 1, wherein the 3D representation of the scene is a voxel grid with features (Thu’s Rendernet uses voxel grids and teaches assigning texture/features to corresponding voxels – Abstract. Texture network generates a 3D representation having the same WxHxD as the shape, thus is substantively similar to assigning a texture value to a corresponding voxel – §4.1).
Regarding claim 8, Thu in view Brook discloses the method of claim 1, wherein the 3D representation of the scene is a tri-plane representation (Brook: One example can identify planes in a semantically-labeled 3D voxel representation of a scene. – Abstract, figs. 1E-1L, shows 3 planes in XYZ format).
Regarding claim 9, Thu in view Brook discloses the method of claim 1, wherein the given viewpoint is defined based on an input camera pose (Thu: Fig. 1 identifies “camera pose”; §3.1 transformation uses azimuth, elevation and distance).
Regarding claim 10, Thu in view Brook discloses the method of claim 1, wherein the given viewpoint is controllable such that different 2D images of the scene are renderable from different given viewpoints, using the 3D representation of the scene (Thu: Fig. 3 shows rendering from different elevations, azimuths, and scaling, and describes outputs from “different views”).
Regarding claim 12, Thu in view Brook discloses the method of claim 1, further comprising, at the device: optimizing the 2D image of the scene (Thu: We train RenderNet using a pixel-space regression loss. We use mean squared error loss for colored images, and binary cross entropy for grayscale images. We use the Adam optimizer [38], with a learning rate of 0.00001. – p. 5, §4 Experiments, last ¶).
Regarding claim 13, Thu in view Brook discloses the method of claim 12, wherein the 2D image of the scene is optimized by a second feed-forward neural network (Thu: the limitation is understood met in the 2D convolution stage having 2 different cascaded stages of neural network, fig. 1).
Regarding claim 14, Thu in view Brook discloses the method of claim 1, wherein the feed-forward neural network generates the 3D representation of the scene from the input in a single feed-forward step (Thu: RenderNet expressly frames renderings as a “feed-forward CNN” amd processes the voxel input through its 3D convolutional network without iterative optimization during inference.).
Regarding clam 15, Thu in view Brook discloses a system (system of fig. 1), comprising:
process an input using a feed-forward neural network to generate a three- dimensional (3D) representation of a scene, wherein the input includes a plurality of labeled voxels describing the scene in 3D: and generate a two-dimensional (2D) image of the scene from a given viewpoint, using the 3D representation of the scene (see substantively similar claim 1 rejection above).
Thu in view Brook as combined is not found disclosing expressly that the system comprising, a non-transitory memory storage comprising instructions; and one or more processors in communication with the memory, wherein the one or more processors execute the instructions to, implement the steps.
However, Brook discloses in another embodiment of the invention a system (see fig. 6) comprises, a non-transitory memory storage comprising instructions; and one or more processors in communication with the memory, wherein the one or more processors execute the instructions to perform the inventive steps described therein (Brook: ¶0077-0080, claim 20, fig. 6).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (AIA ) to implement the RenderNet system of Thu in the system 600 of Brook, to obtain, system comprising, a non-transitory memory storage comprising instructions; and one or more processors in communication with the memory, wherein the one or more processors execute the instructions to, implement the aforementioned steps, because, combining prior art elements ready to be improved according to known method to yield predictable results is obvious (see MPEP §2143.I).
Regarding CRM claim(s) 16, although wording is different, the material is considered substantively equivalent to the system claim(s) 15 as described above.
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over Thu in view of BROOK and further in view of Wan et al. (US 20250148658 A1, hereinafter Wan).
Regarding claim 2, Thu in view Brook discloses the method of claim 1, except, wherein the input description is manually provided by a user.
However, Wan discloses a content generation method and apparatus (abstract), wherein, a user may manually input description keywords of an image (¶0115).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (AIA ) to modify the invention of Thu in view Brook, with the teaching of Wan of manually providing input description of the scene by a user, so as to supervise the semantic segmentation textual description of a scene automatically generated by the CNN (e.g., see fig. 1G of Brook), because, supervision can improve the accuracy of the overall system in scene discernment.
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Thu in view of BROOK and further in view of Hao et al. (US 20220180602 A1, hereinafter Hao).
Regarding claim 11, Thu in view of Brook discloses the method of claim 1, except, wherein the 2D image is generated by projecting the 3D representation of the scene to a 2D feature map via a neural radiance field rendering.
However, Hao discloses, method of image generation using neural networks based on one or more semantic features (abstract), wherein, 2D image is generated by projecting the 3D representation of the scene to a 2D feature map via a neural radiance field rendering (¶0050, 0062-0063).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (AIA ) to modify the invention of Ruggiero, using the teaching of Hao to generate 2D image by projecting the 3D representation of the scene to a 2D feature map via a neural radiance field rendering, because, combining prior art elements according to known method to yield predictable results is obvious. Furthermore, such combination would enhance the versatility of the overall system.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NURUN FLORA whose telephone number is (571)272-5742. The examiner can normally be reached M-F 9:30 am -5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at (571) 272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NURUN FLORA/Primary Examiner, Art Unit 2619