CTNF 18/894,176 CTNF 79238 DETAILED ACTION Notice of Pre-AIA or AIA Status 07-03-aia AIA 15-10-aia The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA. Information Disclosure Statement 06-52 AIA The information disclosure statement (IDS) submitted on 09/24/2024 was filed after the mailing date of the applicati on. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 103 07-06 AIA 15-10-15 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 07-20-aia AIA The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 07-23-aia AIA The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 07-21-aia AIA Claim s 1-3, 5, 7-13, 17, 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Liuet al. (US. Patent App. Pub. No. 2024/0249469, “Liu”, hereinafter) in view of Liu et al. (“PETR: Position Embedding Transformation for Multi-view 3D Object Detection”, “Liu 2”, hereinafter, retrieved from Springer Nature Link, October, 2022, (https://link.springer.com/chapter/10.1007/978-3-031-19812-0_31)) . As per claim 1, as shown in Fig. 2 and 3, Liu teaches a method comprising: obtaining a plurality of input images depicting an object and a set of 3D position embeddings (further addressed below with Liu 2) , wherein each of the plurality of input images depicts the object from a different perspective (¶ [26], obtaining 2D input images of a person viewing from front and back) ; encoding the plurality of input images to obtain a plurality of 2D features corresponding to the plurality of input images, respectively (Fig. 2, step 202, ¶ [26], “At block 202, 2D features of a 2D image are generated on the basis of performing feature extraction on the 2D image”. This is done by encoder 306, 308, Fig. 3 referring to ¶ [31]) ; generating, using a 2D-to-3D transformer, 3D features based on the plurality of 2D features and the set of 3D position embeddings (further addressed below. See Fig. 3, ¶ [34], generating 3D reconstruction features 316, 318) ; and generating a 3D model of the object based on the 3D features (¶ [34], generate final 3D models 324 and 326) . Liu does not expressly teach obtaining a set of 3D position embeddings for use in generating 3D features. However, in a very similar method of generating 3D model of an object based on extracted 2D features (see section Introduction, page 1, and also Fig. 2 on page 5), Liu 2 further teaches the above feature, i.e., obtaining a set of 3D position embeddings for use in generating 3D features (see Fig. 3, page 6, and section 3.3 3D position encoder). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method as taught by Liu 2 into the method as taught by Liu as addressed above, the advantage of which is provide a simple and elegant framework for multi-view 3D object detection (end of page 2 and beginning of page 3). As per claim 2, the combined teachings of Liu and Liu 2 also include generating the 3D features comprises: performing an attention mechanism on the plurality of 2D features and the set of 3D position embeddings (best understood, see Liu 2, Fig. 7, page 14, i.e., attention map). Thus, claim 2 would have been obvious over the combined references for the reason above. As per claim 3, the combined Liu-Liu 2 also teaches generating triplane features based on the 3D features, wherein the 3D model is generated based on the triplane features (interpreted as the three planes created by the 3D coordinates xyz in the 3D world space shown in Fig. 2, and also section 3.2 3D Coordinate Generator on page 5, of Liu 2). Thus, claim 3 would have been obvious over the combined references for the reason above. As per claim 5, as addressed the combined Liu-Liu 2 does also teach obtaining view information for each of the plurality of input images, wherein the plurality of input images are encoded based on the view information (Liu, Fig. 3, ¶ [31], encoding based on the viewing angle of input images) . As per claim 7, the combined Liu-Liu 2 does further teach obtaining a reference view encoding for a first image of the plurality of input images (Liu, Fig. 3 such as image encoding front view of a person 306) and a source view encoding for a second image of the plurality of input images (Liu, Fig. 3 such as image encoding back view of a person 308) , wherein the first image is encoded based on the reference view encoding and the second image is encoded based on the source view encoding (Liu, ¶ [31-32]). As per claim 8, as addressed, the combined Liu-Liu 2 impliedly teach obtaining view intrinsic parameters of each of the plurality of input images, wherein the plurality of input images are encoded based on the view intrinsic parameters (at best understood as the parameter of the image encoder described at ¶ [32] of Liu) . As per claim 9, the combined Liu-Liu 2 does further teach: generating, using the 2D-to-3D transformer, a plurality of image-specific output features corresponding to the plurality of input images, respectively (Liu, ¶ [40], “First, extracted feature 510 is reshaped to convert the 2D feature into 3D feature 512 . For example, a dimension of extracted feature 510 is 512*64*64. After the reshape operation, the dimension becomes 8*64*64*64”) ; and generating pose information for each of the plurality of input images based on the plurality of image-specific output features (Liu, ¶ [40], “In some embodiments, different camera poses can be specified to convert the 2D feature to features in different angles”) . As per claim 10, the combined Liu-Liu 2 also teaches the 2D-to-3D transformer is trained using a training set that includes a plurality of training images depicting different views of a scene (Liu, ¶ [31], training system shown in Fig. 3). Claim 11, which is similar in scope to claims 1 and 10 as addressed above, is thus rejected under the same rationale. As per claim 12, the combined Liu-Liu 2 does also teach wherein training the 2D-to-3D transformer comprises: generating an output image based on the 3D features (see claim 1) ; and computing a reconstruction loss based on the output image and a training image from the plurality of training images (see Liu, ¶ [34], “Based on this characteristic, a loss function between images and models in the same viewing angle can be minimized to optimize an overall 3D model generation process, and each module will be optimized correspondingly”) . As per claim 13, as addressed in claim 12, the combined Liu-Liu 2 does impliedly teach wherein training the 2D-to-3D transformer comprises: computing a perceptual loss based on the output image and the training image (Liu, ¶ [34], at best understood, since the claim language does not specify the scope of a perceptual loss ). Claim 17, which is similar in scope to claim 1 as addressed above with the addition of processor and memory that are also taught by Liu, Fig. 7, is thus rejected under the same rationale. Claim 19, which is similar in scope to claim 3 as addressed above, is thus rejected under the same rationale. Claim 20, which is similar in scope to claim 9 as addressed above, is thus rejected under the same rationale . 07-21-aia AIA Claim s 4, 6, 14, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Liuet al. (US. Patent App. Pub. No. 2024/0249469, “Liu”, hereinafter) in view of Liu 2 (“PETR: Position Embedding Transformation for Multi-view 3D Object Detection”) further in view of Seo et al. (“Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation”, retrieved from arXiv, https://arxiv.org/pdf/2303.07937v1, March 2023, “Seo”, hereinafter) . As per claim 4, the combined Liu-Liu 2 teach generating an output image based on the 3D model (as addressed in claim 1) , but does not expressly teach wherein the output image depicts the object from a perspective different from the plurality of input images . However, in the same field of endeavor as that of the combined Liu-Liu 2 (see Abstract on page 1), Seo teaches this feature, i.e., the output image depicts the object from a perspective different from the plurality of input images (see Fig. 2(b), page 2, output cat image at viewing angle different from the input cat image viewpoints). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to apply the method as taught by Seo to the combined Liu-Liu 2 method, the advantage of which is to ensure semantic consistency throughout all viewpoints of the scene (Abstract, page 1). As per claim 6, although not explicitly taught by the combined Liu-Liu 2, Seo does teach obtaining the plurality of input images comprises: obtaining an input prompt describing the object; and generating the plurality of input images based on the input prompt (Fig. 3, page 4, section 4.2, Semantic code sampling, and further Fig. 5, page 6, output images based on input prompt). As per claim 14, the combined Liu-Liu 2-Seo does teach wherein training the 2D-to-3D transformer comprises: generating pose information for a training image from the plurality of training images (Seo, Fig. 3, page 4, and section 4.3. Incorporating a coarse 3D prior, page 4-5, generating camera pose π ) ; and computing a pose loss based on the pose information (Seo, page 3, right column, starting with “Specifically, let us denote Ө as parameters of NeRF,…”). Thus, claim 14 would have been obvious over the combined references for the reason above. Claim 16, which is similar in scope to claim 6 as addressed above, is thus rejected under the same rationale . 07-21-aia AIA Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Liuet al. (US. Patent App. Pub. No. 2024/0249469, “Liu”, hereinafter) in view of Liu 2 (“PETR: Position Embedding Transformation for Multi-view 3D Object Detection”) further in view of Park et al. (US. Patent App. Pub. No. 2024/0087265, “Park”) . As per claim 15, the combined Liu-Liu 2 teaches wherein training the 2D-to-3D transformer comprises: generating a 3D model based on the 3D features (as addressed in claim 1) but does not explicitly teach computing a 3D model loss based on the 3D model and a ground-truth 3D model of an object. However, Park teaches a very similar method of generating 3D image from different viewpoint images (Fig. 2, ¶ [17]), wherein the method further includes computing a 3D model loss based on the 3D model and a ground-truth 3D model of an object (¶ [102]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method as taught by Park into the method as taught by the combined Liu-Liu 2, the advantage of which is to verify the result of the output images . 07-21-aia AIA 5. Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Liuet al. (US. Patent App. Pub. No. 2024/0249469, “Liu”, hereinafter) in view of Liu 2 (“PETR: Position Embedding Transformation for Multi-view 3D Object Detection”) further in view of Xu et al. (US. Patent App. Pub. No. 2023/0281926, “Xu”) . As per claim 18, the combined Liu-Liu 2 does not expressly teach the 2D-to-3D transformer comprises a cross-attention layer and a self-attention layer . However, Xu teaches a similar method of generating 3D images from input images as shown in Fig. 1 and 3, wherein the method further includes transformer comprises a cross-attention layer and a self-attention layer (see Fig. 6, and ¶ [7-8]). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the method as taught by Xu and apply to the combined Liu-Liu 2 method, the advantage of which is to obtain an efficient method for estimating the shape and pose of a body from a single image (¶ [5]). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Hau H. Nguyen whose telephone number is: 571-272-7787. The examiner can normally be reached on MON-FRI from 8:30-5:30. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard, can be reached on (571) 272-7773. The fax number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /HAU H NGUYEN/Primary Examiner, Art Unit 2611 Application/Control Number: 18/894,176 Page 2 Art Unit: 2611 Application/Control Number: 18/894,176 Page 3 Art Unit: 2611 Application/Control Number: 18/894,176 Page 4 Art Unit: 2611 Application/Control Number: 18/894,176 Page 5 Art Unit: 2611 Application/Control Number: 18/894,176 Page 6 Art Unit: 2611 Application/Control Number: 18/894,176 Page 7 Art Unit: 2611 Application/Control Number: 18/894,176 Page 8 Art Unit: 2611 Application/Control Number: 18/894,176 Page 9 Art Unit: 2611 Application/Control Number: 18/894,176 Page 10 Art Unit: 2611