DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on/after Mar. 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 4, 8-10, 12, and 16-17 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Choy et al. (“3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction”, published 2016, ‘CHOY’).
Regarding claim 1, CHOY discloses a computer-implemented method comprising:
obtaining … [2-D] image(s) (CHOY; p. 10; § 5.3; “We evaluated the performance of our network in single-view reconstruction using real-world images, comparing the performance with that of a recent method by Kar et al. … To make a quantitative comparison, we used images from the PASCAL VOC 2012 dataset … and its corresponding 3D models from the PASCAL 3D+ dataset …”);
determining … feature(s) of the … [2-D] image(s) by processing … portion(s) of the … [2-D] image(s) using a first machine learning technique (CHOY; p. 5; § 3; “The network is made up of three components: a 2D Convolutional Neural Network (2D-CNN), a novel architecture named 3D Convolutional LSTM (3D LSTM), and a 3D Deconvolutional Neural Network (3D-DCNN) (see Fig. 2). Given … image(s) of an object from arbitrary viewpoints, the 2D-CNN first encodes each input image x into low dimensional features T(x) (Section 3.1).” pp. 5-6; § 3.1; “We use CNNs to encode images into features. We designed two different 2D-CNN encoders as shown in Fig. 2: A standard feed-forward CNN and a deep residual variation of it … The encoder output is then flattened and passed to a fully connected layer which compresses the output into a 1024-dimensional feature vector.”);
generating multiple [3-D] visualizations associated with the … portion(s) of the … [2-D] image(s) by processing … portion(s) of the … feature(s) using … generative artificial intelligence technique(s) (CHOY; p. 5; § 3; “… given the encoded input, a set of newly proposed 3D Convolutional LSTM (3D-LSTM) units (Section 3.2) either selectively update their cell states or retain the states by closing the input gate. … the 3D-DCNN decodes the hidden states of the LSTM units and generates a 3D probabilistic voxel reconstruction (Section 3.3).” p. 7; § 3.3; “… we propose a simple decoder network with 5 convolutions and a deep residual version with 4 residual connections followed by a final convolution. After the last layer where the activation reaches the target output resolution, we convert the final activation … to the occupancy probability … of the voxel cell … using voxel-wise SoftMax.”);
selecting … the … [3-D] visualization(s) by processing … portion(s) of the multiple [3-D] visualizations using a second machine learning technique different from the first machine learning technique (CHOY; FIG. 1; p. 3; “Fig.1. (a) Some sample images of the objects we wish to reconstruct- notice that views are separated by a large baseline and objects’ appearance shows little texture and/or are non-Lambertian. (b) An overview of our proposed 3D-R2N2: The network takes a sequence of images (or just one image) from arbitrary (uncalibrated) viewpoints as input ([e.g.], 3 views of the armchair) and generates voxelized 3D reconstruction as an output. The reconstruction is incrementally refined as the network sees more views of the object.”); and
performing … automated action(s) based at least in part on the … selected … [3-D] visualization(s) (CHOY; p. 3; “[A] key attribute of the 3D-R2N2 is that … selectively updates hidden representations by controlling input/forget gates. In training, this mechanism allows the network to adaptively and consistently learn a suitable 3D representation of an object as (potentially conflicting) information from different viewpoints becomes available …” [The Examiner notes that the specification discloses ‘performing … automated action(s) includes automatically training … portion(s) of the … generative [AI] technique(s) using the … selected 3D visualization(s)’ related to step 308 of FIG. 3.]); wherein
the method is performed by … processing device(s) comprising a processor (CHOY; p. 8; § 4; “The networks used in the experiments were trained for 60,000 iterations with a batch size of 36 except for [Res3D-GRU 3] …, which needed a batch size of 24 to fit in an NVIDIA Titan X GPU.”) coupled to a memory ([The Examiner notes that it is commonly-known that a GPU typically has on-chip memory caches, particularly in the case of a dedicated GPU, which uses on board random-access memory.]).
Regarding claim 12, CHOY discloses a non-transitory processor-readable storage medium having stored therein program code of … software program(s) ([The Examiner asserts that the GPU disclosed by CHOY employs a graphics pipeline including a shader, which is a custom program which acts on vertices and primitives (polygons, etc.) and calculates colors, amongst other graphics operations. This program can be stored on on-board random-access memory, which is both ‘non-transitory’ as it is physical hardware and a ‘processor-readable storage medium’ as it can hold data in the form of persistent voltage(s).]), wherein the program code when executed by … processing device(s) (CHOY; p. 8; § 4; “The networks used in the experiments were trained for 60,000 iterations with a batch size of 36 except for [Res3D-GRU 3] …, which needed a batch size of 24 to fit in an NVIDIA Titan X GPU.”) causes the … processing device(s): … ([The remaining limitations are repeated nearly verbatim from those recited in independent claim 1.]).
Regarding claim 17, CHOY discloses an apparatus comprising:
… processing device(s) comprising a processor (CHOY; p. 8; § 4; “The networks used in the experiments were trained for 60,000 iterations with a batch size of 36 except for [Res3D-GRU 3] …, which needed a batch size of 24 to fit in an NVIDIA Titan X GPU.”) coupled to a memory ([The Examiner notes that it is commonly-known that a GPU typically has on-chip memory caches, particularly in the case of a dedicated GPU, which uses on board random-access memory.]);
the … processing device(s) being configured: … ([The remaining limitations are repeated verbatim from those recited in independent claim 12.]).
Regarding claim 4, CHOY discloses the computer-implemented method of claim 1, wherein determining … feature(s) of the … [2-D] image(s) comprises
PNG
media_image1.png
545
1397
media_image1.png
Greyscale
generating … voxel representation(s) of … portion(s) of the … [2-D] image(s) by processing the … portion(s) of the … [2-D] image(s) using … shape generator(s) (CHOY; FIG. 1; p. 3; “An overview of our proposed 3D-R2N2: The network takes a sequence of images (or just one image) from arbitrary (uncalibrated) viewpoints as input (in this example, 3 views of the armchair) and generates voxelized 3D reconstruction as an output. The reconstruction is incrementally refined as the network sees more views of the object.”), wherein
the … shape generator(s) comprise(s) … convolutional neural network-transformer model(s) (CHOY; FIG. 2; p. 5; “The network is made up of three components: a 2D Convolutional Neural Network (2D-CNN), a novel architecture named 3D Convolutional LSTM (3D LSTM), and a 3D Deconvolutional Neural Network (3D-DCNN) (see Fig. 2). Given … image(s) of an object from arbitrary viewpoints, the 2D-CNN first encodes each input image x into low dimensional features T (x) (Section 3.1). Then, given the encoded input, a set of newly proposed 3D Convolutional LSTM (3D-LSTM) units (Section 3.2) either selectively update their cell states or retain the states by closing the input gate. Finally, the 3D-DCNN decodes the hidden states of the LSTM units and generates a 3D probabilistic voxel reconstruction (Section 3.3).”).
Regarding claim 8 and claim 16, CHOY discloses the computer-implemented method of claim 1 and the non-transitory processor-readable storage medium of claim 12, wherein performing … automated action(s) comprises automatically training … portion(s) of the … generative artificial intelligence technique(s) using the … selected [3-D] visualization(s) (CHOY; p. 2; “Instead of trying to match a suitable 3D shape prior to the observation of the object and possibly adapt to it, we use deep convolutional neural networks to learn a mapping from observations to their underlying 3D shapes of objects from a large collection of training data.”).
Regarding claim 9, CHOY discloses the computer-implemented method of claim 1, wherein
the first machine learning technique comprises … deep learning algorithm(s) (CHOY; p. 2; “Instead of trying to match a suitable 3D shape prior to the observation of the object and possibly adapt to it, we use deep convolutional neural networks to learn a mapping from observations to their underlying 3D shapes of objects from a large collection of training data.”), and wherein
the second machine learning technique comprises … active learning technique(s) (CHOY; p. 3; “Our approach requires minimal supervision in training and testing (just bounding boxes, but no segmentation, key points, viewpoint labels, camera calibration, or class labels are needed).” [The Examiner regards the ‘bounding boxes’ as the labeling of data, constituting active learning.]).
Regarding claim 10, CHOY discloses the computer-implemented method of claim 9, wherein selecting … the … [3-D] visualization(s) comprises defining … informativeness measure(s) and … selection function(s) associated with the … active learning technique(s) (CHOY; p. 3; “Our approach requires minimal supervision in training and testing (just bounding boxes, but no segmentation, key points, viewpoint labels, camera calibration, or class labels are needed).” [The Examiner regards the ‘bounding boxes’ as a selection function in that a particular area of the pixelated 2-D image is selected to define a feature within the 2-D image, as well as informativeness measure(s) in that the ‘bounding boxes’ are analogous to data pertaining to the location of the feature within the 2-D image.]).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 USC 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 2-3, 13-14, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over CHOY as applied to claims 1, 12, and 17 above, respectively, and further in view of Das et al. (U.S. PG-PUB 2024/0020844, 'DAS').
Regarding claim 2, claim 13, and claim 18, CHOY discloses the computer-implemented method of claim 1, the non-transitory processor-readable storage medium of claim 12, and the apparatus of claim 17; however, CHOY does not explicitly disclose that determining … feature(s) of the … [2-D] images comprise(s) processing the … portion(s) of the … [2-D] image(s) using … semantic segmentation technique(s), which DAS discloses (DAS; ¶ 0052; “FIG. 2A illustrates a process of pre-training with the use of a transformer model. A segmentation network 203 includes a feature encoder 204 (also referred to as a feature extractor F) and a classifier/prediction head 208. The feature encoder 204 can generate a feature map 206. The feature map 206 is provided to the classifier/prediction head 208. The output of the classifier/prediction head 208 is an unsupervised predicted segmentation mask 209 from which an unsupervised loss can be generated.”).
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the computer-implemented method of claim 1, the non-transitory processor-readable storage medium of claim 12, and the apparatus of claim 17 of CHOY to include the processing the … portion(s) of the … [2-D] image(s) using … semantic segmentation technique(s) of DAS. The motivation for this modification is to exploit a transformer, which is a deep learning model that adopts the concept of self-attention, which involves differentially weighting the significance of each part of the input data. Transformers are designed to process sequential input data. However, the attention mechanism provides context for any input in the input sequence, which enables the ability to perform parallel processing of the input data. The transformer can be trained to learn a mapping between unsupervised and supervised outputs by conditioning on an intermediate feature map from the semantic segmentation model (DAS; ¶ [0006]).
Regarding claim 3, claim 14, and claim 19, CHOY-DAS discloses the computer-implemented method of claim 2, the non-transitory processor-readable storage medium of claim 13, and the apparatus of claim 18, wherein processing the … portion(s) of the … [2-D] image(s) using … semantic segmentation technique(s) comprises processing the … portion(s) of the … [2-D] image(s) using … deep neural network(s) in conjunction with … attention mechanism(s) (DAS; ¶ 0006; “… systems and techniques are described for providing a transformer machine learning model on top of an existing semantic segmentation model. A transformer is a deep learning model that adopts the concept of self-attention, which involves differentially weighting the significance of each part of the input data. Transformers are designed to process sequential input data. However, the attention mechanism provides context for any input in the input sequence, which enables the ability to perform parallel processing of the input data. The transformer can be trained to learn a mapping between unsupervised and supervised outputs by conditioning on an intermediate feature map [‘portion(s) of the … [2-D] image(s)’] from the semantic segmentation model.” ¶ 0043; FIG. 2A; ¶ 0052-53).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over CHOY as applied to claim 1 above, and further in view of Ganapathi et al. ("Graph Based Texture Pattern Classification", published 2022, 'GANAPATHI').
Regarding claim 5, CHOY discloses the computer-implemented method of claim 1; however, CHOY does not explicitly disclose that determining … feature(s) of the … [2-D] image(s) comprises
determining … texture(s) associated with … portion(s) of the … [2-D] image(s) by processing the … portion(s) of the … [2-D] image(s) using … appearance renderer(s), wherein the … appearance renderer(s) comprise(s) …
graph neural network(s) and … deep reinforcement learning technique(s), which GANAPATHI discloses (GANAPATHI; p. 2; “… there is a huge demand for novel algorithms to generate efficient features to identify texture patterns. … we present a deep graph neural network-based technique for 3D meshes. Though classical graph-based algorithms are known for texture classification, this is the first time a 3D mesh texture classification problem has been addressed using a graph neural network [‘appearance renderer(s)’]. The 3D meshes are converted to a graph structure, and each node in the graph is classified. In graph construction, we assume that each facet has three neighboring facets. Each face in a 3D mesh serves as a node, with edge connections created with adjacent facets. At each node, we compute a geometrical feature called local depth, and this feature is used for message passing and aggregation between nodes in the graph network. As a result, the features computed at each node are critical; therefore, the computed geometric surface features should be more efficient than curvature and curvedness features.” FIG. 2; p. 3; “We adopt deep learning-based algorithms [‘appearance renderer(s)’] for 3D mesh texture classification since they are successful in 2D computer vision applications such as recognition, classification, and segmentation. … We approached the texture problem as a graph due to its versatile data structure. The graphs are fed into a graph neural network for node classification to predict the nodes’ class in a graph with only partial node labeling. The following sections illustrate how to create a graph from a [3-D] mesh and use geometrical features to describe each node in the graph for message passing and aggregation during network training.”).
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the computer-implemented method of claim 1 of CHOY to include the determining texture(s) associated with portion(s) of the [2-D] image(s) by processing the portion(s) of the [2-D] image(s) using appearance renderer(s), wherein the appearance renderer(s) comprise(s) graph neural network(s) and deep reinforcement learning technique(s) of GANAPATHI. The motivation for this modification is to implement a graph learning-based approach for classifying the texture of each facet in a 3D mesh. First, a 3-D mesh is transformed into a graph structure in which every node is a facet of a given mesh. Further, each facet is described by a feature vector computed utilizing the neighboring facets within a radius and their geometric properties. The graph structure is then fed into a graph neural network, classifying each node as a texture or non-textured class (GANAPATHI; Abstract).
Claims 6-7, 11, 15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over CHOY as applied to claims 1, 12, and 17 above, respectively, and further in view of Mauldin et al. (U.S. PG-PUB 2024/0312075, 'MAULDIN').
Regarding claim 6, claim 15, and claim 20, CHOY discloses the computer-implemented method of claim 1, the non-transitory processor-readable storage medium of claim 12, and the apparatus of claim 17; however, CHOY does not explicitly disclose that generating multiple [3-D] visualizations associated with the … portion(s) of the … [2-D] image(s) comprises processing the … portion(s) of the … feature(s) using … generative adversarial network(s), which MAULDIN discloses (MAULDIN; ¶ 0036; “A second step 304 of … FIG. 3 involves generating a synthetic image using a trained machine learning model, the synthetic image, in aspects, having the appearance of being acquired by the imaging device of the source image 202 and, in aspects, having … anatomical feature(s) that can be inferred from the source image 202, such as by comparison between the source image 202 and at least one a priori anatomical model. … The trained machine learning model may be comprised of … generative adversarial network(s) (GANs), or any other machine learning model capable of generating … 3D synthetic medical images. The machine learning model architecture may … long-term short memory (LTSM), … as well as any other model architecture element advantageous for the purpose of generating synthetic anatomical features from input images.”).
Before the effective filing date of the claimed invention, it would have been obvious to a person having ordinary skill in the art to modify the computer-implemented method of claim 1, the non-transitory processor-readable storage medium of claim 12, and the apparatus of claim 17 of CHOY to include the processing the … portion(s) of the … feature(s) using … generative adversarial network(s) of MAULDIN. The motivation for this modification is to use machine learning techniques to make inferences of 2-D imagery to generate 3-D voxelized structures based on features detected within the 2-D imagery.
Regarding claim 7, CHOY-MAULDIN disclose the computer-implemented method of claim 6, wherein processing the … portion(s) of the … feature(s) using … generative adversarial network(s) comprises implementing … adversarial loss function(s) in connection with the … generative adversarial network(s) (MAULDIN; FIG. 4; ¶ 0039; “… the machine learning model is trained through a generative adversarial network (GAN) architecture 400 [which] generates a synthetic image 414 from the input image 402 via a generator 412. A discriminator 416 compares the synthetic image 414 and a ground truth image 410 to determine an adversarial loss function 420. Both the generator 412 and discriminator 416 are trained in conjoint with content loss 420 and adversarial loss 420 functions as feedback 422, 424 to improve performance of both the discriminator 416 and generator 412 simultaneously.”).
Regarding claim 11, CHOY discloses the computer-implemented method of claim 1; however, CHOY does not explicitly disclose that performing … automated action(s) comprises automatically outputting, to … user device(s) associated with the … [2-D] image(s), the … selected [3-D] visualization(s), which MAULDIN discloses (MAULDIN; FIG. 5; ¶ 0040; “The augmented image is transferred to the output unit 512, which executes the methods involved with outputting the augmented image 308. … the output unit 512 transfers the augmented image to a display apparatus 514 that conveys the augmented image to a user and/or an observer.”).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JONATHAN M COFINO whose telephone number is (303) 297-4268. The examiner can normally be reached Monday-Friday 10A-4P MT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at 571-272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JONATHAN M COFINO/ Examiner, Art Unit 2614
/KENT W CHANG/ Supervisory Patent Examiner, Art Unit 2614