DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
2. Receipt is acknowledged of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file.
Information Disclosure Statement
3. The information disclosure statement (IDS) submitted on 01/06/2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
7. Claim(s) 1-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nagano et al. (US 2022/0005212A1) in view of Goforth et al. (US 2021/0150228 A1).
8. With reference to claim 1, Nagano teaches An information processing apparatus comprising at least one processor and a memory which is configured to store instructions, the at least one processor executing: an obtaining process of obtaining input data; (“The input unit 10 accepts learning data including an input point cloud which is a set of sparse three-dimensional points representing a part of an object,” [0050] “A system for performing completion of a point cloud that is a measurement result obtained by measuring a set of three-dimensional points on an object, the system comprises: a processor; and a memory storing computer-executable program instructions that when executed by the processor cause the system to: input, by a class identifier, an input point cloud to an identifier that is learned in advance and identifies a class of the object and outputting a class identification feature that is calculated at the time of identification of a class of an object represented by the input point cloud;” claim 6) Nagano also teaches a three-dimensional structure data generating process of generating three-dimensional structure data from the input data; (“The input unit 10 accepts learning data including an input point cloud which is a set of sparse three-dimensional points representing a part of an object,” [0050] “The determination result is a result, having a value on the interval [0,1], of determination by the determiner about a set (x,y) of three-dimensional points obtained by combining together an input point cloud and the teacher point cloud and a result, having a value on the interval [0,1], of determination by the determiner about a set (x,y′) of three-dimensional points obtained by combining together the input point cloud and the generated shape completion point cloud. As for the determiner, learning is performed such that the probability of correctly determining truth or falsity for the set (x,y) of three-dimensional points and the set (x,y′) of three-dimensional points is high (i.e., such that logD(x,y) is maximized)” [0059] “The class identifier then generates a shape completion point cloud which is a set of three-dimensional points serving to complete the point cloud by convoluting the integration result.” [0067]) Nagano further teaches a sampling process of generating sampled three-dimensional structure data by sampling the three-dimensional structure data; a completing process with respect to the sampled three-dimensional structure data with use of a completion model; (“The learning unit 32 optimizes a parameter for a generator and a parameter for a determiner on the basis of the learning data including the input point cloud and the teacher point cloud such that a loss function indicated by Formula (3) above is optimized and stores the parameters in the parameter storage unit 40. The learning unit 32 inputs the input point cloud to the class identifier learned by the pre-learning unit 30, and a class identification feature is output to the generator and to the determiner. The generator is the above-described generator. The generator receives, as input, a point cloud and a class identification feature, gains an integration result obtained by integrating a global feature serving as a global feature based on local features extracted from points of the point cloud with the class identification feature, and convoluting the integration result, thereby generating a shape completion point cloud which is to complete the point cloud. The determiner is the above-described determiner. The determiner receives, as input, a point cloud, a shape completion point cloud, and a class identification feature and determines whether a set of three-dimensional points obtained by combining together the point cloud and the shape completion point cloud is a true point cloud.” [0055-0058]) Nagano teaches a first training process of training the completion model with reference to a first loss value (“The learning unit 32 optimizes a parameter for a generator and a parameter for a determiner on the basis of the learning data including the input point cloud and the teacher point cloud such that a loss function indicated by Formula (3) above is optimized and stores the parameters in the parameter storage unit 40.” [0055] “The parameter optimization is performed on the basis of the loss function indicated by Formula (3) above using stochastic gradient descent. The loss function is represented using a determination result from the determiner and a distance between a shape completion point cloud y′ generated by the generator and a teacher point cloud y. The determination result is a result, having a value on the interval [0,1], of determination by the determiner about a set (x,y) of three-dimensional points obtained by combining together an input point cloud and the teacher point cloud and a result, having a value on the interval [0,1], of determination by the determiner about a set (x,y′) of three-dimensional points obtained by combining together the input point cloud and the generated shape completion point cloud.” [0059])
PNG
media_image1.png
657
448
media_image1.png
Greyscale
Nagano does not explicitly teach a shape estimating process in which an intermediate feature value in the completing process is referred to; and a loss value pertaining to a shape that has been obtained by the shape estimating process. This is what Goforth teaches (“The shared encoder 301 may include a neural network model (i.e., artificial neural network architectures such as e.g., feed-forward neural networks, recurrent neural networks, convolutional neural networks, or the like) that is trained or configured to receive sensor data (e.g., a LIDAR point cloud) corresponding to an object as an input, and generate an output that comprises an encoded or alternative representation of the input 304 (a “code”). Optionally, the code may be a lower dimensional representation of the input point cloud data, and that include defined values of latent variables that each represent a feature of the point cloud (in particular, shape features and/or pose features). The code may include states or feature maps in a vector form or a tensor form corresponding to the received input. The code 304 may serve as a context or conditioning input for the shape decoder 302 and/or the pose decoder 303 for generating outputs including an estimated shape and an estimated pose, respectively, corresponding to the input sensor data.” [0040] “Training may be performed by selecting a batch of training data and, for each partial observation in the training data, inputting the partial observation to the shared encoder and shape decoder to process the input observation with current parameter or weight values of the shared encoder and shape decoder. Training may further include updating the current parameters of the shared encoder and shape decoder based on an analysis of the output completed observation/shape with respect to the ground truth data. In certain embodiments, the training may be constrained by a loss function. Specifically, the shared encoder and shape decoder may be trained to minimize or optimize a loss function between estimated point completion (i.e., shape estimation based on partial observations) and ground truth point completions (i.e., shape estimation based on complete observations). Examples of loss function may include, without limitation, Chamfer Distance, Earth Mover Distance, other distance metric functions, or the like, chosen based on the application and/or required correspondence with point cloud data points. Chamfer distance is a method for measuring total distance between two sets of 3D points. … the encoder may learn to abstract each unaligned partial input into a fixed-length code which captures the object shape in such a way as that the complete shape can be recovered by the shape decoder conditioned on the code in the same (unknown) pose as the partial input.” [0051-0054]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Goforth into Nagano, in order to improve shape completion accuracy.
9. With reference to claim 2, Nagano does not explicitly teach in the shape estimating process, the at least one processor computes one or more distance function values; and the first loss value is a loss value pertaining to the one or more distance function values. This is what Goforth teaches (“The shared encoder 301 may include a neural network model (i.e., artificial neural network architectures such as e.g., feed-forward neural networks, recurrent neural networks, convolutional neural networks, or the like) that is trained or configured to receive sensor data (e.g., a LIDAR point cloud) corresponding to an object as an input, and generate an output that comprises an encoded or alternative representation of the input 304 (a “code”). Optionally, the code may be a lower dimensional representation of the input point cloud data, and that include defined values of latent variables that each represent a feature of the point cloud (in particular, shape features and/or pose features). The code may include states or feature maps in a vector form or a tensor form corresponding to the received input. The code 304 may serve as a context or conditioning input for the shape decoder 302 and/or the pose decoder 303 for generating outputs including an estimated shape and an estimated pose, respectively, corresponding to the input sensor data.” [0040] “Training may be performed by selecting a batch of training data and, for each partial observation in the training data, inputting the partial observation to the shared encoder and shape decoder to process the input observation with current parameter or weight values of the shared encoder and shape decoder. Training may further include updating the current parameters of the shared encoder and shape decoder based on an analysis of the output completed observation/shape with respect to the ground truth data. In certain embodiments, the training may be constrained by a loss function. Specifically, the shared encoder and shape decoder may be trained to minimize or optimize a loss function between estimated point completion (i.e., shape estimation based on partial observations) and ground truth point completions (i.e., shape estimation based on complete observations). Examples of loss function may include, without limitation, Chamfer Distance, Earth Mover Distance, other distance metric functions, or the like, chosen based on the application and/or required correspondence with point cloud data points. Chamfer distance is a method for measuring total distance between two sets of 3D points. … the encoder may learn to abstract each unaligned partial input into a fixed-length code which captures the object shape in such a way as that the complete shape can be recovered by the shape decoder conditioned on the code in the same (unknown) pose as the partial input.” [0051-0054] “Processor 705 is a central processing device of the system, configured to perform calculations and logic operations required to execute programming instructions. As used in this document and in the claims, the terms “processor” and “processing device” may refer to a single processor or any number of processors in a set of processors that collectively perform a set of operations, such as a central processing unit (CPU), a graphics processing unit (GPU), a remote server, or a combination of these. Read only memory (ROM), random access memory (RAM), flash memory, hard drives and other devices capable of storing electronic data constitute examples of memory devices 725.” [0069]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Goforth into Nagano, in order to improve shape completion accuracy.
10. With reference to claim 3, Nagano teaches the sampled three-dimensional structure data. (“The learning unit 32 optimizes a parameter for a generator and a parameter for a determiner on the basis of the learning data including the input point cloud and the teacher point cloud such that a loss function indicated by Formula (3) above is optimized and stores the parameters in the parameter storage unit 40. The learning unit 32 inputs the input point cloud to the class identifier learned by the pre-learning unit 30, and a class identification feature is output to the generator and to the determiner. The generator is the above-described generator. The generator receives, as input, a point cloud and a class identification feature, gains an integration result obtained by integrating a global feature serving as a global feature based on local features extracted from points of the point cloud with the class identification feature, and convoluting the integration result, thereby generating a shape completion point cloud which is to complete the point cloud. The determiner is the above-described determiner. The determiner receives, as input, a point cloud, a shape completion point cloud, and a class identification feature and determines whether a set of three-dimensional points obtained by combining together the point cloud and the shape completion point cloud is a true point cloud.” [0055-0058])
Nagano does not explicitly teach the at least one processor further executes a feature value converting process of converting the intermediate feature value in the completing process into a latent variable, and executes the shape estimating process with reference to the latent variable. This is what Goforth teaches (“The shared encoder 301 may include a neural network model (i.e., artificial neural network architectures such as e.g., feed-forward neural networks, recurrent neural networks, convolutional neural networks, or the like) that is trained or configured to receive sensor data (e.g., a LIDAR point cloud) corresponding to an object as an input, and generate an output that comprises an encoded or alternative representation of the input 304 (a “code”). Optionally, the code may be a lower dimensional representation of the input point cloud data, and that include defined values of latent variables that each represent a feature of the point cloud (in particular, shape features and/or pose features). The code may include states or feature maps in a vector form or a tensor form corresponding to the received input. The code 304 may serve as a context or conditioning input for the shape decoder 302 and/or the pose decoder 303 for generating outputs including an estimated shape and an estimated pose, respectively, corresponding to the input sensor data. Optionally, the shape decoder 302 and/or the pose decoder 303 may be neural network models (i.e., artificial neural network architectures such as e.g., feed-forward neural network, recurrent neural network, convolutional neural network, or the like). In certain scenarios, the shared encoder 301, the shape decode 302 and/or the pose decoder 303 may be embodied as a multi-layer perceptron (MLP) comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more hidden layers, and may utilize any suitable learning algorithm described herein or otherwise known in the art. In a MLP, each node is a feed-forward node, with a number of inputs, a number of weights, a summation point, a non-linear function, and an output port.” [0040] “Training may be performed by selecting a batch of training data and, for each partial observation in the training data, inputting the partial observation to the shared encoder and shape decoder to process the input observation with current parameter or weight values of the shared encoder and shape decoder. Training may further include updating the current parameters of the shared encoder and shape decoder based on an analysis of the output completed observation/shape with respect to the ground truth data. In certain embodiments, the training may be constrained by a loss function. Specifically, the shared encoder and shape decoder may be trained to minimize or optimize a loss function between estimated point completion (i.e., shape estimation based on partial observations) and ground truth point completions (i.e., shape estimation based on complete observations). Examples of loss function may include, without limitation, Chamfer Distance, Earth Mover Distance, other distance metric functions, or the like, chosen based on the application and/or required correspondence with point cloud data points. Chamfer distance is a method for measuring total distance between two sets of 3D points. … the encoder may learn to abstract each unaligned partial input into a fixed-length code which captures the object shape in such a way as that the complete shape can be recovered by the shape decoder conditioned on the code in the same (unknown) pose as the partial input.” [0051-0054] “Processor 705 is a central processing device of the system, configured to perform calculations and logic operations required to execute programming instructions. As used in this document and in the claims, the terms “processor” and “processing device” may refer to a single processor or any number of processors in a set of processors that collectively perform a set of operations, such as a central processing unit (CPU), a graphics processing unit (GPU), a remote server, or a combination of these. Read only memory (ROM), random access memory (RAM), flash memory, hard drives and other devices capable of storing electronic data constitute examples of memory devices 725.” [0069]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Goforth into Nagano, in order to improve shape completion accuracy.
11. With reference to claim 4, Nagano teaches the completion model into which the sampled three-dimensional structure data is inputted and which outputs completed three-dimensional structure data; (“The input unit 10 accepts learning data including an input point cloud which is a set of sparse three-dimensional points representing a part of an object,” [0050] “The learning unit 32 optimizes a parameter for a generator and a parameter for a determiner on the basis of the learning data including the input point cloud and the teacher point cloud such that a loss function indicated by Formula (3) above is optimized and stores the parameters in the parameter storage unit 40. The learning unit 32 inputs the input point cloud to the class identifier learned by the pre-learning unit 30, and a class identification feature is output to the generator and to the determiner. The generator is the above-described generator. The generator receives, as input, a point cloud and a class identification feature, gains an integration result obtained by integrating a global feature serving as a global feature based on local features extracted from points of the point cloud with the class identification feature, and convoluting the integration result, thereby generating a shape completion point cloud which is to complete the point cloud. The determiner is the above-described determiner. The determiner receives, as input, a point cloud, a shape completion point cloud, and a class identification feature and determines whether a set of three-dimensional points obtained by combining together the point cloud and the shape completion point cloud is a true point cloud.” [0055-0058])
Nagano does not explicitly teach an encoder which outputs the intermediate feature value, and a decoder into which the intermediate feature value is inputted; and in the first training process, the at least one processor subjects the encoder to machine learning with reference to the first loss value. This is what Goforth teaches (“The shared encoder 301 may include a neural network model (i.e., artificial neural network architectures such as e.g., feed-forward neural networks, recurrent neural networks, convolutional neural networks, or the like) that is trained or configured to receive sensor data (e.g., a LIDAR point cloud) corresponding to an object as an input, and generate an output that comprises an encoded or alternative representation of the input 304 (a “code”). Optionally, the code may be a lower dimensional representation of the input point cloud data, and that include defined values of latent variables that each represent a feature of the point cloud (in particular, shape features and/or pose features). The code may include states or feature maps in a vector form or a tensor form corresponding to the received input. The code 304 may serve as a context or conditioning input for the shape decoder 302 and/or the pose decoder 303 for generating outputs including an estimated shape and an estimated pose, respectively, corresponding to the input sensor data. Optionally, the shape decoder 302 and/or the pose decoder 303 may be neural network models (i.e., artificial neural network architectures such as e.g., feed-forward neural network, recurrent neural network, convolutional neural network, or the like). In certain scenarios, the shared encoder 301, the shape decode 302 and/or the pose decoder 303 may be embodied as a multi-layer perceptron (MLP) comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more hidden layers, and may utilize any suitable learning algorithm described herein or otherwise known in the art. In a MLP, each node is a feed-forward node, with a number of inputs, a number of weights, a summation point, a non-linear function, and an output port.” [0040] “Training may be performed by selecting a batch of training data and, for each partial observation in the training data, inputting the partial observation to the shared encoder and shape decoder to process the input observation with current parameter or weight values of the shared encoder and shape decoder. Training may further include updating the current parameters of the shared encoder and shape decoder based on an analysis of the output completed observation/shape with respect to the ground truth data. In certain embodiments, the training may be constrained by a loss function. Specifically, the shared encoder and shape decoder may be trained to minimize or optimize a loss function between estimated point completion (i.e., shape estimation based on partial observations) and ground truth point completions (i.e., shape estimation based on complete observations). Examples of loss function may include, without limitation, Chamfer Distance, Earth Mover Distance, other distance metric functions, or the like, chosen based on the application and/or required correspondence with point cloud data points. Chamfer distance is a method for measuring total distance between two sets of 3D points. … the encoder may learn to abstract each unaligned partial input into a fixed-length code which captures the object shape in such a way as that the complete shape can be recovered by the shape decoder conditioned on the code in the same (unknown) pose as the partial input.” [0051-0054] “Processor 705 is a central processing device of the system, configured to perform calculations and logic operations required to execute programming instructions. As used in this document and in the claims, the terms “processor” and “processing device” may refer to a single processor or any number of processors in a set of processors that collectively perform a set of operations, such as a central processing unit (CPU), a graphics processing unit (GPU), a remote server, or a combination of these. Read only memory (ROM), random access memory (RAM), flash memory, hard drives and other devices capable of storing electronic data constitute examples of memory devices 725.” [0069]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Goforth into Nagano, in order to improve shape completion accuracy.
12. With reference to claim 5, Nagano teaches the three-dimensional structure data which has been generated by the three-dimensional structure data generating process. (“The learning unit 32 optimizes a parameter for a generator and a parameter for a determiner on the basis of the learning data including the input point cloud and the teacher point cloud such that a loss function indicated by Formula (3) above is optimized and stores the parameters in the parameter storage unit 40. The learning unit 32 inputs the input point cloud to the class identifier learned by the pre-learning unit 30, and a class identification feature is output to the generator and to the determiner. The generator is the above-described generator. The generator receives, as input, a point cloud and a class identification feature, gains an integration result obtained by integrating a global feature serving as a global feature based on local features extracted from points of the point cloud with the class identification feature, and convoluting the integration result, thereby generating a shape completion point cloud which is to complete the point cloud. The determiner is the above-described determiner. The determiner receives, as input, a point cloud, a shape completion point cloud, and a class identification feature and determines whether a set of three-dimensional points obtained by combining together the point cloud and the shape completion point cloud is a true point cloud.” [0055-0058])
Nagano does not explicitly teach the at least one processor executes a second training process of subjecting the encoder and the decoder to machine learning with reference to a second loss value that indicates a difference between the completed three-dimensional structure data which has been outputted by the decoder and the three-dimensional structure data which has been generated. This is what Goforth teaches (“The shared encoder 301 may include a neural network model (i.e., artificial neural network architectures such as e.g., feed-forward neural networks, recurrent neural networks, convolutional neural networks, or the like) that is trained or configured to receive sensor data (e.g., a LIDAR point cloud) corresponding to an object as an input, and generate an output that comprises an encoded or alternative representation of the input 304 (a “code”). Optionally, the code may be a lower dimensional representation of the input point cloud data, and that include defined values of latent variables that each represent a feature of the point cloud (in particular, shape features and/or pose features). The code may include states or feature maps in a vector form or a tensor form corresponding to the received input. The code 304 may serve as a context or conditioning input for the shape decoder 302 and/or the pose decoder 303 for generating outputs including an estimated shape and an estimated pose, respectively, corresponding to the input sensor data. Optionally, the shape decoder 302 and/or the pose decoder 303 may be neural network models (i.e., artificial neural network architectures such as e.g., feed-forward neural network, recurrent neural network, convolutional neural network, or the like). In certain scenarios, the shared encoder 301, the shape decode 302 and/or the pose decoder 303 may be embodied as a multi-layer perceptron (MLP) comprising at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more hidden layers, and may utilize any suitable learning algorithm described herein or otherwise known in the art. In a MLP, each node is a feed-forward node, with a number of inputs, a number of weights, a summation point, a non-linear function, and an output port. … he first layer may use m input points represented as an m×3 matrix P where each row is the 3D coordinate of a point p.sub.i=(x, y, z). A shared multilayer perceptron (MLP) consisting of two linear layers with ReLU activation may then be used to transform each p.sub.i into a point feature vector f.sub.i to generate a feature matrix F whose rows are the learned point features f.sub.i. Then, a point-wise max pooling may be performed on F to obtain a k-dimensional global feature g, where g.sub.j=max.sub.i=1, . . . , m{F.sub.ij} for j=1, . . . , k. The second deep network layer of the shared encoder 301 may use F and g as input and concatenate g to each f.sub.i to obtain an augmented point feature matrix F.sub.1 whose rows are the concatenated feature vectors [f.sub.i g]. F.sub.1 may be passed through another shared MLP and point-wise max pooling similar to the ones in the first layer, which gives the final feature vector v. … the shape decoder 302 may generate an output point cloud corresponding to an estimated shape of an object from the feature vector v. The shape decoder 302 may generate the output point cloud in two stages. In the first stage, a coarse output Y.sub.coarse of s points may be generated by passing v through a fully-connected network with 3s output units and reshaping the output into an s×3 matrix. In the second stage, for each point q.sub.i in Y.sub.coarse, a patch of t=u.sup.2 points may be generated in local coordinates centered at q.sub.i via a folding operation, and transformed into global coordinates by adding q.sub.i to the output.” [0040-0043] “Training may be performed by selecting a batch of training data and, for each partial observation in the training data, inputting the partial observation to the shared encoder and shape decoder to process the input observation with current parameter or weight values of the shared encoder and shape decoder. Training may further include updating the current parameters of the shared encoder and shape decoder based on an analysis of the output completed observation/shape with respect to the ground truth data. In certain embodiments, the training may be constrained by a loss function. Specifically, the shared encoder and shape decoder may be trained to minimize or optimize a loss function between estimated point completion (i.e., shape estimation based on partial observations) and ground truth point completions (i.e., shape estimation based on complete observations). Examples of loss function may include, without limitation, Chamfer Distance, Earth Mover Distance, other distance metric functions, or the like, chosen based on the application and/or required correspondence with point cloud data points. Chamfer distance is a method for measuring total distance between two sets of 3D points. … the encoder may learn to abstract each unaligned partial input into a fixed-length code which captures the object shape in such a way as that the complete shape can be recovered by the shape decoder conditioned on the code in the same (unknown) pose as the partial input.” [0051-0054] “Processor 705 is a central processing device of the system, configured to perform calculations and logic operations required to execute programming instructions. As used in this document and in the claims, the terms “processor” and “processing device” may refer to a single processor or any number of processors in a set of processors that collectively perform a set of operations, such as a central processing unit (CPU), a graphics processing unit (GPU), a remote server, or a combination of these. Read only memory (ROM), random access memory (RAM), flash memory, hard drives and other devices capable of storing electronic data constitute examples of memory devices 725.” [0069]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Goforth into Nagano, in order to improve shape completion accuracy.
13. Claim 6 is similar in scope to claim 1, and thus is rejected under similar rationale. Nagano additionally teaches a completing process with respect to the three-dimensional structure data with use of a completion model; and an output data generating process of generating output data from the three-dimensional structure data to which the completing process has been applied, the completion model being a model which has been trained (“The learning unit 32 optimizes a parameter for a generator and a parameter for a determiner on the basis of the learning data including the input point cloud and the teacher point cloud such that a loss function indicated by Formula (3) above is optimized and stores the parameters in the parameter storage unit 40. The learning unit 32 inputs the input point cloud to the class identifier learned by the pre-learning unit 30, and a class identification feature is output to the generator and to the determiner. The generator is the above-described generator. The generator receives, as input, a point cloud and a class identification feature, gains an integration result obtained by integrating a global feature serving as a global feature based on local features extracted from points of the point cloud with the class identification feature, and convoluting the integration result, thereby generating a shape completion point cloud which is to complete the point cloud. The determiner is the above-described determiner. The determiner receives, as input, a point cloud, a shape completion point cloud, and a class identification feature and determines whether a set of three-dimensional points obtained by combining together the point cloud and the shape completion point cloud is a true point cloud. The parameter optimization is performed on the basis of the loss function indicated by Formula (3) above using stochastic gradient descent. The loss function is represented using a determination result from the determiner and a distance between a shape completion point cloud y′ generated by the generator and a teacher point cloud y. The determination result is a result, having a value on the interval [0,1], of determination by the determiner about a set (x,y) of three-dimensional points obtained by combining together an input point cloud and the teacher point cloud and a result, having a value on the interval [0,1], of determination by the determiner about a set (x,y′) of three-dimensional points obtained by combining together the input point cloud and the generated shape completion point cloud.” [0055-0059])
14. With reference to claim 7, Nagano does not explicitly teach the at least one processor executes a displaying process of displaying the output data which has been generated by the output data generating process. This is what Goforth teaches (“Training may be performed by selecting a batch of training data and, for each partial observation in the training data, inputting the partial observation to the shared encoder and shape decoder to process the input observation with current parameter or weight values of the shared encoder and shape decoder. Training may further include updating the current parameters of the shared encoder and shape decoder based on an analysis of the output completed observation/shape with respect to the ground truth data. In certain embodiments, the training may be constrained by a loss function. Specifically, the shared encoder and shape decoder may be trained to minimize or optimize a loss function between estimated point completion (i.e., shape estimation based on partial observations) and ground truth point completions (i.e., shape estimation based on complete observations). Examples of loss function may include, without limitation, Chamfer Distance, Earth Mover Distance, other distance metric functions, or the like, chosen based on the application and/or required correspondence with point cloud data points. Chamfer distance is a method for measuring total distance between two sets of 3D points.” [0051-0052] “Processor 705 is a central processing device of the system, configured to perform calculations and logic operations required to execute programming instructions. As used in this document and in the claims, the terms “processor” and “processing device” may refer to a single processor or any number of processors in a set of processors that collectively perform a set of operations, such as a central processing unit (CPU), a graphics processing unit (GPU), a remote server, or a combination of these. Read only memory (ROM), random access memory (RAM), flash memory, hard drives and other devices capable of storing electronic data constitute examples of memory devices 725. … An optional display interface 730 may permit information from the bus 700 to be displayed on a display device 735 in visual, graphic or alphanumeric format.” [0069-0070]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Goforth into Nagano, in order to improve shape completion accuracy.
15. With reference to claim 8, Nagano does not explicitly teach in the obtaining process, the at least one processor obtains a medical image as the input data; and in the displaying process, the at least one processor displays the output data for assisting a healthcare professional in making a decision. This is what Goforth teaches (“The first layer may use m input points represented as an m×3 matrix P where each row is the 3D coordinate of a point p.sub.i=(x, y, z). A shared multilayer perceptron (MLP) consisting of two linear layers with ReLU activation may then be used to transform each p.sub.i into a point feature vector f.sub.i to generate a feature matrix F whose rows are the learned point features f.sub.i. Then, a point-wise max pooling may be performed on F to obtain a k-dimensional global feature g, where g.sub.j=max.sub.i=1, . . . , m{F.sub.ij} for j=1, . . . , k. The second deep network layer of the shared encoder 301 may use F and g as input and concatenate g to each f.sub.i to obtain an augmented point feature matrix F.sub.1 whose rows are the concatenated feature vectors [f.sub.i g]. F.sub.1 may be passed through another shared MLP and point-wise max pooling similar to the ones in the first layer, which gives the final feature vector v.” [0042] “Training may be performed by selecting a batch of training data and, for each partial observation in the training data, inputting the partial observation to the shared encoder and shape decoder to process the input observation with current parameter or weight values of the shared encoder and shape decoder. Training may further include updating the current parameters of the shared encoder and shape decoder based on an analysis of the output completed observation/shape with respect to the ground truth data. In certain embodiments, the training may be constrained by a loss function. Specifically, the shared encoder and shape decoder may be trained to minimize or optimize a loss function between estimated point completion (i.e., shape estimation based on partial observations) and ground truth point completions (i.e., shape estimation based on complete observations). Examples of loss function may include, without limitation, Chamfer Distance, Earth Mover Distance, other distance metric functions, or the like, chosen based on the application and/or required correspondence with point cloud data points. Chamfer distance is a method for measuring total distance between two sets of 3D points.” [0051-0052] “Processor 705 is a central processing device of the system, configured to perform calculations and logic operations required to execute programming instructions. As used in this document and in the claims, the terms “processor” and “processing device” may refer to a single processor or any number of processors in a set of processors that collectively perform a set of operations, such as a central processing unit (CPU), a graphics processing unit (GPU), a remote server, or a combination of these. Read only memory (ROM), random access memory (RAM), flash memory, hard drives and other devices capable of storing electronic data constitute examples of memory devices 725. … An optional display interface 730 may permit information from the bus 700 to be displayed on a display device 735 in visual, graphic or alphanumeric format.” [0069-0070] “An “automated device” or “robotic device” refers to an electronic device that includes a processor, programming instructions, and one or more components that based on commands from the processor can perform at least some operations or tasks with minimal or no human intervention. For example, an automated device may perform one or more automatic functions or function sets. Examples of such operations, functions or tasks may include without, limitation, navigation, transportation, driving, delivering, loading, unloading, medical-related processes, construction-related processes, and/or the like. Example automated devices may include, without limitation, autonomous vehicles, drones and other autonomous robotic devices.” [0075]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Goforth into Nagano, in order to improve shape completion accuracy.
16. Claim 9 is similar in scope to claim 1, and thus is rejected under similar rationale.
17. Claim 10 is similar in scope to claim 1, and thus is rejected under similar rationale. Nagano does not explicitly teach A non-transitory recording medium in which a program for causing a computer to function as the information processing apparatus recited in claim 1 is stored, the program causing the computer to execute. This is what Goforth teaches (“The system may include a processor and a non-transitory computer readable medium for storing data and program code collectively defining a neural network which has been trained to jointly estimate a pose and a shape of a plurality of objects from incomplete point cloud data. … The non-transitory computer readable medium may also include programming instructions that when executed cause the processor to execute the methods for jointly estimating a pose and a shape of an object.” [0006])
18. Claim 11 is similar in scope to claim 10, and thus is rejected under similar rationale.
Conclusion
19. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michelle Chin whose telephone number is (571)270-3697. The examiner can normally be reached on Monday-Friday 8:00 AM-4:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http:/Awww.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Kent Chang can be reached on (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is (571)273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https:/Awww.uspto.gov/patents/apply/patent- center for more information about Patent Center and https:/Awww.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHELLE CHIN/
Primary Examiner, Art Unit 2614