DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites the limitation "the feature value of the associated first three-dimensional point" in line 16. There is insufficient antecedent basis for this limitation in the claim. The claim does previously recite “a feature value of the two-dimensional image”, and thus “the feature value” would appear to have antecedent basis, however, there is no previous “feature value” that has been associated with any “first three-dimensional point”. Thus this makes it unclear as to whether “the feature value” refers to the “a feature of the two-dimensional image” previously recited, or whether it is meant to refer to a feature value derived from a first three-dimensional point or to some inherent or implicit feature value that a three-dimensional point may have. Furthermore, “the associated” implies that some unrecited “association” between a 2D feature and first three-dimensional exists to refer to, thus further leading to a lack of clarity.
Additionally, the detecting step is “detecting…a feature value” in a singular form tied to a singular 2D coordinate value in the singular 2D image, however, the “generating” step requires that “each” of the “plurality of first three-dimensional point data items” has “the feature value of the associated first three-dimensional point” where the associated point would differ from data item to data item. There is no antecedent basis for th plurality of feature values required, and thus it is unclear whether every data item effectively stores the same single detected value or whether each stores a different value.
Furthermore, the “generating” limitation recites a data-item-point association twice, reciting “associated one to one with a plurality of first three-dimensional points included in the first three-dimensional point cloud” and then “each…being associated with corresponding one of first three-dimensional included in the first three-dimensional point cloud.” The second recitation lacks any article for “first three-dimensional points included” and thus is not specifically limited to the previously recited plurality. This makes it unclear as to whether one association or two distinct associations are required, and whether the second association runs to the same plurality of points or to any point of the cloud.
In the interest of compact prosecution, and in order to apply prior art rejections, the Examiner will interpret the claim limitations as if it reads “(ii) the feature value of an associated first three-dimensional point” in line 15 and in line 18, the limitation will be interpreted as “with a corresponding one of the plurality of first three-dimensional points”. Such amendments would render the claims definite.
Note that claim 20 recites the same indefinite claim language and is thus rejected for the same reasons as claim 1 above. In the interest of compact prosecution, for the purposes of applying prior art it will be interpreted in the same manner as claim 1 above.
Additionally note that claims 2-19 carry through the same indefiniteness of the parent claim without curing the deficiency and are each rejected for at least the same reasons as claim 1 above.
Regarding claim 2, the instant claim recites “the plurality of two-dimensional images each being the two-dimensional image” such that this renders the claim indefinite as it is not clear how a plurality of images taken from different positions and/or orientations could actually be a single two-dimensional image. This renders the claim scope indefinite as it is not clear whether some unnamed recitation transforms multiple images into a single 2D image (which has no support in the Specification) or whether it somehow refers to the plurality of 2D images each being of the same type as the captured 2D image. In the interest of compact prosecution, the Examiner will interpret the claim as if line 3 of the claim recites, “a plurality of two-dimensional images including the two-dimensional image are obtained using a camera in different positions and/or orientations” and as if “the plurality of two-dimensional images each being the two-dimensional image” is deleted. This would minimally change the scope of the claim and would retain the intent of the limitation and would render claim definite. Note the claims 3-5 are rejected for carrying through the deficiency of this parent claim and are rejected at least on the same grounds.
Regarding claim 7, the instant claim recites “one or more second three-dimensional points to which the confidence value exceeds the threshold value” and thus recites “the confidence value” as if it has a definite antecedent relating to a 3D point have a confidence value, when claim 1 provides only point data items as having a confidence value. Thus it is unclear as to whether this refers to some unrecited confidence value or whether this is meant to refer to an unrecited association between 3D points being assigned a confidence value as opposed to a point data item. In the interest of compact prosecution the Examiner will interpret the claim as if it recites “a confidence value” which would broaden the manner in which an associated confidence value would be used in the claim and would render the claim definite.
Regarding claim 9, the instant claim recites “determined based on the confidence value,” however no singular confidence value has been previously introduced as rather the “a confidence value” in claim 1 can refer to multiple confidence values for a plurality of point data items in relation to point cloud data. Thus it is unclear whether this refers to some specific singular confidence value or whether it refers to multiple confidence values that have been assigned to the relevant point data item or point cloud item. In the interest of compact prosecution, the Examiner will interpret the claim as if it recites “a confidence value” such that any confidence value may be utilized which would render the claim definite.
Regarding claim 12, the instant claim suffers from a number of defects. First the claim recites “the confidence value that is high” but as noted above, no previous singular confidence value has been designated as high or low or any value. The same problem appears from reciting “the confidence value that is low.” Additionally, the claimed values as “high” or “low” are terms of degree which render the claims indefinite. The term “high” and “low” is not defined by the claim, the specification does not provide a standard for ascertaining the requisite degree, and one of ordinary skill in the art would not be reasonably apprised of the scope of the invention. Rather, threshold values may be used to make determinations about the points, but no such threshold is supplied or disclosed which necessarily says a specific value is “low” or “high”. Additionally, the claim recites the same “the confidence value” as both “high” and “low” as they both refer to the same “first three-dimensional data item having the confidence value” which cannot be true as stated. In the interest of compact prosecution, the Examiner will interpret the claim as if it recites “a confidence value” in both instances which would render the claims definite as a confidence value could be different confidence values depending on the point data item in the point cloud which comprises multiple point data items.
Claim 13 recites, “the first three-dimensional point data item that is selected” as if a first three-dimensional point data has previously been recited as selected, when no such selection has taken place. This lack of antecedent basis renders the claim indefinite as it is not clear whether some unrecited selection is supposed to have taken place or whether this only limits the claim when a selection of a first data item is selected. In the interest of compact prosecution, the Examiner will interpret the claim as if it recites “a first three-dimensional point data item that is selected” which would at least render the claim definite for further examination purposes.
Claim 14 recites, “the first three-dimensional point data item selected” as if a first three-dimensional point data has previously been recited as selected, when no such selection has taken place. This lack of antecedent basis renders the claim indefinite as it is not clear whether some unrecited selection is supposed to have taken place or whether this only limits the claim when a selection of a first data item is selected. In the interest of compact prosecution, the Examiner will interpret the claim as if it recites “a first three-dimensional point data item that is selected” which would at least render the claim definite for further examination purposes. Note that claims 15-16 are also rejected for carrying through this deficiency of their parent claim.
Claim 16 recites, “the confidence value that is higher” and “the confidence value that is lower” as if they exists antecedently. Furthermore, the recitation reads as if the same confidence value is both higher and lower as they both refer to “the confidence value”. Thus the claims are indefinite as they lack antecedent basis and fail to make clear what exactly the confidence values refer to. In the interest of compact prosecution the Examiner will interpret the claim as if it recites “a confidence value” in both instances which would render the claim definite.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1-10 and 12-20 is/are rejected under 35 U.S.C. 102(a)(1)/(a)/(2) as being anticipated by Ilic et al1 (“Ilic”).
Regarding claim 1, as rendered definite as explained above (note the dependent claims are interpreted to carry through the definite rendering as explained above), Ilic teaches a three-dimensional point cloud data processing method for processing, using a processor, three-dimensional point cloud data, the three-dimensional point cloud data processing method comprising (note the preamble requires use of a processor to perform the data processing method, where the method performed is addressed in the rejections below; see Ilic, paragraph 0078 and figure 2 teaching a processor for performing the method as explained below where “computer-executable instructions stored in memory 220, when executed by processor 218, may implement the described image processing techniques” on 3D point cloud data as explained below):
obtaining (note the obtaining of such elements is not required to be performed simultaneously or independently, rather if at any point a 2D image and a 3D point cloud are obtained in any manner this is within the scope of the claim) (i) a two-dimensional image (see Ilic, paragraphs 0141-0142 teaching “a new image frame is acquired with associated data from inertial sensors of the smartphone indicating a position of the smartphone at the time the image frame was captured” where as in paragraph 0123 “each image frame captures a two-dimensional image frame of the object” ) and (ii) a first three- dimensional point cloud (see Ilic, paragraphs 0044-0045 teaching “the result of such processing may be a set of features, represented as a “point cloud,” in which each point may have coordinates in a three dimensional space” and “When image frames capture a common feature from different views, points in the point cloud may provide three dimensional coordinates for features of an object” such that a “point cloud” is obtained from processing; see also paragraphs 0103-0108 teaching “features, combined with positional data determined for image frame 302, may be represented as points in a three dimensional point cloud” and for example, “image frame 302 may be represented as a set of points in the point cloud and may be initially positioned within the point cloud based on the position information of the smartphone at the time image frame 302 was captured. Image frame 302 may be positioned within the point cloud based on a position of a prior image frame within the point cloud” such that the data as in such form combined forms a “three-dimensional point cloud” and is thus obtained; see also paragraphs 0143-0147 teaching “features may be compared to corresponding features in other image frames or already incorporated into a point cloud” and “processing may proceed to block 610 where the local depth map is fused with the depth map. Fusing may include matching points in the localized depth map to points in the depth map” such that these are all examples of obtained point clouds);
detecting, from the two-dimensional image obtained in the obtaining, a feature value of the two-dimensional image, the feature value being associated with a two-dimensional coordinate value in the two-dimensional image (see paragraph 0123 teaching “two image frames 402 and 404 capture a similar feature of object 406. Similar features may be identified by automated processing by, for example, having similar texture in similar locations” and “each image frame captures a two-dimensional image frame of the object” such that here there is detecting of feature values such as “texture” as further explained in paragraph 0089 teaching “analyzing content of image frame 302 to extract features and obtain one or more parameters. Non-limiting examples of the features may comprise lines, edges, corners, colors, junctions and other features. Parameters may comprise sharpness, brightness, contrast, saturation, exposure parameters (e.g., exposure time, aperture, white balance, etc.) and any other parameters. Such parameters may include texture information associated with surfaces of the object at locations near the features” such that these are all examples of feature values where the features and parameters may be considered feature values as they are associated with the features detected; furthermore as in paragraph 0049 it is explained “Texture information, such as color, about the object may be determined from the image frames. The texture information of a particular feature may be determined from the region of the images that depict the feature” such that as these features are detected in 2D images at region locations of the 2D image then these features that can be matched are tied to a 2D coordinate value in the 2D image at the location in any respective 2D image being analyzed; note paragraph 0197 teaching “identification of feature correspondences may include searching, for each point in the image frame, for a corresponding feature in another image frame along a respective epipolar line. The three-dimensional points representing the image frame may be re-projected to establish correspondences to points that may be visible in the current image frame but not in the immediately preceding image frame—e.g., when the current image frame overlaps with a prior image frame other than an immediately preceding image frame in a stream of image frames” such that here in order to conduct such a search for a point in another image frame using such epipolar lines requires a point having a 2D location as the epipolar line would be defined in relation to a 2D point to facilitate matching); and
generating first three-dimensional point cloud data that includes a plurality of first three-dimensional point data items associated one to one with a plurality of first three-dimensional points included in the first three- dimensional point cloud, the plurality of first three-dimensional point data items each being a combination of (note that as a matter of claim interpretation, the step outputs per-point 3D point data item records that have a one to one correspondence with a plurality of the point cloud points and thus each point data item must have one of such records per point, but a “plurality” of first 3D points is a subset so a record is not required for every point in the cloud, and furthermore such data items may contain additional information as the claim language is open-ended but must include the 3D pieces of data recited; see Ilic, paragraph 0127 teaching “features identified in image frames acquired by camera 510 and motion parameter outputs from the sensors 512 may be used to determine three-dimensional coordinates of the features. As image frames are acquired during the scanning of the physical object, more features may be identified and three-dimensional coordinates are included in the three-dimensional content data 514 for the object. The three-dimensional content data may also include accuracy values of the three-dimensional coordinates and/or texture information” such that here “three-dimensional coordinates” that “are included in the three-dimensional content data” are tied to feature values such as texture information as well as “accuracy values” which as explained below correspond to confidence values ) (i) a three-dimensional coordinate value (see Ilic, paragraphs 0044-0045 teaching “the result of such processing may be a set of features, represented as a “point cloud,” in which each point may have coordinates in a three dimensional space. When information about depth of the points relative to the camera is included, this point cloud may serve as a depth map. Estimates of positional uncertainty may be associated with the coordinates. As more image frames are acquired, processing to update the point cloud with new information in the additional image frames may improve certainty for the coordinates” and “When image frames capture a common feature from different views, points in the point cloud may provide three dimensional coordinates for features of an object” such that “coordinates in a three dimensional space” are associated with each detected feature in a point cloud and as in paragraph 0186 “Once features from each of the image frames 1002, 1004, and 1006 are extracted, the image frames may be associated with sets of points representing the features in a three dimensional point cloud space 1009. The point cloud space may be represented in any suitable way, but, in some embodiments, is represented by data stored in computer memory identifying the points and their positions” such that the feature point information related to the image and its 3D position are determined and assigned to a point cloud data item to be stored for each point), (ii) the feature value of an note that “the feature value” refers to the feature value from the 2D image and is not interpreted as an inherent feature value that some associated 3D point may have; see Ilic, paragraph 0186 teaching “Once features from each of the image frames 1002, 1004, and 1006 are extracted, the image frames may be associated with sets of points representing the features in a three dimensional point cloud space 1009. The point cloud space may be represented in any suitable way, but, in some embodiments, is represented by data stored in computer memory identifying the points and their positions” such that here the point cloud data items store “the points” as well as “their positions” such that here the sets of points represent the features meaning the feature value is the point; see also paragraph 0185 teaching that the feature value “may also be processed to improve subsequent feature matching” where “Each image frame may be represented as a set of points representing features extracted from that image frame” such that the points represent features extracted as from the 2D image as explained above where for example such information may be used as in paragraph 0195 teaching “process 1100 may find feature correspondences between pairs of image frames. A succeeding image frame may be compared with one or more previously captured image frames to identify corresponding features between each pair of frames. Such correspondences may be identified based on the nature of the feature characteristics of the image frames surrounding the features or other suitable image characteristics” such that in order to perform such correspondence searching the characteristic or feature value being compared must be stored for that position; see also paragraph 0048 teaching “processing circuitry may maintain an association between the points in the depth map and the image frames from which they were extracted” such that again any point in the depth map point cloud is associated with the image frames that were used to produce it; finally see paragraph 0132 teaching “texture information may correspond to locations of specific features and/or three-dimensional coordinates. Such texture information may be acquired from an image frame depicting those locations. In some embodiments, a representation of surfaces of the object may be defined from the three-dimensional coordinates of points in a point cloud, and texture information may map to points on the surfaces” again teaching that the point cloud stores detected 2D features such as texture information which is associated with 3D locations of points in the point cloud), and (iii) a confidence value (see Ilic, paragraph 0044 teaching “the result of such processing may be a set of features, represented as a “point cloud,” in which each point may have coordinates in a three dimensional space. When information about depth of the points relative to the camera is included, this point cloud may serve as a depth map. Estimates of positional uncertainty may be associated with the coordinates. As more image frames are acquired, processing to update the point cloud with new information in the additional image frames may improve certainty for the coordinates” such that “positional uncertainty” corresponds to a confidence in the position and thus a confidence value where as explained in paragraphs 0144-0148, “In an over constrained system such as may result from multiple image frames showing the same features, some ambiguity may exist in the result, which may be used as an indication of a probability that the correct location of the points representing features has been determined. This information may be used in attaching probabilities to points in a probabilistic depth map” and “Once locations of points associated with features in an image frame are detected, those points may be integrated into a depth map for the object, if warranted” and “data fusion may entail, in addition to updating positions of points in the full depth map, updating the probability/certainty associated with those points. For example, the certainty may increase as the number of image frames processed to determine location of a point increases” such that here a “probability/certainty associated with those points” is a confidence value that is stored for each point and is available for use by other processes for comparison or extraction), each of the plurality of first three-dimensional point data items being associated with a corresponding one of the first three- dimensional points included in the first three-dimensional point cloud (see Ilic, paragraph 0044 as explained above with regard to the coordinates “each point may have coordinates in a three dimensional space. When information about depth of the points relative to the camera is included, this point cloud may serve as a depth map. Estimates of positional uncertainty may be associated with the coordinates” and as in paragraph 0152 for example the features are “associated with each point in the depth map” and as in paragraph 0148, the confidence value explained above is “probability/certainty associated with those points” such that in each case the point cloud data item corresponds to a corresponding one of the first 3D point included in the first 3D point cloud as the probabilistic depth map point cloud that is being updated utilizes the detected features of new images and their assigned 3D positions and confidence values in each update loop to update the point data items),
wherein the confidence value is calculated based on a total number of two-dimensional images in which the associated first three- dimensional point is observed, the confidence value increasing as the total number of two-dimensional images in which the associated first three-dimensional point is observed increases (note that the limitation is interpreted to at least require the confidence value is calculated in any way based on some total number of 2D images in which the associated point is observed where if at any point such total number of images increases a confidence value then the limitations are met, and note that the claim does not require actually calculating a total number of images, as rather the confidence value is positively calculated but may be “based on” a “total number of images in which the associated first three-dimensional point is observed” in any way; see Ilic, paragraphs 0144-0148 as explained above teaching “similar features may be captured in multiple image frames and locations of those features may be determined using geometric analysis” and “In an over constrained system such as may result from multiple image frames showing the same features, some ambiguity may exist in the result, which may be used as an indication of a probability that the correct location of the points representing features has been determined” and teaching “data fusion may entail, in addition to updating positions of points in the full depth map, updating the probability/certainty associated with those points. For example, the certainty may increase as the number of image frames processed to determine location of a point increases” such that this is confidence increasing as the number of image frames processed to determine location of a point increases as the image frame must contain the feature observed in order to determine a location for it which may or may not be used to update the confidence value; see also paragraph 0124 teaching “probabilistic approach may be used to determine a depth estimate. When such a technique is used, a measure of confidence of the depth estimate may be determined. A depth estimate with a high probability may indicate a high measure of confidence in the accuracy of the depth estimate. Although only two frames are shown in FIG. 4, more frames may be used to determine depth information with respect to the smartphone. In some instances, including more image frames that capture the same feature in determining a depth estimate may improve the measure of confidence and reduce the uncertainty associated with the depth estimate” where “more image frames that capture the same feature” leads to a higher “measure of confidence…associated with the depth estimate”; see also paragraph 0044 teaching “Estimates of positional uncertainty may be associated with the coordinates. As more image frames are acquired, processing to update the point cloud with new information in the additional image frames may improve certainty for the coordinates” again showing confidence value tied to more image frames showing a feature and as in paragraph 0128 “If a feature is shown in multiple image frames, and the positional value of the point as determined from each frame is consistent, the accuracy value may be relatively high. Conversely, if the corresponding points in different image frames have different coordinates, the accuracy value for the combined point may be relatively low” again establishing that a higher total of images which depict a feature results in a higher confidence value assigned for the point data items being updated and tracked).
Regarding claim 2, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches the three-dimensional point cloud data processing method according to claim 1, wherein
in the obtaining, a plurality of two-dimensional images including the two-dimensional image are obtained using a camera in different positions and/or orientations, see Ilic, paragraphs 00632-0063 teaching “the object being imaged is three-dimensional and multiple image frames from different perspectives are needed to capture the surface of the object. A user 104 may move the smartphone 102 around object 106 to capture the multiple perspectives. Accordingly, in this example, smartphone 102 is being used in a mode in which it acquires multiple images of an object, such that three dimensional data may be collected” and “Smartphone 200 may include a camera 202” such that here a user may capture multiple perspectives from the camera such that a plurality of 2D images are obtained at different locations and/or orientations corresponding to the perspectives; see also paragraph 0122 teaching “processing within a system 400 to form a data structure containing a three-dimensional representation of an object 406. In this example, image frames 402 and 404 are captured using a smartphone, such as smartphone 102 in FIG. 1. The smartphone is moved to capture different perspectives of object 406 by changing the orientation and position of the smartphone with respect to the object. In this example, image frames 402 and 404 having different orientations and positions to capture different perspectives of object 406” where such frames correspond to a plurality of 2D images obtained),
in the detecting, the feature value is detected for each of the plurality of two-dimensional images obtained in the obtaining (see Ilic, paragraphs 0185-0186 teaching “when an image frame is captured, processing of the image frame includes extracting features. The features may also be processed to improve subsequent feature matching. Each image frame may be represented as a set of points representing features extracted from that image frame” and for example “Once features from each of the image frames 1002, 1004, and 1006 are extracted, the image frames may be associated with sets of points representing the features in a three dimensional point cloud space 1009. The point cloud space may be represented in any suitable way, but, in some embodiments, is represented by data stored in computer memory identifying the points and their positions” such that here for each image obtained there is detecting of the feature value and then it may be associated with a position),
the three-dimensional point cloud data processing method further comprises:
matching feature values associated with two two-dimensional images from the plurality of two-dimensional images using the feature values detected for each of the plurality of two-dimensional images (see Ilic, paragraphs 0194-0196 teaching such matching where “the acquired image frame may be processed by computing that extracts one or more image features from the image frame. The features may be any suitable types of features” and “process 1100 may find feature correspondences between pairs of image frames. A succeeding image frame may be compared with one or more previously captured image frames to identify corresponding features between each pair of frames. Such correspondences may be identified based on the nature of the feature characteristics of the image frames surrounding the features or other suitable image characteristics” where it is further explained that such matching is between pairs or at least two 2D images as in “the processing at block 1110 may involve using a set of features computed for a respective image frame to estimate the epipolar geometry of the pair of the image frames. Each set of features may be represented as a set of points in a three-dimensional space. Thus, the image frame acquired at block 1104 may comprise three-dimensional points projected into a two-dimensional image. When at least one other image frame representing at least a portion of the same object, which may be acquired from a different point of view, has been previously captured, the epipolar geometry that describes the relation between the two resulting views may be estimated” such that for example matching feature values of a feature in two images requires the features to be detected for such comparison); and outputting one or more pairs of the feature values matched (see Ilic, paragraph 0123 and figure 4 teaching an example of a pair of feature values matched and output where “the two image frames 402 and 404 capture a similar feature of object 406. Similar features may be identified by automated processing by, for example, having similar texture in similar locations. Pixels 408 and 410 of image frames 402 and 404, respectively, capture the same region of a feature” and thus these features are matched features which are then output to the next operation where “distance between the two image frames and the location of the similar feature in the image frames may be used to determine depth information” such that this takes the output feature pairs and uses them in the computation of the depth information; see also Ilic, paragraphs 0192-0196 as explained above teaching “building a composite image by representing features of image frames in the three dimensional point cloud” and teaching matching where “the acquired image frame may be processed by computing that extracts one or more image features from the image frame. The features may be any suitable types of features” and “process 1100 may find feature correspondences between pairs of image frames. A succeeding image frame may be compared with one or more previously captured image frames to identify corresponding features between each pair of frames. Such correspondences may be identified based on the nature of the feature characteristics of the image frames surrounding the features or other suitable image characteristics” where it is further explained that such matching is between pairs or at least two 2D images as in “the processing at block 1110 may involve using a set of features computed for a respective image frame to estimate the epipolar geometry of the pair of the image frames. Each set of features may be represented as a set of points in a three-dimensional space. Thus, the image frame acquired at block 1104 may comprise three-dimensional points projected into a two-dimensional image. When at least one other image frame representing at least a portion of the same object, which may be acquired from a different point of view, has been previously captured, the epipolar geometry that describes the relation between the two resulting views may be estimated” such that here then the pairs of the feature values for two “resulting views” that image the same object feature are output as used to determine the 3D position of the feature in the point cloud).
Regarding claim 3, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 2 above and further teaches the three-dimensional point cloud data processing method according to claim 2, further comprising generating a second three-dimensional point cloud including one or more second three-dimensional points each having the feature value, the one or more second three-dimensional points being generated using one or more feature values and one of one or more first three-dimensional points from among the plurality of first three-dimensional points forming the first three-dimensional point cloud (see Ilic, paragraphs 0143-0146 where creation of the localized depth map for each new image frame corresponds to generating a second 3D point cloud including one or more second 3D points each having the feature value as taught where “the new image frame and associated inertial data may be used to calculate a probabilistic depth map. To calculate the depth map, features of the object may be identified in the new image frame. Those features may be compared to corresponding features in other image frames or already incorporated into a point cloud. Positions of the points representing features in the object may be determined in three dimensions” such that here in order to determine the 3D location in the local depth map point cloud, the points of the local depth map are generated using the feature values that are “corresponding features in other image frames or already incorporated into a point cloud” such that this generates a second 3D point cloud as recited; additionally, Ilic teaches another instance of generation of a second 3D point cloud as recited as in paragraphs 0147-0149 teaching “processing may proceed to block 610 where the local depth map is fused with the depth map. Fusing may include matching points in the localized depth map to points in the depth map. The locations of points in the depth map may be adjusted to minimize an error function between the positions of points in both the local depth map and the full depth map. That error function may be a weighted average of positions based on the probabilities associated with each point or the number of image frames used to compute the location of corresponding points in each of the local and full depth maps” and “when the depth map based on the new image frame is determined to be relevant, local fusion of consecutive depth maps may occur by block 610. Depth maps may be fused or merged to form a combined depth map in any suitable way” such that here the combined depth map is generated using the local map’s points and the corresponding pre-existing points of the full depth map which forms what may be considered a second point cloud based on the merging of corresponding points according to the features detected in each image).
Regarding claim 4, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 3 above and further teaches wherein the one or more feature values are each detected from the plurality of two-dimensional images (see Ilic, paragraphs 0185-0186 teaching “when an image frame is captured, processing of the image frame includes extracting features. The features may also be processed to improve subsequent feature matching. Each image frame may be represented as a set of points representing features extracted from that image frame” and for example “Once features from each of the image frames 1002, 1004, and 1006 are extracted, the image frames may be associated with sets of points representing the features in a three dimensional point cloud space 1009. The point cloud space may be represented in any suitable way, but, in some embodiments, is represented by data stored in computer memory identifying the points and their positions” such that here for each image obtained there is detecting of the feature value and then it may be associated with a position; see also paragraphs 0123-0124 teaching “two image frames 402 and 404 capture a similar feature of object 406. Similar features may be identified by automated processing by, for example, having similar texture in similar locations. Pixels 408 and 410 of image frames 402 and 404, respectively, capture the same region of a feature. Although each image frame captures a two-dimensional image frame of the object, the two image frames combined with camera parameters may be used to determine a distance between the object and the smartphone. The camera parameters may include focus, motion, and/or positional information of the camera, which may be used to determine a distance between the two image frames. The distance between the two image frames and the location of the similar feature in the image frames may be used to determine depth information. Such depth information may be an estimate of the distance between the object and the smartphone's position when capturing one of the image frames” such that here a plurality of images capture the same feature which is detected in the plurality of images).
Regarding claim 5, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 3 above and further teaches wherein the one or more feature values include a plurality of feature values, each of the plurality of feature values being a different attribute type (see Ilic, paragraphs 0123-0124 teaching “the two image frames 402 and 404 capture a similar feature of object 406. Similar features may be identified by automated processing by, for example, having similar texture in similar locations. Pixels 408 and 410 of image frames 402 and 404, respectively, capture the same region of a feature. Although each image frame captures a two-dimensional image frame of the object, the two image frames combined with camera parameters may be used to determine a distance between the object and the smartphone. The camera parameters may include focus, motion, and/or positional information of the camera, which may be used to determine a distance between the two image frames. The distance between the two image frames and the location of the similar feature in the image frames may be used to determine depth information” such that here the features include an attribute type corresponding to “texture” and other values related to the feature corresponding to attributes of a different type such as “camera parameters” associated with the image feature; see also paragraph 0089 teaching “acquired image frame 302 may be pre-processed to prepare image frame 302 for further analysis. This may comprise improving quality of image frame 302. The pre-processing 304 may also include analyzing content of image frame 302 to extract features and obtain one or more parameters. Non-limiting examples of the features may comprise lines, edges, corners, colors, junctions and other features. Parameters may comprise sharpness, brightness, contrast, saturation, exposure parameters (e.g., exposure time, aperture, white balance, etc.) and any other parameters. Such parameters may include texture information associated with surfaces of the object at locations near the features” such that again the feature may comprise different feature descriptors which are all different attribute types and may also include features such as the camera parameters for the relevant feature being analyzed).
Regarding claim 6, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein in the detecting, a feature quantity calculated for an area from among a plurality of areas forming the two-dimensional image obtained is detected as the feature value of the two-dimensional image associated with the one of the plurality of first three-dimensional points included in the first three-dimensional point cloud (see Ilic, paragraph 0123 teaching for example “the two image frames 402 and 404 capture a similar feature of object 406. Similar features may be identified by automated processing by, for example, having similar texture in similar locations. Pixels 408 and 410 of image frames 402 and 404, respectively, capture the same region of a feature” such that here areas among a plurality of areas forming the 2D image correspond to pixels of the 2D image at locations or areas corresponding to the features and these are the 2D image features associated with points of the point cloud ).
Regarding claim 7, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches receiving a threshold value from a client device (note that the claim does not define or limit what a “client device” is nor does such receiving require any user to provide a threshold such that if any device acts as a client such as by receiving something from another element, then such is a client device receiving a threshold value; see Ilic, paragraphs 0045-0047 teaching “each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features” and “As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here a threshold is received by the device which is acting as a client device and is used to extract only certain points to combine with other sets of points have a similar level of accuracy or confidence);
extracting, from the plurality of first three-dimensional points, one or more second three-dimensional points to which a confidence value exceeds the threshold value received (see Ilic, paragraphs 0045-0047 as explained above “each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features” and “As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here a threshold is received by the device which is acting as a client device and is used to extract only certain points to combine with other sets of points have a similar level of accuracy or confidence such that these points exceeding the threshold are extracted for the combination); and
transmitting, to the client device, a second three-dimensional point cloud including the one or more second three-dimensional points extracted (see Ilic, paragraphs 0045-0047 as explained above where after determining the points to be extracted they are transmitted to the combination stage of the client device which combines the point clouds).
Regarding claim 8, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein the confidence value is calculated by using three-dimensional coordinates obtained by a distance sensor (note that the claim does not recite how the confidence value is calculated by using 3D coordinates obtained by a distance sensor and furthermore does not limit or define what type of sensor may be considered a distance sensor, and a distance sensor is any sensor which is used to sense a distance in any manner from what is sensed such that this does not limit a distance sensor to any particular type of sensor such as a LIDAR, or structured light sensor but would include any device that senses a distance utilizing sensors which would include for example 2D cameras arranged in a system that provide depth values of the objects sensed by the images; thus see Ilic, paragraphs 0045-0046 teaching “image frames capture a common feature from different views, points in the point cloud may provide three dimensional coordinates for features of an object. Sets of points, each set representing features extracted from an image frame, may be positioned within the depth map. Initially, the sets may be positioned within the depth map based on position information of the smartphone at the time the associated image frame was captured. This positional information may include information such as the direction in which the camera on the phone was facing, the distance between the camera and the object being imaged, the focus and/or zoom of the camera at the time each image frame was captured and/or other information that may be provided by sensors or other components on the smart phone” such that “distance between the camera and the object being imaged” may give a depth map position for the object imaged and “Sets of points in different image frames, representing features in an object being imaged, may be compared in any suitable way. As each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features. In some embodiments, the adjustment may be based on projecting points associated with multiple image frames into a common frame of reference, which may be a plane or volume. When there is overlap between the portions of the object being imaged represented in different image frames, adjacent sets of points will likely include points corresponding to the same image features. By adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved. As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here then the confidence value is based on these depth measurements which can give information about the points to be processed in comparison to the other information to determine the confidence value for a point; see also paragraphs 0123-0124 teaching the cameras acting as a depth sensor as well to determine a confidence value associated with a point where “Although each image frame captures a two-dimensional image frame of the object, the two image frames combined with camera parameters may be used to determine a distance between the object and the smartphone. The camera parameters may include focus, motion, and/or positional information of the camera, which may be used to determine a distance between the two image frames. The distance between the two image frames and the location of the similar feature in the image frames may be used to determine depth information. Such depth information may be an estimate of the distance between the object and the smartphone's position when capturing one of the image frames” and “A probabilistic approach may be used to determine a depth estimate. When such a technique is used, a measure of confidence of the depth estimate may be determined. A depth estimate with a high probability may indicate a high measure of confidence in the accuracy of the depth estimate. Although only two frames are shown in FIG. 4, more frames may be used to determine depth information with respect to the smartphone. In some instances, including more image frames that capture the same feature in determining a depth estimate may improve the measure of confidence and reduce the uncertainty associated with the depth estimate” ).
Regarding claim 9, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein an order of encoding the plurality of first three-dimensional points is determined based on the confidence value (note that the manner of encoding is not limited, nor is the order of encoding utilized for any specified purpose and encoding is any act of converting or transforming data into another form as data converted or processed or changed into a different form then such data would be encoded; see Ilic, paragraph 0046 teaching “As each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features. In some embodiments, the adjustment may be based on projecting points associated with multiple image frames into a common frame of reference, which may be a plane or volume. When there is overlap between the portions of the object being imaged represented in different image frames, adjacent sets of points will likely include points corresponding to the same image features. By adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved. As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here based on the confidence value for a 3D point of the plurality of points, an order of encoding these points so that they are “combined with other sets of points having a similar level of accuracy” is determined as they must follow the order of processing and determining confidence values to encode the plurality of first 3D points into the combined set of points).
Regarding claim 10, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein the feature value is a combination of a plurality of elements (see Ilic, Ilic, paragraphs 0123-0124 teaching “the two image frames 402 and 404 capture a similar feature of object 406. Similar features may be identified by automated processing by, for example, having similar texture in similar locations. Pixels 408 and 410 of image frames 402 and 404, respectively, capture the same region of a feature. Although each image frame captures a two-dimensional image frame of the object, the two image frames combined with camera parameters may be used to determine a distance between the object and the smartphone. The camera parameters may include focus, motion, and/or positional information of the camera, which may be used to determine a distance between the two image frames. The distance between the two image frames and the location of the similar feature in the image frames may be used to determine depth information” such that here the features include a plurality of elements corresponding to “texture” and other values related to the feature corresponding to attributes of a different type such as “camera parameters” associated with the image feature; see also paragraph 0089 teaching “acquired image frame 302 may be pre-processed to prepare image frame 302 for further analysis. This may comprise improving quality of image frame 302. The pre-processing 304 may also include analyzing content of image frame 302 to extract features and obtain one or more parameters. Non-limiting examples of the features may comprise lines, edges, corners, colors, junctions and other features. Parameters may comprise sharpness, brightness, contrast, saturation, exposure parameters (e.g., exposure time, aperture, white balance, etc.) and any other parameters. Such parameters may include texture information associated with surfaces of the object at locations near the features” such that again the feature may comprise a plurality of different elements of different feature descriptors which are all different attribute types and may also include features such as the camera parameters for the relevant feature being analyzed).
Regarding claim 12, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein out of the plurality of first three- dimensional point data items, a first three-dimensional point data item having a confidence value that is high is made more likely to be selected than a first three-dimensional point data item having a confidence value that is low (note that the claim does not define or limit what the “to be selected” is selected for and as noted above does not provide any ascertainable scope for what is “high” vs “low” as a confidence value, rather leaving them as relative terms of degree whose scope is unclear; see Ilic, paragraphs 0046-0053 teaching “each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features” and “adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved” and “When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy. In this way, a coarse alignment of image frames, associated with the sets of points, may be achieved” and “processing circuitry may maintain an association between the points in the depth map and the image frames from which they were extracted. Once the relative position, orientation, zoom and/or other positional characteristics are determined with respect to a common reference for the sets of points, those points may be used to identify structural features of an object. Points may be grouped, for example, to represent structural features. The structural features may be represented as a convex hull” and “Structures may be identified based on the point cloud in any suitable way, which may include filtering or other pre-processing. The depth map, for example, may be smoothed and/or filtered before determining a three-dimensional volume representing the object. As an example of filtering, a threshold value on the positional uncertainty may be used to select the points to include in the volume. Points having an accuracy level above the threshold may be included while points below the threshold may be discarded” such that here the first 3D point data items correspond to the processed points having their features, positions and confidence values determined in relation to the information from all of the images and the 3D data point items having a high confidence or accuracy level assigned to the points are more likely to be selected for outputting to the next stage to build the convex hull whereas lower level confidence points are discarded).
Regarding claim 13, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein the confidence value of the first three- dimensional point data item that is selected is greater than or equal to a threshold value (see Ilic, paragraphs 0046-0053 as explained above, teaching “each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features” and “adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved” and “When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy. In this way, a coarse alignment of image frames, associated with the sets of points, may be achieved” and “processing circuitry may maintain an association between the points in the depth map and the image frames from which they were extracted. Once the relative position, orientation, zoom and/or other positional characteristics are determined with respect to a common reference for the sets of points, those points may be used to identify structural features of an object. Points may be grouped, for example, to represent structural features. The structural features may be represented as a convex hull” and “Structures may be identified based on the point cloud in any suitable way, which may include filtering or other pre-processing. The depth map, for example, may be smoothed and/or filtered before determining a three-dimensional volume representing the object. As an example of filtering, a threshold value on the positional uncertainty may be used to select the points to include in the volume. Points having an accuracy level above the threshold may be included while points below the threshold may be discarded” such that here the first 3D point data items correspond to the processed points having their features, positions and confidence values determined in relation to the information from all of the images and the 3D data point items having a high confidence or accuracy level assigned to the points are more likely to be selected for outputting to the next stage to build the convex hull whereas lower level confidence points are discarded, and this involves a “level above the threshold”).
Regarding claim 14, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein an encoding process is performed on a first three-dimensional point data item selected (note that “encoding process” here is interpreted under its broadest reasonable interpretation similarly to claim 9 above and is any conversion or transformation of the state or purpose of some data into another form or use; see Ilic, paragraphs paragraph 0046 teaching “As each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features. In some embodiments, the adjustment may be based on projecting points associated with multiple image frames into a common frame of reference, which may be a plane or volume. When there is overlap between the portions of the object being imaged represented in different image frames, adjacent sets of points will likely include points corresponding to the same image features. By adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved. As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here based on the confidence value for a 3D point of the plurality of point data items, an encoding of these points so that they are “combined with other sets of points having a similar level of accuracy” is determined as they must determine confidence values to encode the plurality of first 3D points into the combined set of points), and a resultant first three-dimensional point data item is output (see Ilic, paragraph 0046 as explained above teaching “As each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features. In some embodiments, the adjustment may be based on projecting points associated with multiple image frames into a common frame of reference, which may be a plane or volume. When there is overlap between the portions of the object being imaged represented in different image frames, adjacent sets of points will likely include points corresponding to the same image features. By adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved. As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here the combined sets of points is a resultant 3D point data item), the resultant first three-dimensional point data item resulting from performing the encoding process on the first three-dimensional point data item selected (see Ilic, paragraph 0046 as explained above teaching “As each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features. In some embodiments, the adjustment may be based on projecting points associated with multiple image frames into a common frame of reference, which may be a plane or volume. When there is overlap between the portions of the object being imaged represented in different image frames, adjacent sets of points will likely include points corresponding to the same image features. By adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved. As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here the resulting combined sets of points is based on encoding the 3D point data item selected to be in its combined form).
Regarding claim 15, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 14 above and further teaches wherein an order of selecting the first three- dimensional point data item to be output from among the plurality of first three-dimensional point data items is determined based on the confidence value, the encoding process is performed based on the order, and the resultant first three-dimensional point data item is output (note that the claim does not define or limit how the order of selecting is based on the confidence value and does not require a specific type of order so long as any order in which the selection occurs is somehow based on the confidence value; see Ilic, paragraph 0046 as explained above teaching “As each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features. In some embodiments, the adjustment may be based on projecting points associated with multiple image frames into a common frame of reference, which may be a plane or volume. When there is overlap between the portions of the object being imaged represented in different image frames, adjacent sets of points will likely include points corresponding to the same image features. By adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved. As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here in order to select the first 3D point data item to be output to be combined with other points of similar accuracy at the same point, the confidence value of the 3D point data item must be determined such that then any order of selection for outputting to the encoding stage requires a confidence value and then the encoding to combine the points is based on that order which is any order of points so long as they have a confidence value and are able to be combined and the result combined first 3D point data item is output based on this encoding of the point into the combined point for each corresponding point analyzed).
Regarding claim 16, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein out of the plurality of first three- dimensional point data items, a first three-dimensional point data item having a confidence value that is higher is preferentially encoded to a first three-dimensional point data item having a confidence value that is lower (see Ilic, paragraph 0046 teaching “As each set of points is added to the depth map, its three-dimensional position may be adjusted to ensure consistency with sets of points containing points representing an overlapping set of features. In some embodiments, the adjustment may be based on projecting points associated with multiple image frames into a common frame of reference, which may be a plane or volume. When there is overlap between the portions of the object being imaged represented in different image frames, adjacent sets of points will likely include points corresponding to the same image features. By adjusting the three dimensional position associated with each set of points to achieve coincidence in the frame of reference between points representing the same features, accuracy of the coordinates of the points can be improved. As more image frames are gathered, the accuracy of the points in the depth map may be improved. When points corresponding to a particular feature or region of the frame of reference have a measure of accuracy exceeding a threshold or meeting some other criteria, those points may be combined with other sets of points having a similar level of accuracy” such that here the system prefers to combine and thus encode first 3D point data items that have a higher level to those that are lower and do not pass a confidence threshold).
Regarding claim 17, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein the confidence value indicates certainty of (a) correspondence between (i) a three-dimensional coordinate value of the associated first three-dimensional point and (ii) the two-dimensional coordinate value in the two- dimensional image (note that here the correspondence is not defined by any particular metric or technique such that any confidence value that indicates certainty that some point’s 3D coordinate and the image location at issue are the same or in some correspondence; see Ilic, paragraphs 0122-0128 teaching “processing within a system 400 to form a data structure containing a three-dimensional representation of an object 406. In this example, image frames 402 and 404 are captured using a smartphone, such as smartphone 102 in FIG. 1. The smartphone is moved to capture different perspectives of object 406 by changing the orientation and position of the smartphone with respect to the object. In this example, image frames 402 and 404 having different orientations and positions to capture different perspectives of object 406” and “the two image frames 402 and 404 capture a similar feature of object 406. Similar features may be identified by automated processing by, for example, having similar texture in similar locations. Pixels 408 and 410 of image frames 402 and 404, respectively, capture the same region of a feature. Although each image frame captures a two-dimensional image frame of the object, the two image frames combined with camera parameters may be used to determine a distance between the object and the smartphone” and “distance between the two image frames and the location of the similar feature in the image frames may be used to determine depth information. Such depth information may be an estimate of the distance between the object and the smartphone's position when capturing one of the image frames” and “probabilistic approach may be used to determine a depth estimate. When such a technique is used, a measure of confidence of the depth estimate may be determined. A depth estimate with a high probability may indicate a high measure of confidence in the accuracy of the depth estimate. Although only two frames are shown in FIG. 4, more frames may be used to determine depth information with respect to the smartphone. In some instances, including more image frames that capture the same feature in determining a depth estimate may improve the measure of confidence and reduce the uncertainty associated with the depth estimate” and “features identified in image frames acquired by camera 510 and motion parameter outputs from the sensors 512 may be used to determine three-dimensional coordinates of the features. As image frames are acquired during the scanning of the physical object, more features may be identified and three-dimensional coordinates are included in the three-dimensional content data 514 for the object. The three-dimensional content data may also include accuracy values of the three-dimensional coordinates and/or texture information” and “accuracy value may indicate a level of uncertainty of a three-dimensional coordinate and may be used in determining whether the coordinate is included in forming a composite image or representation of the object. The accuracy values, for example, may be computed based on a distribution of points in a point cloud and may represent the certainty or likelihood that a point accurately represents a feature of an object. If a feature is shown in multiple image frames, and the positional value of the point as determined from each frame is consistent, the accuracy value may be relatively high. Conversely, if the corresponding points in different image frames have different coordinates, the accuracy value for the combined point may be relatively low” such that here the accuracy value measures the agreement or correspondence between the point’s 3D coordinate and its point’s appearance in each image frame such that when the 3D position derived from a given frame’s observation agrees with others, then the accuracy is higher, meaning that the value indicates a certainty that a stored 3D coordinate corresponds to the image location at which the feature was observed), or (b) correspondence between (i) the three-dimensional coordinate value of the associated first three-dimensional point out of the plurality of first three-dimensional points and (ii) the feature value of the associated first three-dimensional point (note that that is functionally a value that reflects how reliable it is that the point’s 3D coordinate really is the location of the feature whose value the record carries; see Ilic, paragraph 0128 teaching “accuracy value may indicate a level of uncertainty of a three-dimensional coordinate and may be used in determining whether the coordinate is included in forming a composite image or representation of the object. The accuracy values, for example, may be computed based on a distribution of points in a point cloud and may represent the certainty or likelihood that a point accurately represents a feature of an object” such that here the confidence value is “the certainty or likelihood that a point accurately represents a feature of an object”).
Regarding claim 18, as rendered definite as explained above, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein the confidence value is calculated using a matching error of multiple feature points corresponding to the associated first three-dimensional point (note that a matching error is not limited to any specific form or technique but simply must be some value representing a degree to which matching fails or is less than optimal which could include a matching error of 0 indicating matching for example, and here the error must be from two or more image feature points that correspond to a single point in the cloud; see Ilic, paragraph 0152 as referred to in the rejection of claim 1 addressing confidence value calculation and teaching “the accuracy of the points in the depth map may be used when processing the depth map. The accuracy of the points may be determined when the probabilistic depth map was calculated in block 606. A threshold level of accuracy may be applied to the points in the depth map. Points that are more accurate may be kept while points that are less accurate may be discarded. Such a smoothing process may remove points from the model that are unlikely to be correct and/or produce artifacts in the smooth model” such that the creation of this depth map with the confidence values corresponds to the confidence values of the probabilistic depth maps as explained in paragraphs 0142-0145 teaching “For example, similar features may be captured in multiple image frames and locations of those features may be determined using geometric analysis. Detected motion of the portable electronic device between capture of those image frames may serve as known values in a system of equations with the locations of the points in the point cloud representing those features as the unknowns for which a solution may be determined using linear algebraic techniques or other suitable computations. In an over constrained system such as may result from multiple image frames showing the same features, some ambiguity may exist in the result, which may be used as an indication of a probability that the correct location of the points representing features has been determined. This information may be used in attaching probabilities to points in a probabilistic depth map” and “a probability may be associated with a position of a point representing a feature in an image frame, the probability that the features in multiple image frames used to compute location of the point actually represent the same feature of the object may be determined. That determination may be made using a cross-correlation to determine regions of similarity between image frames. A matching score for a depth may be determined based on the cross-correlation between image frames. For example, when two image frames have a region of similarity with a high cross-correlation, the matching score may be high. An accurate depth estimate may be indicated by a high matching score” such that here a “matching score” corresponds to a matching error where a low matching score corresponds to a higher error in matching and high matching indicates a lower error in matching and the matching is between “features in multiple image frames used to compute the location of the point” and the score is what determines the confidence value; see also paragraph 0128 and figure 5 teaching “accuracy value may indicate a level of uncertainty of a three-dimensional coordinate and may be used in determining whether the coordinate is included in forming a composite image or representation of the object. The accuracy values, for example, may be computed based on a distribution of points in a point cloud and may represent the certainty or likelihood that a point accurately represents a feature of an object. If a feature is shown in multiple image frames, and the positional value of the point as determined from each frame is consistent, the accuracy value may be relatively high. Conversely, if the corresponding points in different image frames have different coordinates, the accuracy value for the combined point may be relatively low” such that here matching is attempted between locations of similar image features in frames which yields positions of the features which are compared to determine if they match and if not they are given a low confidence or accuracy rating).
Regarding claim 19, Ilic teaches all that is required as applied to claim 1 above and further teaches wherein the confidence value is calculated using a matching error between the associated first three-dimensional point and feature points corresponding to the associated first three-dimensional point (note the claim introduces the new term “feature points corresponding to the associated first three-dimensional point” such that it is left undefined what such feature points are or where they come from or necessarily what they refer to, however, as they do not specifically refer to antecedent feature points they can be interpreted as introduced such that the feature point could refer to feature values or attributes of the 2D image or could refer to other types of feature points that correspond to an associated first 3D point in any manner, and note that this would include a point at the image location where it was detected and would also include a point representing that feature in space as determined from a single image frame; furthermore note that “the associated first three-dimensional point” corresponds to the single point cloud point that the confidence value belongs to such that a matching error between this point and feature points would for example correspond to any error measured in 3D space between the point and the per-frame positions determined for that feature; thus see Ilic, paragraphs 0127-0128 teaching “features identified in image frames acquired by camera 510 and motion parameter outputs from the sensors 512 may be used to determine three-dimensional coordinates of the features. As image frames are acquired during the scanning of the physical object, more features may be identified and three-dimensional coordinates are included in the three-dimensional content data 514 for the object. The three-dimensional content data may also include accuracy values of the three-dimensional coordinates and/or texture information” and “accuracy value may indicate a level of uncertainty of a three-dimensional coordinate and may be used in determining whether the coordinate is included in forming a composite image or representation of the object. The accuracy values, for example, may be computed based on a distribution of points in a point cloud and may represent the certainty or likelihood that a point accurately represents a feature of an object. If a feature is shown in multiple image frames, and the positional value of the point as determined from each frame is consistent, the accuracy value may be relatively high. Conversely, if the corresponding points in different image frames have different coordinates, the accuracy value for the combined point may be relatively low” such that the associated point corresponds to the combined point resulting from the feature observations being combined being assigned a confidence/accuracy value and the feature points corresponding to it are the corresponding points in different image frames which are positions determined for that same feature from one frame and the matching error between them is the extent to which those corresponding points have different coordinates such that consistent determinations give a higher confidence value).
Regarding claim 20, as rendered definite as explained above, the instant claim recites an apparatus as a “three-dimensional point cloud data processing device for processing three-dimensional point cloud data, the three-dimensional point cloud data processing device comprising a processor, wherein the processor” performs the same steps as performed by the processor in the method of claim 1. Ilic teaches such a method as in claim as a device performing the method (see Ilic, paragraph 0078 and figure 2 teaching a processor for performing the method as explained below where “computer-executable instructions stored in memory 220, when executed by processor 218, may implement the described image processing techniques” on 3D point cloud data as explained in the rejection of claim 1). In light of this, the limitations of claim 20 correspond to the limitations of claim 1; thus it is rejected on the same grounds as claim 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Ilic in view of Sinha et al2 (“Sinha”).
Regarding claim 11, Ilic teaches all that is required as applied to claim 1 above but fails to specifically teach wherein the feature value is expressed by a 256-bit data string. Ilic does teach the feature value that can be feature characteristics that are computed for the purpose of feature matching (see Ilic, paragraph 0185-0195 teaching “when an image frame is captured, processing of the image frame includes extracting features. The features may also be processed to improve subsequent feature matching. Each image frame may be represented as a set of points representing features extracted from that image frame. FIG. 10 illustrates schematically exemplary image frames 1002, 1004 and 1006 each representing a portion of a scene and acquired as the scene is being scanned by a smartphone. In this example, image frames 1002, 1004 and 1006 represent different portions of the same object, an apple” and “process 1100 may find feature correspondences between pairs of image frames. A succeeding image frame may be compared with one or more previously captured image frames to identify corresponding features between each pair of frames. Such correspondences may be identified based on the nature of the feature characteristics of the image frames surrounding the features or other suitable image characteristics” and these are retained per point for the point cloud as in paragraph 0197 teaching “identification of feature correspondences may include searching, for each point in the image frame, for a corresponding feature in another image frame along a respective epipolar line. The three-dimensional points representing the image frame may be re-projected to establish correspondences to points that may be visible in the current image frame but not in the immediately preceding image frame—e.g., when the current image frame overlaps with a prior image frame other than an immediately preceding image frame in a stream of image frames” such that here it can be seen that the features for the images are retained as associated with the 3D points in a point data item). However, Ilic fails to teach actually storing a feature value as a 256-bit data string and thus stands as a base device upon which the claimed invention can be seen as an improvement in terms of storing such type of information useable as claimed in a compact 256-bit data string.
In the same field of endeavor relating to building a 3D point cloud from images whose detected features are matched to determine 3D locations of points imaged in the images, Sinha teaches that it is known to represent a feature value as a 256-bit data string when performing matching operations between features detected in different images of the same feature to determine 3D locations of an imaged point in a point cloud (see Sinha, paragraphs 0042-0043 teaching “for each consecutive image frame input during the image-based localization process, keypoints are identified in the current image frame and identified in one or more previously-input image frames. In this way, identified keypoints are tracked from frame to frame” and “for a μ×μ pixel, square patch around each keypoint, a 256-bit Binary Robust Independent Elementary Feature (BRIEF) descriptor is computed. BRIEF descriptors lack rotational and scale invariance, but can be computed very fast. The tracked keypoints from the prior frame are compared to all the keypoint in the current frame, within a ρ×ρ search window around its respective positions in the prior frame” such that here a feature value is stored as a “256-bit Binary Robust Independent Elementary Feature (BRIEF) descriptor” for comparison with other features which is a 256 bit data string). Thus Sinha teaches a known technique applicable to the base system of Ilic above.
Therefore it would have been obvious for one of ordinary skill in the art to modify Ilic to apply the teachings of Sinha as doing so would be no more than application of a known technique to a base device ready for improvement where the modification would yield predictable results and would result in an improved system. The predictable result of combining Sinha’s technique with Ilic would be that the feature value identified in Ilic as explained above would be expressed as a 256-bit data string which Ilic’s matching step would compare bitwise to find the same frame to frame correspondences already being found. This would result in an improved system because the matching could be performed quickly and in a fixed, compact per-point storage space as suggested by Sinha (see Sinha, paragraph 0042-0043 teaching “BRIEF descriptors lack rotational and scale invariance, but can be computed very fast. The tracked keypoints from the prior frame are compared to all the keypoint in the current frame, within a ρ×ρ search window around its respective positions in the prior frame. BRIEF descriptors are compared using Hamming distance (computed with bitwise XOR followed bit-counting using the parallel bit-count algorithm)” and paragraph 0023 teaching “embodiments of the technique involve efficiently interleaving a fast keypoint tracker that uses inexpensive binary feature descriptors with an approach for direct 2D-to-3D matching”).
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 12189037. Although the claims at issue are not identical, they are not patentably distinct from each other as explained below.
Note the following table showing the independent claims with similarities rendered in bold.
Conflicting US Patent No. 12189037
Pending Application 18213040
Claim 1.
A three-dimensional point cloud data processing method for processing, using a processor, three-dimensional point cloud data, the three-dimensional point cloud data processing method comprising:
obtaining (i) a two-dimensional image obtained using a camera and (ii) a first three-dimensional point cloud obtained using a distance sensor;
detecting, from the two-dimensional image obtained in the obtaining, a feature value of the two-dimensional image, the feature value being associated with a two-dimensional coordinate value in a two-dimensional image and one of a plurality of first three-dimensional points included in the first three-dimensional point cloud;
generating first three-dimensional point cloud data that includes a plurality of first three-dimensional point data items associated one to one with the plurality of first three-dimensional points included in the first three-dimensional point cloud, the plurality of first three-dimensional point data items each being a combination of (i) a three-dimensional coordinate value of an associated first three-dimensional point out of the plurality of first three-dimensional points, (ii) the feature value of the associated first three-dimensional point, and (iii) a confidence value of the three-dimensional coordinate value of the associated first three-dimensional point, each of the plurality of first three-dimensional point data items being associated with each of the plurality of first three-dimensional points included in the first three-dimensional point cloud; and
selecting, based on the confidence value, a first three-dimensional point from among the plurality of first three-dimensional point cloud to output a first three-dimensional point data item associated to the selected first three-dimensional point,
wherein the confidence value indicates certainty of a coordinate position of the first three-dimensional point, and
wherein the confidence value is calculated based on a total number of two-dimensional images in which the associated three-dimensional point is observed, the confidence value increasing as the total number of two-dimensional images in which the associated three-dimensional point is observed increases.
Claim 1.
A three-dimensional point cloud data processing method for processing, using a processor, three-dimensional point cloud data, the three-dimensional point cloud data processing method comprising:
obtaining (i) a two-dimensional image and (ii) a first three-dimensional point cloud;
detecting, from the two-dimensional image obtained in the obtaining, a feature value of the two-dimensional image, the feature value being associated with a two-dimensional coordinate value in the two-dimensional image; and
generating first three-dimensional point cloud data that includes a plurality of first three-dimensional point data items associated one to one with a plurality of first three-dimensional points included in the first three-dimensional point cloud, the plurality of first three-dimensional point data items each being a combination of (i) a three-dimensional coordinate value, (ii) the feature value of the associated first three-dimensional point, and (iii) a confidence value, each of the plurality of first three-dimensional point data items being associated with corresponding one of first three-dimensional points included in the first three-dimensional point cloud,
wherein the confidence value is calculated based on a total number of two-dimensional images in which the associated first three-dimensional point is observed, the confidence value increasing as the total number of two-dimensional images in which the associated first three-dimensional point is observed increases.
Thus it can be seen that the instant claims 1 and 20 are a broader version of the conflicting patented claim and thus are effectively a genus to the species recited in claim 1 of the conflicting patent and thus the species anticipates the broader genus claim. The remaining features of the dependent claims of the pending application correspond to the remaining dependent claims of the conflicting patent. Thus they are rejected as also anticipated by the corresponding species claim of the conflicting patent.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SCOTT E SONNERS whose telephone number is (571)270-7504. The examiner can normally be reached Mon-Friday 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SCOTT E SONNERS/Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613
1 US PGPUB No. 2017/0085733
2 US PGPUB No. 20140010407