DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Allowable Subject Matter
Claims 3-8 are objected to as being dependent upon a rejected base claim 2 but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 10-15 are objected to as being dependent upon a rejected base claim 9 but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 17-21 are objected to as being dependent upon a rejected base claim 16 but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Response to Preliminary Amendment
Applicant’s preliminary amendment filed 4/22/2025 has been entered. The claim 1 has been cancelled. The claims 2-21 have been newly added. The claims 2-21 are pending in the current application.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2, 9 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Schonberger et al. US-PGPUB No. 2020/0372672 (hereinafter Schonberger) in view of Garud et al. US-PGPUB No. 2020/0334857 (hereinafter Garud); Barnes et al. US-PGPUB No. 2021/0056668 (hereinafter Barnes); Li et al. US-PGPUB No. 2020/0226782 (hereinafter Li).
Re Claim 2:
Shonberger teaches a method of aligning a first set of features, derived from at least one image collected on a portable electronic device, at least partially to a second set of features in a stored map (Schonberger teaches at Paragraph 0055 estimating the pose of a camera device may involve determining correspondences between image features detected at 2D pixel positions of images of a real-world environment and map features having 3D map positions in a digital environment map and at Paragraph 0011 that the digital environment map is essentially a spatial database including geometric data), the method comprising:
identifying a plurality of pairs of matched features from the first and second sets of features, each of the plurality of pairs of matched features comprising a first feature of the first set and a second feature of the second set (Schonberger teaches at Paragraph 0058 correspondences are identified between image features and map features at 414 and at Paragraph 0021 that the image 110A includes a plurality of features 112 and at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0061 that the remote device may identify a subset of correspondences from the overall set of determined 2D image features to 3D map feature correspondences. Schonberger teaches at Paragraph 0040 that image features are detected by the camera device. Each image feature includes a descriptor 506, as well as a 2D pixel position 508 at which the image feature was detected in image 502. Each descriptor 506 will include some representation of the visual content of image 502 at which a corresponding image feature was detected—for instance, taking the form of a multi-dimensional feature vector.);
for each of the plurality of pairs of matched features, computing a quality metric (Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0067 that the remote device may determine how many of the determined correspondences would be inliers and how many would be outliers); and
using the plurality of pairs of matched features and respective quality metrics for aligning the first set of features to the second set of features by randomly selecting and processing subsets of matched feature pairs until a stop condition is reached (Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0067 that the remote device may determine how many of the determined correspondences would be inliers and how many would be outliers and at Paragraph 0062 a plurality of subsets of 2D data point pairs may be selected and in a RANSCA framework, these subsets are selected at random.
Stop condition:
Schonberger teaches at Paragraph 0072 that the remote device may optionally continue generating pose estimates until a pose that exceeds the preliminary confidence threshold is estimated or the process is stopped for another reason and at Paragraph 0076 that once the final estimated pose is received, the camera device may in some cases terminate the image-based localization process. For instance, the camera device may discontinue capturing new images of the real-world environment),
wherein:
the quality metric for a pair of matched features indicates a likelihood of finding an alignment between the first set of features and the second set of features with a subset of the plurality of pairs of matched features if the pair of matched features are included in the subset of the plurality of pairs of matched features (Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0067 that the remote device may determine how many of the determined correspondences would be inliers and how many would be outliers and at Paragraph 0067 that in the case of RANSAC, this confidence value may be proportional to the inlier/outlier ratio described above, in which pose estimates that have more inliers (e.g., have more correspondences consistent with the pose than inconsistent with the pose) have relatively higher confidence values); and
randomly selecting comprises a random selection weighted based on the quality metrics of the plurality of pairs (and at Paragraph 0062 a plurality of subsets of 2D data point pairs may be selected and in a RANSCA framework, these subsets are selected at random. If the minimal solver is applied to each subset (i.e., pair of 2D points), different potential solutions may be found from each subset. However, some potential solutions (2D lines) will be more consistent with the input data than other potential solutions. In other words, some potential lines will pass through more 2D points in the dataset than others. In RANSAC, for each potential solution identified by a minimal solver, the solution (i.e., 2D line) is compared to other 2D data points in the dataset to determine which points are consistent with the proposed solution (i.e., “inliers”) vs. which points are not consistent with the proposed solution (i.e., “outliers”). Any solution that has the highest inlier to outlier ratio, or has at least a threshold number of total inliers, may be accepted as the actual solution—e.g., 2D the line that best fits the set of 2D points.
Thus the subsets of correspondences are weighted based on the highest inlier to outlier ratio, meaning that the outlier/incorrect correspondences are weighted 0 and the inlier/correct correspondences are weighted 1).
For the above reasons, Schonberger implicitly teaches quality metrics and the selection weighted while Garud explicitly teaches quality metric and the selection weighted.
Garud teaches at Paragraph 0029 that 0029] At block 612, a set of candidate matches between the set of image feature points and the obtained map feature points are determined. For example, the matching may be performed using a cost function in conjunction with 2-way correspondence to determine a set of candidate matching feature points. In certain cases, candidate matches may be determined by matching an image feature descriptor of an image feature point against the obtained map feature descriptors to determine a 2D-3D matched feature point pair, matching a map feature descriptor of the candidate matched feature point against the image feature descriptors of the set of image feature points to determine a 3D-2D matched feature point pair, and determining a candidate match based on quality of the match and comparison between the 2D-3D matched feature point pairs and the 3D-2D matched feature point pairs. The quality of the match may refer to how well the image feature points and the obtained map feature points match. For example, where a cost function is used such as SAD, the output of the cost function reflects the quality of the match. At block 614, a pose of the camera may be determined from a set of potential poses of the camera from candidate matches from the set of candidate matches and an associated reprojection error estimated for the remaining points to select a first pose of the set of potential poses based on having a lowest associated error. For example, a random sample consensus (RANSAC) algorithm may be applied. In certain cases, repeatedly estimating the pose may include randomly selecting a subset of point pairs from the set of candidate matches, determining a camera pose based on the selected feature points, generating a 2D projection of the remaining matched map feature points based on the determined camera pose, determine the Euclidean distance i.e. reprojection error value associated with the generated 2D projections and matching image feature point locations to, repeating the steps of randomly selecting feature points, determining a camera pose, generating a 2D projection, and matching map feature points to generate a set of error values, and selecting the camera pose associated with the lowest error value.
It would have been obvious to one of the ordinary skill in the art to have incorporated Garud’s quality metric of the match and comparison between the 2D-3D matched feature point pairs of Schonberger to have determined the camera pose from the candidate matches from the set of candidate matches to have provided the lowest error value in estimating the camera pose. One of the ordinary skill in the art would have determined the candidate matches based on the quality of the match for the determination of the camera pose and the random selection of the subsets of the matching features from the candidate matches based on the quality of the matches, meaning that the matching pairs of the candidate matches are weighted 1 and the matching pairs of the non-candidate matches are weighted 0.
Barnes further teaches at Paragraph 0022 that 0022] Subsequently, a first subset (e.g., some, but not all) of matched feature pairs of the plurality of matched feature pairs are selected randomly, or based on some pre-established criteria (e.g., sparse objects that are relatively large are selected over sparse objects that are relatively small). For example, an appropriate RANdom SAmple Consensus algorithm, or “RANSAC” algorithm, is used for the selection. If a selected second matched feature pair includes a second feature from the primary image and a second feature from the auxiliary image, then the geometric transformations aim to align the first and second features from the primary image with the respective first and second features from the auxiliary image. Subsequently, a score is generated that indicates the quality of the transformed matches. The score can be a function of, for example, how evenly the selected matched feature pairs are distributed throughout the image, how close the selected matched feature pairs are to the region to be filed, and/or the geometric transformations performed. If the score is less than a threshold, the process is repeated. For example, if the score is less than the threshold, a second subset of matched feature pairs of the plurality of matched feature pairs are selected randomly, one or more geometric transformations are selected and performed, and another score is generated. This process continues until a termination condition has been satisfied.
Thus the selection is more weighted based on the matched feature pairs of relatively large sparse objects than the matched feature pairs of the relatively small sparse objects.
It would have been obvious to one of the ordinary skill in the art to have incorporated Barnes’s weighted selection of the matched feature pairs based on the quality metric into Schonberger’s selection of the matched feature pairs to have determined the camera pose from the candidate matches from the set of candidate matches to have provided the lowest error value in estimating the camera pose. One of the ordinary skill in the art would have determined the subset of the matching pairs based on the quality of the match for the determination of the camera pose and the random selection of the subsets of the matching features from the candidate matches based on the quality of the matches, meaning that the matching pairs of the candidate matches are weighted 1 and the matching pairs of the non-candidate matches are weighted 0.
Li teaches at Paragraph 0139 obtaining a conversion relationship between the coordinate system of the visual map and the coordinate system of the global grid map according to a relative installation position of the video camera device.
It would have been obvious to one of the ordinary skill in the art before the filing date of the instant application to have incorporated Li’s conversion of the coordinate system of the visual map and the coordinate system of the global map (world map or environment map) to have determined the correspondences of Schonberger of the matching pairs. One of the ordinary skill in the art would have been motivated to have compared the image features of the two different coordinate systems.
Re Claim 9:
Schonberger teaches an electronic system that supports specification of a position of virtual content relative to stored maps in a database of stored maps (Schonberger teaches at Paragraph 0055 estimating the pose of a camera device may involve determining correspondences between image features detected at 2D pixel positions of images of a real-world environment and map features having 3D map positions in a digital environment map and at Paragraph 0011 that the digital environment map is essentially a spatial database including geometric data), the system comprising:
a communication component configured to receive from a portable electronic device information about a first set of features in a three-dimensional (3D) environment of the portable electronic device (Schonberger teaches at Paragraph 0019-0022 that the camera device 106 takes the form of an HMD and the images 110A-110C are captured by camera device 106 to include a plurality of image features represented by black circles 112); and
a computing component, connected to the communication component, the computing component configured to:
identify a plurality of pairs of matched features from the first set of features and a second set of features in a stored map of the database of stored maps, each of the plurality of pairs of matched features comprising a first feature of the first set and a second feature of the second set (Schonberger teaches at Paragraph 0058 correspondences are identified between image features and map features at 414 and at Paragraph 0021 that the image 110A includes a plurality of features 112 and at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0061 that the remote device may identify a subset of correspondences from the overall set of determined 2D image features to 3D map feature correspondences. Schonberger teaches at Paragraph 0040 that image features are detected by the camera device. Each image feature includes a descriptor 506, as well as a 2D pixel position 508 at which the image feature was detected in image 502. Each descriptor 506 will include some representation of the visual content of image 502 at which a corresponding image feature was detected—for instance, taking the form of a multi-dimensional feature vector);
for each of the plurality of pairs of matched features, compute a quality metric based on feature information indicating a position of the first feature in a first coordinate frame of the portable electronic device and a position of the second feature in a second coordinate frame of the stored map (
Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0056 that not all of the determined correspondences are necessarily correct. In practice, some number of incorrect correspondences may be determined.
Schonberger teaches at Paragraph 0067 that the remote device may determine how many of the determined correspondences would be inliers and how many would be outliers and at Paragraph 0067 that in the case of RANSAC, this confidence value may be proportional to the inlier/outlier ratio described above, in which pose estimates that have more inliers (e.g., have more correspondences consistent with the pose than inconsistent with the pose) have relatively higher confidence values.
Schonberger teaches that the 3D map features are in the 3D word coordinate system and the 2D image features are in the 2D pixel-level image coordinate system.
Schonberger teaches at Paragraph 0054 that the 3D map may include suitable information describing the real-world environment and includes lines, planes and high-level objects, GPS coordinates and at Paragraph 0055 that estimating the pose of a camera device may involve determining correspondences between image features (first coordinate frame such as the image coordinate frame or the camera coordinate frame) detected at 2D pixel positions of images of a real-world environment and map features (second coordinate frame---the world coordinate frame) having 3D map positions in a digital environment map
Schonberger teaches at Paragraph 0043 that the “relative pose” therefore differs from the “estimated pose” or “absolute pose” output by the remote device, which defines the position and orientation of the camera device relative to a different frame of reference, e.g., a world-locked coordinate system corresponding to the real-world environment, or another suitable frame of reference); and
provide the plurality of pairs of matched features and respective quality metrics for aligning the first set of features at least partially to the second set of features (
providing the plurality of pairs of matched features:
Schongerger teaches at Paragraph [0059] that, the remote device determines whether there are additional image features to be matched with map features in the digital environment map. If yes, the remote device matches such image features at 414. If no, the remote device discontinues matching of image features at 420. Notably, discontinuing matching of image features at 420 does not represent an end to the overall image-based localization process, as any matched image features may be used to estimate a pose of the camera device. Furthermore, should additional images and/or image features be received from the camera device, the remote device may resume matching image features at 414.Schonberger teaches at Paragraph 0060 that once correspondences are identified between image features and map features, such correspondences may be used at 422 to estimate a pose of the camera device.
Providing respective quality metrics:
Schonberger teaches at Paragraph 0066 that the remote device may determine, for each candidate camera device pose, how many of the determined correspondences would be inliers and how many would be outliers, and the best overall camera device pose may be identified on this basis.
Schonberger teaches at Paragraph [0062] Continuing with the 2D line fitting example, a plurality of subsets of 2D data point pairs may be selected. In a RANSAC framework, these subsets are selected at random. If the minimal solver is applied to each subset (i.e., pair of 2D points), different potential solutions may be found from each subset. However, some potential solutions (2D lines) will be more consistent with the input data than other potential solutions. In other words, some potential lines will pass through more 2D points in the dataset than others. In RANSAC, for each potential solution identified by a minimal solver, the solution (i.e., 2D line) is compared to other 2D data points in the dataset to determine which points are consistent with the proposed solution (i.e., “inliers”) vs. which points are not consistent with the proposed solution (i.e., “outliers”). Any solution that has the highest inlier to outlier ratio, or has at least a threshold number of total inliers, may be accepted as the actual solution—e.g., 2D the line that best fits the set of 2D points.
Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0056 that not all of the determined correspondences are necessarily correct. In practice, some number of incorrect correspondences may be determined.
Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map. This may be done by searching the 3D map for the feature descriptors associated with each image feature and identifying the 3D map features having the most similar feature descriptors in the 3D point cloud. In other words, determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map. As a result, 2D points detected in images of the real-world environment correspond to 3D points associated with 3D map features, giving a set of 2D point to 3D point correspondences. This feature descriptor matching step can be implemented using one of many nearest neighbor matching techniques. The L2-distance between the descriptor vectors may be used to calculate the pairwise feature similarity.
Pair matching at the remote device.
Schonberger teaches at Paragraph 0061 that the remote device may identify a subset of correspondences from the overall set of determined 2D image feature to 3D map feature correspondences and at Paragraph 0067 that the remote device may determine how many of the determined correspondences would be inliers and how many would be outliers and at Paragraph 0067 that in the case of RANSAC, this confidence value may be proportional to the inlier/outlier ratio described above, in which pose estimates that have more inliers (e.g., have more correspondences consistent with the pose than inconsistent with the pose) have relatively higher confidence values).
For the above reasons, Schonberger implicitly teaches quality metrics while Garud explicitly teaches quality metric.
Garud teaches at Paragraph 0029 that 0029] At block 612, a set of candidate matches between the set of image feature points and the obtained map feature points are determined. For example, the matching may be performed using a cost function in conjunction with 2-way correspondence to determine a set of candidate matching feature points. In certain cases, candidate matches may be determined by matching an image feature descriptor of an image feature point against the obtained map feature descriptors to determine a 2D-3D matched feature point pair, matching a map feature descriptor of the candidate matched feature point against the image feature descriptors of the set of image feature points to determine a 3D-2D matched feature point pair, and determining a candidate match based on quality of the match and comparison between the 2D-3D matched feature point pairs and the 3D-2D matched feature point pairs. The quality of the match may refer to how well the image feature points and the obtained map feature points match. For example, where a cost function is used such as SAD, the output of the cost function reflects the quality of the match. At block 614, a pose of the camera may be determined from a set of potential poses of the camera from candidate matches from the set of candidate matches and an associated reprojection error estimated for the remaining points to select a first pose of the set of potential poses based on having a lowest associated error. For example, a random sample consensus (RANSAC) algorithm may be applied. In certain cases, repeatedly estimating the pose may include randomly selecting a subset of point pairs from the set of candidate matches, determining a camera pose based on the selected feature points, generating a 2D projection of the remaining matched map feature points based on the determined camera pose, determine the Euclidean distance i.e. reprojection error value associated with the generated 2D projections and matching image feature point locations to, repeating the steps of randomly selecting feature points, determining a camera pose, generating a 2D projection, and matching map feature points to generate a set of error values, and selecting the camera pose associated with the lowest error value.
It would have been obvious to one of the ordinary skill in the art to have incorporated Garud’s quality metric of the match and comparison between the 2D-3D matched feature point pairs of Schonberger to have determined the camera pose from the candidate matches from the set of candidate matches to have provided the lowest error value in estimating the camera pose. One of the ordinary skill in the art would have determined the candidate matches based on the quality of the match for the determination of the camera pose and the random selection of the subsets of the matching features from the candidate matches based on the quality of the matches, meaning that the matching pairs of the candidate matches are weighted 1 and the matching pairs of the non-candidate matches are weighted 0.
Barnes further teaches at Paragraph 0022 that 0022] Subsequently, a first subset (e.g., some, but not all) of matched feature pairs of the plurality of matched feature pairs are selected randomly, or based on some pre-established criteria (e.g., sparse objects that are relatively large are selected over sparse objects that are relatively small). For example, an appropriate RANdom SAmple Consensus algorithm, or “RANSAC” algorithm, is used for the selection. If a selected second matched feature pair includes a second feature from the primary image and a second feature from the auxiliary image, then the geometric transformations aim to align the first and second features from the primary image with the respective first and second features from the auxiliary image. Subsequently, a score is generated that indicates the quality of the transformed matches. The score can be a function of, for example, how evenly the selected matched feature pairs are distributed throughout the image, how close the selected matched feature pairs are to the region to be filed, and/or the geometric transformations performed. If the score is less than a threshold, the process is repeated. For example, if the score is less than the threshold, a second subset of matched feature pairs of the plurality of matched feature pairs are selected randomly, one or more geometric transformations are selected and performed, and another score is generated. This process continues until a termination condition has been satisfied.
Thus the selection is more weighted based on the matched feature pairs of relatively large sparse objects than the matched feature pairs of the relatively small sparse objects.
It would have been obvious to one of the ordinary skill in the art to have incorporated Barnes’s weighted selection of the matched feature pairs based on the quality metric into Schonberger’s selection of the matched feature pairs to have determined the camera pose from the candidate matches from the set of candidate matches to have provided the lowest error value in estimating the camera pose. One of the ordinary skill in the art would have determined the subset of the matching pairs based on the quality of the match for the determination of the camera pose and the random selection of the subsets of the matching features from the candidate matches based on the quality of the matches, meaning that the matching pairs of the candidate matches are weighted 1 and the matching pairs of the non-candidate matches are weighted 0.
Li teaches at Paragraph 0139 obtaining a conversion relationship between the coordinate system of the visual map and the coordinate system of the global grid map according to a relative installation position of the video camera device.
It would have been obvious to one of the ordinary skill in the art before the filing date of the instant application to have incorporated Li’s conversion of the coordinate system of the visual map and the coordinate system of the global map (world map or environment map) to have determined the correspondences of Schonberger of the matching pairs. One of the ordinary skill in the art would have been motivated to have compared the image features of the two different coordinate systems.
Re Claim 16:
Schonberger teaches a non-transitory computer-readable medium storing computer executable instructions configured to, when executed by at least one processor (
Schonberger teaches at Paragraph 0081 that logic subsystem 602 includes one or more physical devices configured to execute instructions. For example, the logic subsystem may be configured to execute instructions that are part of one or more applications, services, or other logical constructs. The logic subsystem may include one or more hardware processors configured to execute software instructions. Additionally, or alternatively, the logic subsystem may include one or more hardware or firmware devices configured to execute hardware or firmware instructions and at Paragraph 0082 that Storage subsystem 604 includes one or more physical devices configured to temporarily and/or permanently hold computer information such as data and instructions executable by the logic subsystem), perform a method of aligning a first set of features, derived from at least one image collected on a portable electronic device, at least partially to a second set of features in a stored map (Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0067 that the remote device may determine how many of the determined correspondences would be inliers and how many would be outliers and at Paragraph 0067 that in the case of RANSAC, this confidence value may be proportional to the inlier/outlier ratio described above, in which pose estimates that have more inliers (e.g., have more correspondences consistent with the pose than inconsistent with the pose) have relatively higher confidence values), the method comprising:
identifying a plurality of pairs of matched features from the first and second sets of features, each of the plurality of pairs of matched features comprising a first feature of the first set and a second feature of the second set (
Schonberger teaches at Paragraph 0058 correspondences are identified between image features and map features at 414 and at Paragraph 0021 that the image 110A includes a plurality of features 112 and at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0061 that the remote device may identify a subset of correspondences from the overall set of determined 2D image features to 3D map feature correspondences. Schonberger teaches at Paragraph 0040 that image features are detected by the camera device. Each image feature includes a descriptor 506, as well as a 2D pixel position 508 at which the image feature was detected in image 502. Each descriptor 506 will include some representation of the visual content of image 502 at which a corresponding image feature was detected—for instance, taking the form of a multi-dimensional feature vector);
for each of the plurality of pairs of matched features, computing a quality metric based on feature information indicating a position of the first feature in a first coordinate frame of the portable electronic device and a position of the second feature in a second coordinate frame of the stored map (Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0067 that the remote device may determine how many of the determined correspondences would be inliers and how many would be outliers and at Paragraph 0067 that in the case of RANSAC, this confidence value may be proportional to the inlier/outlier ratio described above, in which pose estimates that have more inliers (e.g., have more correspondences consistent with the pose than inconsistent with the pose) have relatively higher confidence values.
Schonberger teaches that the 3D map features are in the 3D word coordinate system and the 2D image features are in the 2D pixel-level image coordinate system.
Schonberger teaches at Paragraph 0054 that the 3D map may include suitable information describing the real-world environment and includes lines, planes and high-level objects, GPS coordinates and at Paragraph 0055 that estimating the pose of a camera device may involve determining correspondences between image features (first coordinate frame such as the image coordinate frame or the camera coordinate frame) detected at 2D pixel positions of images of a real-world environment and map features (second coordinate frame---the world coordinate frame) having 3D map positions in a digital environment map
Schonberger teaches at Paragraph 0043 that the “relative pose” therefore differs from the “estimated pose” or “absolute pose” output by the remote device, which defines the position and orientation of the camera device relative to a different frame of reference, e.g., a world-locked coordinate system corresponding to the real-world environment, or another suitable frame of reference); and
providing the plurality of pairs of matched features and respective quality metrics for aligning the first set of features at least partially to the second set of features (providing the plurality of pairs of matched features:
Schongerger teaches at Paragraph [0059] that, the remote device determines whether there are additional image features to be matched with map features in the digital environment map. If yes, the remote device matches such image features at 414. If no, the remote device discontinues matching of image features at 420. Notably, discontinuing matching of image features at 420 does not represent an end to the overall image-based localization process, as any matched image features may be used to estimate a pose of the camera device. Furthermore, should additional images and/or image features be received from the camera device, the remote device may resume matching image features at 414.Schonberger teaches at Paragraph 0060 that once correspondences are identified between image features and map features, such correspondences may be used at 422 to estimate a pose of the camera device.
Providing respective quality metrics:
Schonberger teaches at Paragraph 0066 that the remote device may determine, for each candidate camera device pose, how many of the determined correspondences would be inliers and how many would be outliers, and the best overall camera device pose may be identified on this basis.
Schonberger teaches at Paragraph [0062] Continuing with the 2D line fitting example, a plurality of subsets of 2D data point pairs may be selected. In a RANSAC framework, these subsets are selected at random. If the minimal solver is applied to each subset (i.e., pair of 2D points), different potential solutions may be found from each subset. However, some potential solutions (2D lines) will be more consistent with the input data than other potential solutions. In other words, some potential lines will pass through more 2D points in the dataset than others. In RANSAC, for each potential solution identified by a minimal solver, the solution (i.e., 2D line) is compared to other 2D data points in the dataset to determine which points are consistent with the proposed solution (i.e., “inliers”) vs. which points are not consistent with the proposed solution (i.e., “outliers”). Any solution that has the highest inlier to outlier ratio, or has at least a threshold number of total inliers, may be accepted as the actual solution—e.g., 2D the line that best fits the set of 2D points.
Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map and at Paragraph 0056 that not all of the determined correspondences are necessarily correct. In practice, some number of incorrect correspondences may be determined.
Schonberger teaches at Paragraph 0055 that determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map. This may be done by searching the 3D map for the feature descriptors associated with each image feature and identifying the 3D map features having the most similar feature descriptors in the 3D point cloud. In other words, determining the correspondences may include identifying a set of image features having feature descriptors that match feature descriptors of 3D map features in the 3D map. As a result, 2D points detected in images of the real-world environment correspond to 3D points associated with 3D map features, giving a set of 2D point to 3D point correspondences. This feature descriptor matching step can be implemented using one of many nearest neighbor matching techniques. The L2-distance between the descriptor vectors may be used to calculate the pairwise feature similarity.
Pair matching at the remote device.
Schonberger teaches at Paragraph 0061 that the remote device may identify a subset of correspondences from the overall set of determined 2D image feature to 3D map feature correspondences and at Paragraph 0067 that the remote device may determine how many of the determined correspondences would be inliers and how many would be outliers and at Paragraph 0067 that in the case of RANSAC, this confidence value may be proportional to the inlier/outlier ratio described above, in which pose estimates that have more inliers (e.g., have more correspondences consistent with the pose than inconsistent with the pose) have relatively higher confidence values).
For the above reasons, Schonberger implicitly teaches quality metrics while Garud explicitly teaches quality metric.
Garud teaches at Paragraph 0029 that 0029] At block 612, a set of candidate matches between the set of image feature points and the obtained map feature points are determined. For example, the matching may be performed using a cost function in conjunction with 2-way correspondence to determine a set of candidate matching feature points. In certain cases, candidate matches may be determined by matching an image feature descriptor of an image feature point against the obtained map feature descriptors to determine a 2D-3D matched feature point pair, matching a map feature descriptor of the candidate matched feature point against the image feature descriptors of the set of image feature points to determine a 3D-2D matched feature point pair, and determining a candidate match based on quality of the match and comparison between the 2D-3D matched feature point pairs and the 3D-2D matched feature point pairs. The quality of the match may refer to how well the image feature points and the obtained map feature points match. For example, where a cost function is used such as SAD, the output of the cost function reflects the quality of the match. At block 614, a pose of the camera may be determined from a set of potential poses of the camera from candidate matches from the set of candidate matches and an associated reprojection error estimated for the remaining points to select a first pose of the set of potential poses based on having a lowest associated error. For example, a random sample consensus (RANSAC) algorithm may be applied. In certain cases, repeatedly estimating the pose may include randomly selecting a subset of point pairs from the set of candidate matches, determining a camera pose based on the selected feature points, generating a 2D projection of the remaining matched map feature points based on the determined camera pose, determine the Euclidean distance i.e. reprojection error value associated with the generated 2D projections and matching image feature point locations to, repeating the steps of randomly selecting feature points, determining a camera pose, generating a 2D projection, and matching map feature points to generate a set of error values, and selecting the camera pose associated with the lowest error value.
It would have been obvious to one of the ordinary skill in the art to have incorporated Garud’s quality metric of the match and comparison between the 2D-3D matched feature point pairs of Schonberger to have determined the camera pose from the candidate matches from the set of candidate matches to have provided the lowest error value in estimating the camera pose. One of the ordinary skill in the art would have determined the candidate matches based on the quality of the match for the determination of the camera pose and the random selection of the subsets of the matching features from the candidate matches based on the quality of the matches, meaning that the matching pairs of the candidate matches are weighted 1 and the matching pairs of the non-candidate matches are weighted 0.
Barnes further teaches at Paragraph 0022 that 0022] Subsequently, a first subset (e.g., some, but not all) of matched feature pairs of the plurality of matched feature pairs are selected randomly, or based on some pre-established criteria (e.g., sparse objects that are relatively large are selected over sparse objects that are relatively small). For example, an appropriate RANdom SAmple Consensus algorithm, or “RANSAC” algorithm, is used for the selection. If a selected second matched feature pair includes a second feature from the primary image and a second feature from the auxiliary image, then the geometric transformations aim to align the first and second features from the primary image with the respective first and second features from the auxiliary image. Subsequently, a score is generated that indicates the quality of the transformed matches. The score can be a function of, for example, how evenly the selected matched feature pairs are distributed throughout the image, how close the selected matched feature pairs are to the region to be filed, and/or the geometric transformations performed. If the score is less than a threshold, the process is repeated. For example, if the score is less than the threshold, a second subset of matched feature pairs of the plurality of matched feature pairs are selected randomly, one or more geometric transformations are selected and performed, and another score is generated. This process continues until a termination condition has been satisfied.
Thus the selection is more weighted based on the matched feature pairs of relatively large sparse objects than the matched feature pairs of the relatively small sparse objects.
It would have been obvious to one of the ordinary skill in the art to have incorporated Barnes’s weighted selection of the matched feature pairs based on the quality metric into Schonberger’s selection of the matched feature pairs to have determined the camera pose from the candidate matches from the set of candidate matches to have provided the lowest error value in estimating the camera pose. One of the ordinary skill in the art would have determined the subset of the matching pairs based on the quality of the match for the determination of the camera pose and the random selection of the subsets of the matching features from the candidate matches based on the quality of the matches, meaning that the matching pairs of the candidate matches are weighted 1 and the matching pairs of the non-candidate matches are weighted 0.
Li teaches at Paragraph 0139 obtaining a conversion relationship between the coordinate system of the visual map and the coordinate system of the global grid map according to a relative installation position of the video camera device.
It would have been obvious to one of the ordinary skill in the art before the filing date of the instant application to have incorporated Li’s conversion of the coordinate system of the visual map and the coordinate system of the global map (world map or environment map) to have determined the correspondences of Schonberger of the matching pairs. One of the ordinary skill in the art would have been motivated to have compared the image features of the two different coordinate systems.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIN CHENG WANG whose telephone number is (571)272-7665. The examiner can normally be reached Mon-Fri 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Poon can be reached at 571-270-0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JIN CHENG WANG/Primary Examiner, Art Unit 2617