DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant's claim for foreign priority based on an application filed in France on 12/20/2023. It is noted, however, that applicant has not filed a certified copy of the FR2314650 application as required by 37 CFR 1.55. An attempt by the Office to electronically retrieve, under the priority document exchange program, the foreign application 2314650 to which priority is claimed has FAILED on 05/20/2025.
Receipt is acknowledged of certified copies of papers of application EP24305954 required by 37 CFR 1.55.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-6, 10, 12-15, and 48-50 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US 20220148219 A1), and in view of Streem et al. (US 20230400327 A1).
Regarding Claim 1, Kim discloses A computer-implemented method for reconstructing a scene in three dimensions from a plurality of images of one or more viewpoints of the scene acquired using an imaging device (ABST reciting “A visual localization method includes generating a first feature point map by using first map data calculated on the basis of a first viewpoint; generating a second feature point map by using second map data calculated on the basis of a second viewpoint different from the first viewpoint; constructing map data for localization having the first and second feature point maps integrated with each other, by compensating for a position difference between a point of the first feature point map and a point of the second feature point map;” ¶26 reciting “images captured by the mobile device or the autonomous driving device”), comprising:
receiving the plurality of images without receiving extrinsic or intrinsic properties of the imaging device; and
processing the plurality of images to produce a plurality of pointmaps of the scene that correspond to the plurality of images and that are aligned in a common coordinate frame,
(¶26 reciting “to perform visual localization by using the map data for localization and images captured by the mobile device or the autonomous driving device, wherein the map data for localization is provided with a first feature point map and a second feature point map, wherein the first feature point map is generated by using first map data calculated based on a first viewpoint, wherein the second feature point map is generated by using second map data calculated based on a second viewpoint different from the first viewpoint, and wherein the first and second feature point maps are matched with each other by using a difference of a camera pose therebetween.” Further, ¶127 disclosing the first and the second point maps are aligned in a common coordinate frame, and reciting “information such as a camera pose may be consistent with each other on the same coordinates of the first and second feature point maps.” Furthermore, ¶145 reciting “a camera position and a camera pose may be estimated”, therefore the input image does not have properties of the camera.)
wherein each pointmap is a one-to-one mapping between pixels of one of the plurality of images and three-dimensional points of the scene. (¶148 reciting “The operation by the PnP solver is performed to obtain correspondences between 2D pixel coordinates and 3D points on the map through local feature matching”. In addition, ¶118 and ¶145 disclosing the 2D input image.)
Kim discloses that a neural network is used to estimate a dense depth image. However, Kim does not explicitly disclose processing the plurality of images to produce a plurality of pointmaps of the scene using a neural network.
Streem teaches “ a localization processing service for enabling localization of a navigation network-restricted subsystem, as well as to an observed scene reconstruction service.” (¶3). More specifically, Fig. 7 shows a system process 701; and ¶108 recites “In the localization process, all camera streamings (e.g., all four camera streamings 704), and an inertial navigation system (“INS”) orientation 714, which may be embedded in IMU hardware, to create a 360° image 706.” Further, ¶135 recites “The use of one or more suitable models or engines or neural networks or the like (e.g., model 220, 224, 290, etc.) may enable estimation or any suitable determination of a localization of a mobile subsystem in an environment.”
It would have been obvious to one with ordinary skill, before the effective filing date of the claimed invention, to modify the method (taught by Kim) to use a neural network to do the lactonization and scene reconstruction (taught by Streem). The suggestions/motivations would have been that “ Such models (e.g., neural networks) running on any suitable processing units (e.g., graphical processing units (“GPUs”) that may be available to system 1) provide significant speed improvements in efficiency and accuracy with respect to estimation over other types of algorithms and human-conducted analysis of data, as such models can provide estimates in a few milliseconds or less, thereby improving the functionality of any computing device on which they may be run. ” (¶135), and to apply a known technique to a known device (method, or product) ready for improvement to yield predictable results.
Regarding Claim 2, Kim in view of Streem discloses The computer-implemented method according to claim 1, wherein the processing the plurality of images using the neural network further includes:
processing the plurality of images using the neural network to produce a plurality of local feature maps that correspond to each of the plurality of images. (Kim, ABST reciting “generating a first feature point map by using first map data calculated on the basis of a first viewpoint; generating a second feature point map by using second map data”. ¶76 reciting “one feature point map for visual localization is generated by using the street view image”; and ¶77 reciting “The feature point map is a map having data on a 3D feature point, which may be referred to as a feature map”)
Regarding Claim 3, Kim in view of Streem discloses The computer-implemented method according to claim 2, further comprising performing one of the following applications using the plurality of pointmaps of the scene:
rendering a pointcloud of the scene for a given camera pose;
recovering camera parameters of the scene;
recovering depth maps of the scene for a given camera pose; and
recovering three dimensional meshes of the scene.
(Kim, ¶145 reciting “a camera position and a camera pose are estimated through local feature matching.” Further, ¶147 reciting “More specifically, the deepfeature serving server 160 extracts a local feature point, and a 3D value of the first feature point map matching with the feature point of the 2D image is detected through 2D-3D matching, by using a local feature descriptor. Then, the PnP solver performs an operation based on the 3D value and a pixel value of the 2D input image, thereby estimating a camera pose.”)
Regarding Claim 4, Kim in view of Streem discloses The computer-implemented method according to claim 3, further comprising performing visual localization in the scene using the recovered camera parameters. (Kim, ¶12 reciting “a visual localization method and system capable of generating map data for localization by matching a first feature point map and a second feature point map generated from data at different viewpoints with each other, by using a difference of a camera pose.”)
Regarding Claim 5, Kim in view of Streem discloses The computer-implemented method according to claim 1, wherein the processing the plurality of images using the neural network further includes:
processing a plurality of image subsets of the plurality of images using the neural network, wherein each image subset of the plurality of image subsets includes different ones of the plurality of images; and
aligning pointmaps from the plurality of image subsets into the plurality of pointmaps that are aligned in the common coordinate frame using a global aligner that performs regression based alignment.
(Kim, ABST reciting “constructing map data for localization having the first and second feature point maps integrated with each other, by compensating for a position difference between a point of the first feature point map and a point of the second feature point map”. Further, ¶12 reciting “generating map data for localization by matching a first feature point map and a second feature point map generated from data at different viewpoints with each other, by using a difference of a camera pose.” Furthermore, ¶27 reciting “constructing map data for localization having the first and second feature point maps integrated with each other, by compensating for a position difference between a point of the first feature point map and a point of the second feature point map; and performing visual localization by using the map data for localization.” In addition, ¶127 reciting “Through the aforementioned processes, information such as a camera pose may be consistent with each other on the same coordinates of the first and second feature point maps.” Streem, Fig. 7 shows four camera streamings 704, and Kim discloses integrating 2 input images from different viewpoints. It would have been obvious to one with ordinary skill all the input images can be processed by subsets. The suggestions/motivations would have been the same as that of Claim 1 rejections.)
Regarding Claim 6, Kim in view of Streem discloses The computer-implemented method according to claim 2, wherein the processing the plurality of images using the neural network further includes:
processing a plurality of image subsets of the plurality of images using the neural network, wherein each image subset of the plurality of image subsets includes different ones of the plurality of images; and
aligning pointmaps from the plurality of image subsets into the plurality of pointmaps that are aligned in the common coordinate frame using an alignment module that performs pixel correspondence based alignment using the plurality of local feature maps.
(See Claims 1, 2 and 5 rejections for detailed analysis.)
Regrading Claim 10, Kim in view of Streem discloses The computer-implemented method according to claim 1, wherein each pointmap represents a two-dimensional field of three-dimensional points of the scene (Kim, ¶148 reciting “The operation by the PnP solver is performed to obtain correspondences between 2D pixel coordinates and 3D points on the map through local feature matching”. In addition, ¶118 and ¶145 disclosing the 2D input image.), and wherein the processing the plurality of images using the neural network further includes generating a confidence score map for each pointmap. (Kim, ¶194 reciting “For optimization of the 3D point, the nodes may be set as the pose of the street view images and the position of the 3D point, and the edges may be set as a plurality of errors related to the nodes.”, where errors read on confidence scores.)
Regarding Claim 12, Kim in view of Streem discloses The computer-implemented method according to claim 1, further comprising, based on the plurality of pointmaps of the scene, determining extrinsic or intrinsic parameters of the imaging device. (Kim, ¶145 reciting “a camera position and a camera pose are estimated through local feature matching.” Further, ¶147 reciting “More specifically, the deepfeature serving server 160 extracts a local feature point, and a 3D value of the first feature point map matching with the feature point of the 2D image is detected through 2D-3D matching, by using a local feature descriptor. Then, the PnP solver performs an operation based on the 3D value and a pixel value of the 2D input image, thereby estimating a camera pose.”)
Regarding Claim 13, Kim in view of Streem discloses The computer-implemented method according to claim 12, wherein at least one of: the extrinsic parameters of the imaging device include rotation and translation of the imaging device; and the intrinsic parameters of the imaging device include skew and focal length. (Kim, ¶145 reciting “a camera position and a camera pose are estimated through local feature matching.”)
Regarding Claim 14, Kim in view of Streem discloses The computer-implemented method according to claim 1, wherein the neural network performs cross attention between views of the scene. (Kim, ¶183 reciting “a transform of homography between two images is calculated by putative matching for matching entire feature points of two images one by one. ”)
Regarding Claim 15, Kim in view of Streem discloses The computer program product comprising code instructions which, when the program is executed by a computer, cause the computer to carry out the method according to claim 1. (Kim. ¶27 reciting “ a computer-readable medium stores computer-executable program instructions that, when executed by a processor, cause the processor to perform operations ”)
Regarding Claim 48. Kim in view of Streem discloses A system for reconstructing a scene in three dimensions from a plurality of images of one or more viewpoints of the scene acquired using an imaging device, comprising:
one or more processors; and
memory including code that, when executed by the one or more processors, perform to:
receive the plurality of images without receiving extrinsic or intrinsic properties of the imaging device; and process the plurality of images using a neural network to produce a plurality of pointmaps of the scene that correspond to the plurality of images and that are aligned in a common coordinate frame, wherein each pointmap is a one-to-one mapping between pixels of one of the plurality of images and three-dimensional points of the scene.
(See Claim 1 and Claim 15 rejections.)
Regarding Claim 49. Kim in view of Streem discloses The system according to claim 48, wherein the code that, when executed by the one or more processors, further perform to: process the plurality of images using the neural network to produce a plurality of local feature maps that correspond to each of the plurality of images. (See Claim 48 and Claim 2 rejections)
Regarding Claim 50. Kim in view of Streem discloses The system according to claim 49, wherein the code that, when executed by the one or more processors, further perform one of the following applications using the plurality of pointmaps of the scene to:
render a pointcloud of the scene for a given camera pose;
(ii) recover camera parameters of the scene;
(iii) recover depth maps of the scene for a given camera pose; and
(iv) recover three dimensional meshes of the scene.
(See Claim 48 and Claim 3 rejections)
Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kim et al. (US 20220148219 A1), and in view of Streem et al. (US 20230400327 A1), and further in view of Tsintotas et al. "The revisiting problem in simultaneous localization and mapping: A survey on visual loop closure detection." IEEE Transactions on Intelligent Transportation Systems 23.11 (2022): 19929-19953.
Regarding Claim 11, Kim in view of Streem discloses The computer-implemented method according to claim 1.
However, Kim in view of Streem does not explicitly disclose wherein processing the plurality of images further includes expressing the plurality of images using a co-visibility graph and processing the co-visibility graph using the neural network to produce the plurality of pointmaps of the scene that correspond to the plurality of images and that are aligned in the common coordinate frame.
Co-visibility graphs are well known in the art. In addition, Tsintotas teaches “co-visibility graphs, generated from learned features, could boost the invariance to viewpoint changes” (p. 19943, left column. Lines 15-17).
It would have been obvious to one with ordinary skill, before the effective filing date of the claimed invention, to modify the method (taught by Kim in view of Streem) to adapt the co-visibility graph method (taught by Tsintotas). The suggestions/motivations would have been to boost the invariance to view point changes, and to apply a known technique to a known device (method, or product) ready for improvement to yield predictable results.
Allowable Subject Matter
Claims 20-46 and 47 are allowed.
The following is an examiner’s statement of reasons for allowance:
Independent Claim 20 is distinguished from the closest known prior art alone or in reasonable combination, in consideration of the claim as a whole, particularly the limitations similar to “encoding the images into features corresponding to the images, respectively; determining similarities between pairs of the images based on the features, the similarity of each pair being determined based on the features of the images of that pair; filtering out ones of the pairs of the images based on the similarities; generating a graph of the scene based on the pairs not filtered out; and determining pointmaps of the scene that correspond to ones of the images of the pairs not filtered out and that are aligned in a common coordinate frame” in combination with the remaining aspects of the claim. Independent Claim 47 is similar in scope to Claim 20, and therefore also contains allowable subject matter. Claims 21-46 depend from claim 20, and therefore also contain allowable subject matter.
The closest prior art (Kim, Streem) teaches A computer-implemented method for reconstructing a scene in three dimensions from a plurality of images of one or more viewpoints of the scene acquired using one or more imaging devices, the method comprising: receiving the plurality of images without receiving extrinsic or intrinsic properties of the one or more imaging devices; and determining pointmaps of the scene and that are aligned in a common coordinate frame, wherein each of the pointmaps is a one-to-one mapping between pixels of one of the plurality of images and three-dimensional points of the scene. (See Claim 1 rejections) However, the closest prior fails to teach each and every limitations for Claim 20.
Claims 7-8 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding Claim 7, closest Kim in view of Streem discloses The computer-implemented method according to claim 1.
However, the closest art Kim in view of Streem and Tsintotas does not teaches wherein the processing the plurality of images using the neural network further includes:
(a) for each of the plurality of images:
(i) generating patches with a pre-encoder;
(ii) encoding the generated patches with a transformer encoder to define token encodings that represent the generated patches; and
(iii) decoding the token encodings with a transformer decoder to generate token decodings;
(b1) for one of the token decodings, generating a pointmap corresponding to one of the plurality of images with a first regression head that produces pointmaps in a coordinate frame of the one of the plurality of images; and
(c1) for each of other of the token decodings, generating a pointmap corresponding to each of the other of the plurality of images with a second regression head that produces pointmaps in the coordinate frame of the one of the plurality of images.
Claim 8 depends from Claim 7 and therefore also contain allowable subject matter.
Claim 9 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding Claim 9, closest art Kim in view of Streem discloses The computer-implemented method according to claim 1.
However, the closest art Kim in view of Streem does not explicitly teach wherein the processing the plurality of images using the neural network further includes: (a) initializing the plurality of pointmaps; (b) for each of the plurality of images: (i) generating image patches with a pre-encoder; and (ii) encoding the generated image patches with a transformer encoder to define image token encodings that represent the generated image patches; (c) for each of the plurality of pointmaps: (i) generating pointmap patches with a pre-encoder; and (ii) encoding the generated pointmap patches with a transformer encoder to define pointmap token encodings that represent the pointmap generated patches; (d) for each of the plurality of image token encodings and corresponding of the plurality of pointmap token encodings, aggregating each pair with a mixer to generate mixed token encodings; (e) for each of the generated mixed token encodings, decoding the mixed token encodings with a transformer decoder to generate mixed token decodings; (f) for each of the mixed token decodings, replacing the plurality of pointmaps corresponding to the plurality of images with pointmaps generated by a regression head that produces pointmaps in a coordinate frame that is common to the plurality of images; and (g) repeating (c)-(f) for a predetermined number of iterations.
Claims 16-19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding Claim 16, closest art Kim in view of Streem discloses The computer-implemented method of claim 1.
However, the closest art Kim in view of Streem fails to teach wherein the processing includes:
encoding the images into features corresponding to the images, respectively;
determining similarities between pairs of the images based on the features, the similarity of each pair being determined based on the features of the images of that pair;
generating a graph of the scene based on ones of the pairs; and
determining the pointmaps of the scene that correspond to ones of the images of ones of the pairs and that are aligned in a common coordinate frame.
Claims 17-19 are directly or indirectly dependent from claim 16 and they also contain allowable subject matter.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YI WANG whose telephone number is (571)272-6022. The examiner can normally be reached 9am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at (571)272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YI WANG/Primary Examiner, Art Unit 2619