DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) because the claim limitations use a generic placeholder “unit” that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: an image frame acquisition unit, an image frame processing unit, and a 3D point cloud map creation unit in claim 11.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f), they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f), applicant may: (1) amend the claim limitations to avoid them being interpreted under 35 U.S.C. 112(f) (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid them being interpreted under 35 U.S.C. 112(f).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-7, and 9-14 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 2020/0394410; hereinafter “Zhang”) in view of Wrenninge (US 2022/0237410).
Regarding claim 1, Zhang discloses A method for creating a 3D point cloud map (“create a descriptive point cloud from the results of 3D feature-based tracking,” para. 208), comprising: acquiring image frames from an image frame sequence of an external environment captured by a target camera (“the camera that captured the image data,” para. 208; “two image frames, taken at times t and t+Δt,” para. 139) of a vehicle (“autonomous vehicle guidance technology,” para. 64); processing an image frame in the image frame sequence using a feature point recognition model to identify a set of feature points in the image frame (“performs feature detection upon image frames,” para. 94), along with a corresponding set of descriptors for the set of feature points (“image preprocessing, feature detection, and feature descriptor preparation,” para. 134), wherein the feature point recognition model is a neural network model (“the feature detection engine and the feature description engine can be combined, as demonstrated in Convolutional Neural Networks (CNNs),” para. 366) trained using a plurality of sample images (“After training for days or even weeks on a large data set, a CNN can be capable of in real-time perception of key features in images,” para. 366) under different lighting intensities (“A labeled data set is required to represent all possible driving environments … night, day,” para. 367); and creating the 3D point cloud map of the external environment based on the set of descriptors (“create a descriptive point cloud from the results of 3D feature-based tracking,” para. 208).
Zhang does not specifically recite training the neural network using a plurality of sample images of a same scene under different lighting intensities.
In the same art of training a neural network for feature detection in autonomous vehicles, Wrenninge teaches training a neural network using a plurality of sample images of a same scene under different lighting intensities (“generating and using synthetic images to train a machine learning model … a distribution of environmental attributes may be selected for a training purpose, such as selecting variation to create variations in … illumination for the same 3D scene,” para. 114; “The synthetic image dataset may be used for applications such as training a … neural network model … training or evaluating a model used to control an autonomous vehicle,” para. 115).
Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Wrenninge to Zhang. The motivation would have been “to optimize model convergence and/or accuracy” (Wrenninge, para. 18).
Regarding claim 3, the combination of Zhang and Wrenninge renders obvious wherein the image frame is a first image frame, the set of feature points is a set of first feature points, and the set of descriptors is a set of first descriptors, the method further comprising: processing a second image frame in the image frame sequence using the feature point recognition model to determine a set of second feature points in the second image frame and a set of second descriptors corresponding to the set of second feature points (“calculate the motion between two image frames, taken at times t and t+Δt,” Zhang, para. 139; “The first image frame is selected as a keyrig, and the device coordinate frame at that timestamp become the coordinates of the descriptive point cloud … triangulate the new features from images captured from the device at this timestamp to obtain the points in the device's coordinate frame,” Zhang, paras. 214-215).
Regarding claim 4, the combination of Zhang and Wrenninge renders obvious determining a set of matching feature points between the first image frame and the second image frame based on the set of first descriptors and the set of second descriptors (“Feature description can be used to compare and match feature points between different images,” Zhang, para. 364); and creating the 3D point cloud map based on the set of matching feature points ("geometric information obtained by a 3D feature-based tracking process is represented using a descriptive point cloud representation,” Zhang, para. 202).
Regarding claim 5, the combination of Zhang and Wrenninge renders obvious determining a first pose of the vehicle corresponding to the first image frame and a second pose of the vehicle corresponding to the second image frame (“The pose information includes a location of the mobile device and view of the camera that captured the image data,” Zhang, para. 208); and creating the 3D point cloud map based on the first pose, the second pose, and the set of matching feature points (“a descriptive point cloud from the 3D feature-based tracking is refined by obtaining an ‘optimal’ (i) pose of each keyrig,” Zhang, para. 217).
Regarding claim 6, the combination of Zhang and Wrenninge renders obvious determining the first pose and the second pose based on at least one of (i) data from an odometer of the vehicle, or (ii) data from an inertial measurement unit of the vehicle (“extrapolate an initial pose using inertial data from the multi-axis IMU,” Zhang, para. 166).
Regarding claim 7, the combination of Zhang and Wrenninge renders obvious based on the first pose and the second pose, determining a position of map points corresponding to the matching feature points in the set of matching feature points; and generating the 3D point cloud map by optimizing the positions of the map points (“a descriptive point cloud from the 3D feature-based tracking is refined by obtaining an ‘optimal’ (i) pose of each keyrig and (ii) 3D locations of all the points,” Zhang, para. 217; “Optical flow techniques are used to determine 2D correspondences in the images, enabling matching together features in different images … The 2D features points from the first 360-degrees image and the 360-degrees second image will be triangulated to determine the depth of the feature point,” Zhang, para. 457).
Regarding claim 9, the combination of Zhang and Wrenninge renders obvious acquiring a target image frame captured by the target camera after generating the 3D point cloud map (“creates, updates, and refines descriptive point cloud using feature descriptors,” para. 260); determining a set of target feature points in the target image frame and a corresponding set of target descriptors for the set of target feature points; based on the set of target descriptors, determining a set of matching target feature points between the target image frame and a third image frame in the image frame sequence (“The first image frame is selected as a keyrig, and the device coordinate frame at that timestamp become the coordinates of the descriptive point cloud … triangulate the new features from images captured from the device at this timestamp to obtain the points in the device's coordinate frame,” Zhang, paras. 214-215; and determining a position of the vehicle in the 3D point cloud map based on the position of a set of map points corresponding to the set of matching target feature points (“trajectory of a device between two consecutive keyrigs (i.e. from keyrig k to keyrig k+1) in a descriptive point cloud is estimated … Initialize the image frame that creates keyrig k to be at its pose stored in the descriptive point cloud … Use the ‘3D feature-based tracking process’ as described herein under section heading ‘Tracking’ to track the image frames between the two keyrigs,” Zhang, paras. 236-240).
Regarding claim 10, the combination of Zhang and Wrenninge renders obvious wherein: the vehicle has a plurality of cameras, the plurality of cameras include the target camera, and the plurality of cameras capture a plurality of image frame sequences directed at the external environment (“Cameras 208, 210 [of Fig. 2] include at least partially overlapping fields of view,” Zhang, para. 73; “the multi-ocular sensor includes four interfaces to couple with four cameras,” Zhang, para. 332); and creating the 3D point cloud map of the external environment based on the set of descriptors comprises: creating the 3D point cloud map of the external environment based on the plurality of sets of descriptors corresponding to a plurality of image frames in each image frame of a plurality of image frame sequences (“images by the quadocular system can be utilized to build a 3D map. Therefore, the pair of 360-degrees frames can be compared by extracting and matching key features in the frames,” Zhang, paras. 332-333).
Regarding claim 11, it is rejected using the same citations and rationales described in the rejection of claim 1.
Regarding claim 12, the combination of Zhang and Wrenninge renders obvious A computing device, comprising: at least one processor; and a non-transitory memory, coupled to the at least one processor, and having instructions stored thereon that, when the instructions are executed by the at least one processor, cause the computing device to perform the method according to claim 1 (“a non-transitory computer readable storage medium storing instructions executable by a processor to perform any of the methods described above,” Zhang, para. 290).
Regarding claim 13, the combination of Zhang and Wrenninge renders obvious A vehicle, comprising: a camera; and a computing device according to claim 12 (“the technology disclosed can be applied to autonomous vehicle guidance technology,” Zhang, para. 64; “the device is mounted on a vehicle,” Zhang, para. 218; Fig. 2 of Zhang illustrates cameras).
Regarding claim 14, the combination of Zhang and Wrenninge renders obvious A non-transitory machine-readable storage medium storing machine-executable instructions, wherein the machine-executable instructions are executed by a processor to implement the method according to claim 1 (“a non-transitory computer readable storage medium storing instructions executable by a processor to perform any of the methods described above,” Zhang, para. 290).
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Zhang and Wrenninge, and further in view of Pugh et al. (US 2021/0142497; hereinafter “Pugh”).
Regarding claim 2, the combination of Zhang and Wrenninge renders obvious wherein each descriptor in the set of descriptors is determined based on an entire image frame (“performs feature detection upon image frames,” para. 94; “feature extractor analyzes images (e.g., image frames captured via cameras 208, 210 [of Fig. 2]),” para. 122; “new features are extracted from the new frame,” para. 298).
The combination of Zhang and Wrenninge does not disclose that a feature descriptor indicates a color gradient of a corresponding feature point in the set of feature points.
In the same art of extracting features from images, Pugh teaches a descriptor that indicates a color gradient of a corresponding feature point in the set of feature points (“determining a photogrammetry point cloud from the image (e.g., using SLAM, SFM, MVS, depth sensors, etc.),” para. 21; “estimating visual information from each image … to determine features,” para. 45; “The visual information can include two-dimensional features, three-dimensional features, or additionally or alternatively neural network features,” para. 47; “Two-dimensional features that can be extracted can include … color gradients,” para. 48).
Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Pugh to the combination of Zhang and Wrenninge. The motivation would have been for “generating a dense, scaled, accurate point cloud” (Pugh, para. 21).
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over the combination of Zhang and Wrenninge, and further in view of He et al. (US 2019/0318502; hereinafter “He”).
Regarding claim 8, the combination of Zhang and Wrenninge does not disclose determining the set of matching feature points based on a distance between the descriptors in the set of first descriptors and the descriptors in the set of second descriptors; or determining the set of matching feature points by applying a neural network model used for feature point matching to the set of first descriptors and the set of second descriptors.
In the same art of matching feature descriptors, He teaches determining the set of matching feature points by applying a neural network model used for feature point matching to the set of first descriptors and the set of second descriptors (“feature descriptor matching using the trained feature descriptor neural network and the trained feature descriptor matching model,” para. 67).
Before the effective filing date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teaching of He to the combination of Zhang and Wrenninge. The motivation would have been to “provide quick and efficient matching” (He, para. 30).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ryan McCulley whose telephone number is (571)270-3754. The examiner can normally be reached Monday through Friday, 8:00am - 4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571) 272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RYAN MCCULLEY/Primary Examiner, Art Unit 2611