Prosecution Insights
Last updated: October 04, 2026
Application No. 19/200,000

Machine-Learned Monocular Depth Estimation and Semantic Segmentation for 6-DOF Absolute Localization of a Delivery Drone

Non-Final OA §103§112§DOUBLEPATENT
Filed
May 06, 2025
Priority
Jul 18, 2022 — continuation of 12/307,710
Examiner
MOLINA, NIKKI MARIE M
Art Unit
Tech Center
Assignee
Wing Aviation LLC
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 2m
Est. Remaining
82%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
87 granted / 111 resolved
+18.4% vs TC avg
Minimal +4% lift
Without
With
+4.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
17 currently pending
Career history
144
Total Applications
across all art units

Statute-Specific Performance

§101
14.2%
-25.8% vs TC avg
§103
44.6%
+4.6% vs TC avg
§102
13.6%
-26.4% vs TC avg
§112
26.6%
-13.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 111 resolved cases

Office Action

§103 §112 §DOUBLEPATENT
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This is a Non-final Office Action on the merits. Claims 1-20 are currently pending and are addressed below. Information Disclosure Statement The information disclosure statement(s) (IDS) submitted on 05/06/2025 and 04/13/2026 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement(s) is/are being considered by the examiner. Specification The disclosure is objected to because of the following informalities: Lines 3 & 5-6 of [0027] recites “semantically labeled 3D point” in which the underlined portion appears to be a possible typographical error, since line 2 recites “semantically labeled three-dimensional (3D) point cloud”. Lines 6 & 10-11 of [0027] recites “absolution position”, in which the underlined portion appears to be a typographical error. Line 2 of [0065] recites “in either reaching the specific target delivery location”, in which the underlined portion appears to be grammatically incorrect. Line 6 of [0107] recites “trained machine learning model(s) 432 can be trained, reside, and execute to provide inferences”, which is unclear, and the underlined portion appears to be grammatically incorrect. Appropriate correction is required. Claim Objections Claims 1 and 20 objected to because of the following informalities: Claim 1 (and claim 20 by reciting analogous limitations) recites “…a unmanned aerial vehicle…”, which appears to be grammatically incorrect. Appropriate correction is required. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 1-20 rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-19 of U.S. Patent No. 12307710 B2 (hereinafter “Patent ‘710”). Although the claims at issue are not identical, they are not patentably distinct from each other because representative claims 1-18 (and claims 19-20 by reciting analogous limitations) of the instant application are encompassed by the subject matter of claims representative claims 1-19 of Patent ‘710, as illustrated in the tables below, where differences in the claim sets are bolded. Present Application 19/200,000 1,19-20 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 U.S. Patent No. 12307710 B2 1,18-19 2 3 4 5 6 7 8 1 9 10 11 12 13 14 15 18 17 Present Application 19/200,000 Claims 1-18 U.S. Patent No. 12307710 B2 Claims 1-17 1. A method comprising: receiving an image captured by a camera on a unmanned aerial vehicle (UAV) and representative of an environment of the UAV; 1. A method comprising: receiving a two-dimensional (2D) image captured by a camera on a unmanned aerial vehicle (UAV) and representative of an environment of the UAV; determining, based on the image, a semantic image of the environment and a depth image of the environment; applying a trained machine learning model to the 2D image to produce a semantic image of the environment and a depth image of the environment, wherein the machine learning model has been trained with a semantics branch to produce the semantic image and a depth branch to produce the depth image, and wherein the semantic image comprises one or more semantic labels; retrieving reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels; retrieving reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels; and determining a location of the UAV in the environment, wherein determining the location of the UAV in the environment comprises aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data; and determining a location of the UAV in the environment, wherein determining the location of the UAV in the environment comprises aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating the one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data; and controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment. controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment. 2. The method of Claim 1, further comprising: controlling the UAV to navigate in the environment using a Global Navigation Satellite System (GNSS) system; detecting a disruption in service from the GNSS system, wherein the location of the UAV in the environment is determined responsive to detecting the disruption in service from the GNSS system; and subsequent to detecting the disruption in service from the GNSS system, controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment. 2. The method of claim 1, further comprising: controlling the UAV to navigate in the environment using a Global Navigation Satellite System (GNSS) system; detecting a disruption in service from the GNSS system, wherein the location of the UAV in the environment is determined responsive to detecting the disruption in service from the GNSS system; and subsequent to detecting the disruption in service from the GNSS system, controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment. 3. The method of Claim 1, further comprising: controlling the UAV to navigate in the environment using a GNSS system; and using the determined location of the UAV in the environment to cross-check location data from the GNSS system. 3. The method of claim 1, further comprising: controlling the UAV to navigate in the environment using a GNSS system; and using the determined location of the UAV in the environment to cross-check location data from the GNSS system. 4. The method of Claim 1, further comprising: determining a GNSS location of the UAV in the environment using a GNSS system; determining a refined location of the UAV in the environment based on the GNSS location of the UAV in the environment and the determined location of the UAV in the environment; and controlling the UAV to navigate in the environment based on the refined location of the UAV in the environment. 4. The method of claim 1, further comprising: determining a GNSS location of the UAV in the environment using a GNSS system; determining a refined location of the UAV in the environment based on the GNSS location of the UAV in the environment and the determined location of the UAV in the environment; and controlling the UAV to navigate in the environment based on the refined location of the UAV in the environment. 5. The method of Claim 1, wherein the aligning the depth image of the environment with the reference depth data representative of the environment to determine the location of the UAV in the environment comprises using an iterative closest point (ICP) algorithm. 5. The method of claim 1, wherein the aligning the depth image of the environment with the reference depth data representative of the environment to determine the location of the UAV in the environment comprises using an iterative closest point (ICP) algorithm. 6. The method of Claim 5, wherein the ICP algorithm aligns points from the reference depth data with points from the depth image such that the reference semantic labels from the reference depth data correspond to the one or more semantic labels from the semantic image. 6. The method of claim 5, wherein the ICP algorithm aligns points from the reference depth data with points from the depth image such that the reference semantic labels from the reference depth data correspond to the one or more semantic labels from the semantic image. 7. The method of Claim 1, wherein the semantic image of the environment and the depth image of the environment have the same dimensions. 7. The method of claim 1, wherein the semantic image of the environment and the depth image of the environment produced by the machine learning model have the same dimensions. 8. The method of Claim 1, wherein the camera on the UAV faces downward, and wherein the image captured by the camera is representative of a terrain in the environment below the UAV. 8. The method of claim 1, wherein the camera on the UAV faces downward, and wherein the 2D image captured by the camera is representative of a terrain in the environment below the UAV. 9. The method of claim 1, wherein the semantic image and the depth are determined using a machine learning model. 1. A method comprising… …applying a trained machine learning model to the 2D image to produce a semantic image of the environment and a depth image of the environment, wherein the machine learning model has been trained with a semantics branch to produce the semantic image and a depth branch to produce the depth image… 10. The method of Claim 9, wherein a semantics branch and a depth branch of the machine learning model operate on a commonly generated feature set. 9. The method of claim 1, wherein the semantics branch and the depth branch of the machine learning model operate on a commonly generated feature set. 11. The method of Claim 9, wherein the machine learning model has been trained based on ground truth depth data, wherein the ground truth depth data is based on performance of a structure from motion (SfM) algorithm on images captured by one or more UAVs. 10. The method of claim 1, wherein the machine learning model has been trained based on ground truth depth data, wherein the ground truth depth data is based on performance of a structure from motion (SfM) algorithm on images captured by one or more UAVs. 12. The method of Claim 9, wherein the machine learning model has been trained based on ground truth semantic data, wherein the ground truth semantic data is based on operator labeling of images captured by one or more UAVs. 11. The method of claim 1, wherein the machine learning model has been trained based on ground truth semantic data, wherein the ground truth semantic data is based on operator labeling of images captured by one or more UAVs. 13. The method of Claim 9, wherein the machine learning model has been trained using a scale invariant loss. 12. The method of claim 1, wherein the machine learning model has been trained using a scale invariant loss for training of the depth branch. 14. The method of Claim 13, further comprising applying a scale factor to the depth image, wherein the scale factor comprises a ratio of an altitude of the UAV above ground level relative to a median of a monocular depth map, wherein the monocular depth map is based on the reference depth data. 13. The method of claim 12, further comprising applying a scale factor to the depth image, wherein the scale factor comprises a ratio of an altitude of the UAV above ground level relative to a median of a monocular depth map, wherein the monocular depth map is based on the reference depth data. 15. The method of Claim 13, further comprising applying a scale factor to the depth image, wherein the scale factor comprises a ratio of an altitude of the UAV above ground level relative to an above ground level estimate from a monocular depth map, wherein the monocular depth map is based on the reference depth data. 14. The method of claim 12, further comprising applying a scale factor to the depth image, wherein the scale factor comprises a ratio of an altitude of the UAV above ground level relative to an above ground level estimate from a monocular depth map, wherein the monocular depth map is based on the reference depth data. 16. The method of Claim 1, wherein the one or more semantic labels are selected from a predetermined set of labels, wherein the predetermined set of labels comprises at least the following labels: building, road, vegetation, vehicle, driveway, lawn, and sidewalk. 15. The method of claim 1, wherein the one or more semantic labels are selected from a predetermined set of labels, wherein the predetermined set of labels comprises at least the following labels: building, road, vegetation, vehicle, driveway, lawn, and sidewalk. 17. The method of Claim 1, further comprising retrieving the reference depth data in advance of a flight of the UAV, wherein the reference depth data is selected based on a planned flight path of the UAV. 16. The method of claim 1, further comprising retrieving the reference depth data in advance of a flight of the UAV, wherein the reference depth data is selected based on a planned flight path of the UAV. 18. The method of Claim 1, further comprising applying a Kalman filter to the determined location of the UAV in the environment to control navigation of the UAV in the environment. 17. The method of claim 1, further comprising applying a Kalman filter to the determined location of the UAV in the environment to control navigation of the UAV in the environment. A difference is that claim 1 of the instant application recites “an image” and generic steps of “determining, based on the image, a semantic image of the environment and a depth image of the environment” and “wherein the semantic image and the depth are determined using a machine learning model”, whereas claim 1 of Patent ‘710 recites “a two-dimensional (2D) image” and a specific step of “applying a trained machine learning model to the 2D image to produce a semantic image of the environment and a depth image of the environment”. Although the instant application claims a broader scope, it is fully encompassed by the narrower, specific scope already granted in the reference patent. Because the instant application omits these specific limitations to claim the underlying inventive concept in its generic form, the claims are directed to patentably indistinct variations of the same invention. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 9-15 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 9 recites the limitation "the depth" in line 1. There is insufficient antecedent basis for this limitation in the claim. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 7-8, 16-17, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena of AU 2021204188 A1, published 07/15/2021, hereinafter “Abeywardena”, in view of Schonberger of “Semantic Visual Localization”, published 12/15/2017, hereinafter “Schonberger”. Regarding claim 1, Abeywardena teaches: A method comprising: receiving an image captured by a camera on a unmanned aerial vehicle (UAV) and representative of an environment of the UAV; (See at least [0083]: “Figure 3 illustrates a UAV 300 equipped with a backup navigation system, according to an exemplary embodiment. As illustrated in Figure 3, the backup navigation system can include an image capture device 304 that is configured to capture images of the UAV's environment” & [0091]: “In an embodiment, the backup navigation system can detect unique visual features in the environment by processing captured images of the environment and detecting visual features in the images. For instance, the backup navigation system can cause the image capture device 304 to periodically capture images of the UAV 300's environment as the UAV 300 flies along the flight path. Then, the backup navigation system can process the captured images to detect visual features in the images. Within examples, the portion of the environment that is captured in the image depends on a location of the UAV 300, specifications of the image capture device 304, and the orientation of the image capture device 304.”) determining, based on the image, a semantic image of the environment(See at least [0129]: “Figure 5 illustrates an image 500 of an area 502 that has a hilly or mountainous terrain, according to an exemplary embodiment. In an example, the image 500 is captured by an image capture device of a UAV that is flying over the area 502 en route to a destination. Within examples, the backup navigation system of the UAV can process the image 500 to detect features of the area 502. For example, the backup navigation system can detect features 504 and 506 of the area 502. As illustrated in Figure 5, the features 504, 506 can be features of the terrain that have unique characteristics. For instance, the characteristics of a hilly terrain feature can be a slope, a ridge, a valley, a hill, a depression, a saddle, a cliff, among other characteristics. As also illustrated in Figure 5, the backup navigation system can mark the detected features 504, 506 with a dotted line that encompasses the feature.”) retrieving reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels; (See at least [0096-0097]: “The backup navigation system can then store the localized feature in a memory of the UAV 300. In particular, the backup navigation system can store data indicative of identifiable characteristics of the localized feature, such as a respective portion of each image that includes the localized feature or any other data that can be used to identify the feature. Additionally, the backup navigation system can store data indicative of a location of the localized feature…In an embodiment, the backup navigation system can use the stored data to generate a visual log or map of the flight path. That is, the backup navigation system can generate a map that includes the features that are localized during the UAV 300's flight. In the generated map, each feature can be placed at its respective location, and therefore, the map can be representative of the environment along the flight path to the destination” & [0118]: “The backup navigation system can then use the localized features to generate and/or update a map of the flight path 480. For example, the map can include the building 412 that has been localized. The map can indicate the location of the building 412 in the city 402 and can also indicate unique characteristics of the building 412 that can be used to identify the building 412.”) determining a location of the UAV in the environment, wherein determining the location of the UAV in the environment comprises (See at least [0101-0103]: “In an embodiment, the backup navigation system can use the generated map of the flight path to the destination (i.e., initial flight path) to determine the UAV 300's location. To use the map, the backup navigation system can capture an image of the UAV 300's current environment, and can process the captured image to detect visual features of the environment. The backup navigation system can then compare the detected visual features to features in the map (i.e., features that were localized when the backup navigation system was operating in the feature localization mode). If the backup navigation system detects a relationship between the detected visual features and the localized features, the system can determine the location of the detected feature. In an example, a relationship between the detected feature and the localized feature can be that the detected feature and the localized feature are the same feature in the environment. That is, the backup navigation system captured an image of a feature that it has localized before. In such an example, the detected feature has the same location as the localized features…Once the backup navigation system determines a location of a detected feature, the system can then use the location of the feature to determine a location of the UAV 300. As explained above, in practice, determining a location from two-dimensional images can require two known locations in the images in order to use the known locations to calculate the unknown location. Therefore, in an embodiment, the backup navigation system can determine a respective location of two or more detected features in the captured image in order to determine a location of the UAV 300. By way of example, the backup navigation system can determine that two identified features match two respective localized features in the map, and can determine the location of the two identified features based on the locations of the respective localized features in the map. Then, the system can calculate the location of the UAV 300 using the locations of the two identified features. For instance, the system can use the two known locations to triangulate the location of the UAV 300.”) controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment. (See at least [0100]: “In an embodiment, the backup navigation system can navigate the UAV 300 to the location from which the UAV 300 began its current flight (i.e., the starting or launch point). In particular, responsive to detecting the failure event, the backup navigation system can transition from operating in the feature localization mode to operating in a UAV localization mode in which the system can determine a location of the UAV and can navigate the UAV to the safe zone. In order to navigate the UAV 300 to the safe zone, the backup navigation system can determine a current location of the UAV 300” & [0105]: “Once the backup navigation system has determined the current location of the UAV 300, the system can determine a flight path to the safe zone. This flight path is also referred to herein as a return flight path.”) Abeywardena does not explicitly teach: determining, based on the image…a depth image of the environment; …aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data… Schonberger teaches: determining, based on the image…a depth image of the environment; (See at least Section 3-3.1: “The input to our system is a set of color images with associated depth maps I = {Ii} and, for database images, their respective camera poses P ={Pi}with Pi ∈ SE(3)…We first compute dense pixelwise semantic segmentations S = {Si} for all input images, where each pixel of Si is assigned a semantic class label l ∈{1,...,L}. Next, we fuse the images into semantic 3D voxel maps MD and MQ for the database and query images [22,25]”, Supplementary Material Fig. 4 & Supplementary Material Section 2: “The depth maps were computed by two-view stereo between the left and right camera using semi-global matching [1]. The images, depth maps, and semantic segmentations are jointly fused into semantic 3D maps, which are stored in an efficient Octree data structure at a maximum leaf node resolution of 0.3m.”) …aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data… (See at least Sections 3.4-3.5: “To find matches between a given query image and the database, we find the top K =5nearest database words D for each query word f(vj) ∈F(MQ) by traversing the vocabulary tree and finding nearest neighbors in Hamming space. Since our descriptors are rotation variant, as they are trained on aligned subvolumes (see Section 3.2), and, generally, we have no a priori knowledge about the orientation of the query, we perform the same query for a fixed set of orientation hypotheses θ ∈ SO(3) while the database remains fixed. The set of putative matches D(θ) for the different orientations provide evidence for the location of the query. The next section details how to accurately localize the query based on this evidence using a joint semantic map alignment and verification. Given the putative matches D(θ) from Section 3.4, we seek to find the transformation P ∈ SE(3) that best aligns the query to the database map. Specifically, a good alignment is established if both the geometry (i.e., occupancy) as well as the semantics agree. Due to the rotation variance of our descriptors, a single 3D-3D match between the query and the database defines a transformation hypothesis P, which is composed of the rotation defined by θ and the translation t ∈ R3 defined by the spatial offset of the corresponding subvolumes. We exhaustively enumerate all transformation hypotheses defined by the matches. To verify a single transformation hypothesis, we then align the query to the database map using P and count the number of correctly aligned voxels of the query map. A correctly aligned voxel matches both in terms of geometry and semantics, i.e., an occupied voxel in the aligned query map must also be occupied in the spatially closest voxel in the database map. In addition, the spatial distance of their voxel centers must be smaller than κ and their semantic class labels must match exactly.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena’s method with Schonberger’s technique of determining, based on the image, a depth image of the environment and aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data. Doing so would be obvious so that “The resulting 3D descriptors are robust to missing observations by encoding high-level 3D geometric and semantic information”, to achieve “Accurate camera pose estimation under strong viewpoint changes and illumination/seasonal changes”, and because “robustness against such changes is important, e.g., for AR devices or robots to re-localize robustly in a changing environment”. (See Abstract and Section 1 of Schonberger). Regarding claim 2, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Abeywardena additionally teaches: further comprising: controlling the UAV to navigate in the environment using a Global Navigation Satellite System (GNSS) system; (See at least [0088]: “In an embodiment, the primary navigation system can navigate the UAV 300 along the flight path to the destination. In some examples, the primary navigation system can rely on a satellite or other remote computing device for positional information of the UAV 300, which the system can then use to determine a flight path of the UAV 300. By way of example, the primary navigation system can rely on GPS that can provide the system with the UAV's GPS coordinates. The system can then use the UAV's GPS coordinates to determine the UAV's location with respect to the flight path, and can make adjustments to the flight path and/or the UAV's operation (e.g., speed, flight altitude, etc.) as necessary.”) detecting a disruption in service from the GNSS system, wherein the location of the UAV in the environment is determined responsive to detecting the disruption in service from the GNSS system; and (See at least [0099-0100]: “As explained above, encountering such failure events can render the primary navigation system inoperable, and the primary navigation system may not be able to navigate the UAV 300 to the destination. Accordingly, to avoid undesirable consequences, the UAV 300, responsive to the UAV 300 and/or the primary navigation system encountering a failure event, can determine to return to a safe landing zone…In particular, responsive to detecting the failure event, the backup navigation system can transition from operating in the feature localization mode to operating in a UAV localization mode in which the system can determine a location of the UAV and can navigate the UAV to the safe zone. In order to navigate the UAV 300 to the safe zone, the backup navigation system can determine a current location of the UAV 300.”) subsequent to detecting the disruption in service from the GNSS system, controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment. (See at least [0100]: “In particular, responsive to detecting the failure event, the backup navigation system can transition from operating in the feature localization mode to operating in a UAV localization mode in which the system can determine a location of the UAV and can navigate the UAV to the safe zone. In order to navigate the UAV 300 to the safe zone, the backup navigation system can determine a current location of the UAV 300. And based on the current location of the UAV 300, the system can determine or update a return flight path from the location at which the UAV 300 was located when the failure event occurred to the safe zone.”) Regarding claim 3, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Abeywardena additionally teaches: further comprising: controlling the UAV to navigate in the environment using a GNSS system; and (See at least [0088]: “In an embodiment, the primary navigation system can navigate the UAV 300 along the flight path to the destination. In some examples, the primary navigation system can rely on a satellite or other remote computing device for positional information of the UAV 300, which the system can then use to determine a flight path of the UAV 300. By way of example, the primary navigation system can rely on GPS that can provide the system with the UAV's GPS coordinates. The system can then use the UAV's GPS coordinates to determine the UAV's location with respect to the flight path, and can make adjustments to the flight path and/or the UAV's operation (e.g., speed, flight altitude, etc.) as necessary.”) using the determined location of the UAV in the environment to cross-check location data from the GNSS system. (See at least [0041]: “In addition to providing backup navigation, the backup navigation system disclosed herein can also perform other functions, such as checking the integrity or accuracy of the navigation provided by the primary navigation system…In an implementation of this function, the backup navigation system can periodically determine the UAV's location (e.g., using the methods described herein), and can then compare the UAV location to the location determined by the primary navigation system.”) Regarding claim 7, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Schonberger additionally teaches: wherein the semantic image of the environment and the depth image of the environment have the same dimensions. (See at least Supplementary Material Fig. 4, which shows an input image and its corresponding depth map and semantic segmentation in the left column.) Regarding claim 8, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Abeywardena additionally teaches: wherein the camera on the UAV faces downward, and wherein the image captured by the camera is representative of a terrain in the environment below the UAV. (See at least [0111]: “Figure 4B illustrates a representation of an image 420 of the area 410, according to an exemplary embodiment. In this scenario, the configuration of the image capture device 304 is such that image capture device 304 captures a top view image of the environment (e.g., the city 402). Accordingly, the vantage point of the image 420 is a top view (i.e., bird's eye view) of the area 410.”) Regarding claim 16, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Schonberger additionally teaches: wherein the one or more semantic labels are selected from a predetermined set of labels, (See at least Section 3.1: “Given a robust semantic classifier, e.g., trained specifically for different seasons, the semantic maps are inherently invariant to large illumination changes and geometric variations up to the voxel resolution. Note that using semantics, it is easy to determine reliable classes and, e.g., to ignore dynamic objects such as cars”, 3.4: “We establish correspondence between subvolumes in the query and database using nearest neighbor search in the descriptor space f(v) ∈ RN using the Euclidean metric. For efficient semantic word matching, we build a semantic vocabulary [47] in an offline procedure using the bag of semantic words of the training dataset” & Supplementary Material Section 1: “The visual vocabulary is represented by 216 visual words embedded in a NB = 64 dimensional Hamming space and using a hierarchical branching factor of 256. Using these settings, we obtain several thousand descriptors per image.”) Abeywardena and Schonberger in combination do not explicitly teach: wherein the predetermined set of labels comprises at least the following labels: building, road, vegetation, vehicle, driveway, lawn, and sidewalk. However, Schonberger does teach building a semantic vocabulary in a training procedure, as discussed above, and a classifier that “segments the scene into L = 19 semantic classes and we only consider the maximum activation per pixel and discard any pixels with sky labels” (See at least Section 4.2). Supplementary Material Fig. 4 of Schonberger also shows semantic segmentation images that are color-coded with respect to different semantic classes, such as the building, road, vegetation, vehicle, and sidewalk. Since Schonberger teaches semantically segmenting a scene with respect to semantic classes, it would be obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to select specific labels as desired by the user, such as labels for a building, road, vegetation, vehicle, driveway, lawn, and sidewalk, which provides the benefit of “it is easy to determine reliable classes and, e.g., to ignore dynamic objects such as cars” (See Section 3.1 of Schonberger). Regarding claim 17, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Abeywardena additionally teaches: further comprising retrieving the reference depth data in advance of a flight of the UAV, wherein the reference depth data is selected based on a planned flight path of the UAV. (See at least [0038]: “In some instances, the backup navigation system can determine to continue travelling to the original destination (i.e., continue the mission). In other instances, for various reasons (e.g., fuel considerations), the backup navigation system can determine to fly the UAV to a safe zone that is not in an area that was traversed by the UAV. In either instance, since the UAV did not traverse the area, the UAV may have not generated a map of the area, and therefore the backup navigation system may not be able to navigate to the desired safe zone. To overcome this obstacle, also disclosed herein is a method of compiling a database of localized features from images captured by a plurality of UAVs that are operating in a given area over a given amount of time. In particular, a controller of the database can process the images (and the features extracted from thereof) to generate one or more maps of the given area. Then any UAV whose flight path overlaps the given area will receive a copy of the one or maps so that the UAV's backup navigation system can use the maps to navigate to a safe zone in the given area.”) Regarding claim 19, Abeywardena teaches: An unmanned aerial vehicle (UAV), comprising: a camera; and (See at least [0083]: “Figure 3 illustrates a UAV 300 equipped with a backup navigation system, according to an exemplary embodiment. As illustrated in Figure 3, the backup navigation system can include an image capture device 304 that is configured to capture images of the UAV's environment.”) a control system configured to: receive an image captured by the camera on the UAV and representative of an environment of the UAV; (See at least [0091]: “In an embodiment, the backup navigation system can detect unique visual features in the environment by processing captured images of the environment and detecting visual features in the images. For instance, the backup navigation system can cause the image capture device 304 to periodically capture images of the UAV 300's environment as the UAV 300 flies along the flight path. Then, the backup navigation system can process the captured images to detect visual features in the images.”) determine, based on the image, a semantic image of the environment(See at least [0129]: “Figure 5 illustrates an image 500 of an area 502 that has a hilly or mountainous terrain, according to an exemplary embodiment. In an example, the image 500 is captured by an image capture device of a UAV that is flying over the area 502 en route to a destination. Within examples, the backup navigation system of the UAV can process the image 500 to detect features of the area 502. For example, the backup navigation system can detect features 504 and 506 of the area 502. As illustrated in Figure 5, the features 504, 506 can be features of the terrain that have unique characteristics. For instance, the characteristics of a hilly terrain feature can be a slope, a ridge, a valley, a hill, a depression, a saddle, a cliff, among other characteristics. As also illustrated in Figure 5, the backup navigation system can mark the detected features 504, 506 with a dotted line that encompasses the feature.”) retrieve reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels; (See at least [0096-0097]: “The backup navigation system can then store the localized feature in a memory of the UAV 300. In particular, the backup navigation system can store data indicative of identifiable characteristics of the localized feature, such as a respective portion of each image that includes the localized feature or any other data that can be used to identify the feature. Additionally, the backup navigation system can store data indicative of a location of the localized feature…In an embodiment, the backup navigation system can use the stored data to generate a visual log or map of the flight path. That is, the backup navigation system can generate a map that includes the features that are localized during the UAV 300's flight. In the generated map, each feature can be placed at its respective location, and therefore, the map can be representative of the environment along the flight path to the destination” & [0118]: “The backup navigation system can then use the localized features to generate and/or update a map of the flight path 480. For example, the map can include the building 412 that has been localized. The map can indicate the location of the building 412 in the city 402 and can also indicate unique characteristics of the building 412 that can be used to identify the building 412.”) determine a location of the UAV in the environment, wherein determining the location of the UAV in the environment comprises labels from the reference depth data; and (See at least [0101-0103]: “In an embodiment, the backup navigation system can use the generated map of the flight path to the destination (i.e., initial flight path) to determine the UAV 300's location. To use the map, the backup navigation system can capture an image of the UAV 300's current environment, and can process the captured image to detect visual features of the environment. The backup navigation system can then compare the detected visual features to features in the map (i.e., features that were localized when the backup navigation system was operating in the feature localization mode). If the backup navigation system detects a relationship between the detected visual features and the localized features, the system can determine the location of the detected feature. In an example, a relationship between the detected feature and the localized feature can be that the detected feature and the localized feature are the same feature in the environment. That is, the backup navigation system captured an image of a feature that it has localized before. In such an example, the detected feature has the same location as the localized features…Once the backup navigation system determines a location of a detected feature, the system can then use the location of the feature to determine a location of the UAV 300. As explained above, in practice, determining a location from two-dimensional images can require two known locations in the images in order to use the known locations to calculate the unknown location. Therefore, in an embodiment, the backup navigation system can determine a respective location of two or more detected features in the captured image in order to determine a location of the UAV 300. By way of example, the backup navigation system can determine that two identified features match two respective localized features in the map, and can determine the location of the two identified features based on the locations of the respective localized features in the map. Then, the system can calculate the location of the UAV 300 using the locations of the two identified features. For instance, the system can use the two known locations to triangulate the location of the UAV 300.”) control the UAV to navigate in the environment based on the determined location of the UAV in the environment. (See at least [0100]: “In an embodiment, the backup navigation system can navigate the UAV 300 to the location from which the UAV 300 began its current flight (i.e., the starting or launch point). In particular, responsive to detecting the failure event, the backup navigation system can transition from operating in the feature localization mode to operating in a UAV localization mode in which the system can determine a location of the UAV and can navigate the UAV to the safe zone. In order to navigate the UAV 300 to the safe zone, the backup navigation system can determine a current location of the UAV 300” & [0105]: “Once the backup navigation system has determined the current location of the UAV 300, the system can determine a flight path to the safe zone. This flight path is also referred to herein as a return flight path.”) Abeywardena does not explicitly teach the underlined: determine, based on the image…a depth image of the environment; …aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data… Schonberger teaches: determine, based on the image…a depth image of the environment; (See at least Section 3-3.1: “The input to our system is a set of color images with associated depth maps I = {Ii} and, for database images, their respective camera poses P ={Pi}with Pi ∈ SE(3)…We first compute dense pixelwise semantic segmentations S = {Si} for all input images, where each pixel of Si is assigned a semantic class label l ∈{1,...,L}. Next, we fuse the images into semantic 3D voxel maps MD and MQ for the database and query images [22,25]”, Supplementary Material Fig. 4 & Supplementary Material Section 2: “The depth maps were computed by two-view stereo between the left and right camera using semi-global matching [1]. The images, depth maps, and semantic segmentations are jointly fused into semantic 3D maps, which are stored in an efficient Octree data structure at a maximum leaf node resolution of 0.3m.”) …aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data… (See at least Sections 3.4-3.5: “To find matches between a given query image and the database, we find the top K =5nearest database words D for each query word f(vj) ∈F(MQ) by traversing the vocabulary tree and finding nearest neighbors in Hamming space. Since our descriptors are rotation variant, as they are trained on aligned subvolumes (see Section 3.2), and, generally, we have no a priori knowledge about the orientation of the query, we perform the same query for a fixed set of orientation hypotheses θ ∈ SO(3) while the database remains fixed. The set of putative matches D(θ) for the different orientations provide evidence for the location of the query. The next section details how to accurately localize the query based on this evidence using a joint semantic map alignment and verification. Given the putative matches D(θ) from Section 3.4, we seek to find the transformation P ∈ SE(3) that best aligns the query to the database map. Specifically, a good alignment is established if both the geometry (i.e., occupancy) as well as the semantics agree. Due to the rotation variance of our descriptors, a single 3D-3D match between the query and the database defines a transformation hypothesis P, which is composed of the rotation defined by θ and the translation t ∈ R3 defined by the spatial offset of the corresponding subvolumes. We exhaustively enumerate all transformation hypotheses defined by the matches. To verify a single transformation hypothesis, we then align the query to the database map using P and count the number of correctly aligned voxels of the query map. A correctly aligned voxel matches both in terms of geometry and semantics, i.e., an occupied voxel in the aligned query map must also be occupied in the spatially closest voxel in the database map. In addition, the spatial distance of their voxel centers must be smaller than κ and their semantic class labels must match exactly.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena’s method with Schonberger’s technique of determining, based on the image, a depth image of the environment and aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data. Doing so would be obvious so that “The resulting 3D descriptors are robust to missing observations by encoding high-level 3D geometric and semantic information”, to achieve “Accurate camera pose estimation under strong viewpoint changes and illumination/seasonal changes”, and because “robustness against such changes is important, e.g., for AR devices or robots to re-localize robustly in a changing environment”. (See Abstract and Section 1 of Schonberger). Regarding claim 20, Abeywardena teaches: A non-transitory computer readable medium comprising program instructions executable by one or more processors to perform operations, the operations comprising: receiving an image captured by a camera on a unmanned aerial vehicle (UAV) and representative of an environment of the UAV; (See at least [0083]: “Figure 3 illustrates a UAV 300 equipped with a backup navigation system, according to an exemplary embodiment. As illustrated in Figure 3, the backup navigation system can include an image capture device 304 that is configured to capture images of the UAV's environment” & [0091]: “In an embodiment, the backup navigation system can detect unique visual features in the environment by processing captured images of the environment and detecting visual features in the images. For instance, the backup navigation system can cause the image capture device 304 to periodically capture images of the UAV 300's environment as the UAV 300 flies along the flight path. Then, the backup navigation system can process the captured images to detect visual features in the images. Within examples, the portion of the environment that is captured in the image depends on a location of the UAV 300, specifications of the image capture device 304, and the orientation of the image capture device 304.”) determining, based on the image, a semantic image of the environment(See at least [0129]: “Figure 5 illustrates an image 500 of an area 502 that has a hilly or mountainous terrain, according to an exemplary embodiment. In an example, the image 500 is captured by an image capture device of a UAV that is flying over the area 502 en route to a destination. Within examples, the backup navigation system of the UAV can process the image 500 to detect features of the area 502. For example, the backup navigation system can detect features 504 and 506 of the area 502. As illustrated in Figure 5, the features 504, 506 can be features of the terrain that have unique characteristics. For instance, the characteristics of a hilly terrain feature can be a slope, a ridge, a valley, a hill, a depression, a saddle, a cliff, among other characteristics. As also illustrated in Figure 5, the backup navigation system can mark the detected features 504, 506 with a dotted line that encompasses the feature.”) retrieving reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels; (See at least [0096-0097]: “The backup navigation system can then store the localized feature in a memory of the UAV 300. In particular, the backup navigation system can store data indicative of identifiable characteristics of the localized feature, such as a respective portion of each image that includes the localized feature or any other data that can be used to identify the feature. Additionally, the backup navigation system can store data indicative of a location of the localized feature…In an embodiment, the backup navigation system can use the stored data to generate a visual log or map of the flight path. That is, the backup navigation system can generate a map that includes the features that are localized during the UAV 300's flight. In the generated map, each feature can be placed at its respective location, and therefore, the map can be representative of the environment along the flight path to the destination” & [0118]: “The backup navigation system can then use the localized features to generate and/or update a map of the flight path 480. For example, the map can include the building 412 that has been localized. The map can indicate the location of the building 412 in the city 402 and can also indicate unique characteristics of the building 412 that can be used to identify the building 412.”) determining a location of the UAV in the environment, wherein determining the location of the UAV in the environment comprises (See at least [0101-0103]: “In an embodiment, the backup navigation system can use the generated map of the flight path to the destination (i.e., initial flight path) to determine the UAV 300's location. To use the map, the backup navigation system can capture an image of the UAV 300's current environment, and can process the captured image to detect visual features of the environment. The backup navigation system can then compare the detected visual features to features in the map (i.e., features that were localized when the backup navigation system was operating in the feature localization mode). If the backup navigation system detects a relationship between the detected visual features and the localized features, the system can determine the location of the detected feature. In an example, a relationship between the detected feature and the localized feature can be that the detected feature and the localized feature are the same feature in the environment. That is, the backup navigation system captured an image of a feature that it has localized before. In such an example, the detected feature has the same location as the localized features…Once the backup navigation system determines a location of a detected feature, the system can then use the location of the feature to determine a location of the UAV 300. As explained above, in practice, determining a location from two-dimensional images can require two known locations in the images in order to use the known locations to calculate the unknown location. Therefore, in an embodiment, the backup navigation system can determine a respective location of two or more detected features in the captured image in order to determine a location of the UAV 300. By way of example, the backup navigation system can determine that two identified features match two respective localized features in the map, and can determine the location of the two identified features based on the locations of the respective localized features in the map. Then, the system can calculate the location of the UAV 300 using the locations of the two identified features. For instance, the system can use the two known locations to triangulate the location of the UAV 300.”) controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment. (See at least [0100]: “In an embodiment, the backup navigation system can navigate the UAV 300 to the location from which the UAV 300 began its current flight (i.e., the starting or launch point). In particular, responsive to detecting the failure event, the backup navigation system can transition from operating in the feature localization mode to operating in a UAV localization mode in which the system can determine a location of the UAV and can navigate the UAV to the safe zone. In order to navigate the UAV 300 to the safe zone, the backup navigation system can determine a current location of the UAV 300” & [0105]: “Once the backup navigation system has determined the current location of the UAV 300, the system can determine a flight path to the safe zone. This flight path is also referred to herein as a return flight path.”) Abeywardena does not explicitly teach: determining, based on the image…a depth image of the environment; …aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data… Schonberger teaches: determining, based on the image…a depth image of the environment; (See at least Section 3-3.1: “The input to our system is a set of color images with associated depth maps I = {Ii} and, for database images, their respective camera poses P ={Pi}with Pi ∈ SE(3)…We first compute dense pixelwise semantic segmentations S = {Si} for all input images, where each pixel of Si is assigned a semantic class label l ∈{1,...,L}. Next, we fuse the images into semantic 3D voxel maps MD and MQ for the database and query images [22,25]”, Supplementary Material Fig. 4 & Supplementary Material Section 2: “The depth maps were computed by two-view stereo between the left and right camera using semi-global matching [1]. The images, depth maps, and semantic segmentations are jointly fused into semantic 3D maps, which are stored in an efficient Octree data structure at a maximum leaf node resolution of 0.3m.”) …aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data… (See at least Sections 3.4-3.5: “To find matches between a given query image and the database, we find the top K =5nearest database words D for each query word f(vj) ∈F(MQ) by traversing the vocabulary tree and finding nearest neighbors in Hamming space. Since our descriptors are rotation variant, as they are trained on aligned subvolumes (see Section 3.2), and, generally, we have no a priori knowledge about the orientation of the query, we perform the same query for a fixed set of orientation hypotheses θ ∈ SO(3) while the database remains fixed. The set of putative matches D(θ) for the different orientations provide evidence for the location of the query. The next section details how to accurately localize the query based on this evidence using a joint semantic map alignment and verification. Given the putative matches D(θ) from Section 3.4, we seek to find the transformation P ∈ SE(3) that best aligns the query to the database map. Specifically, a good alignment is established if both the geometry (i.e., occupancy) as well as the semantics agree. Due to the rotation variance of our descriptors, a single 3D-3D match between the query and the database defines a transformation hypothesis P, which is composed of the rotation defined by θ and the translation t ∈ R3 defined by the spatial offset of the corresponding subvolumes. We exhaustively enumerate all transformation hypotheses defined by the matches. To verify a single transformation hypothesis, we then align the query to the database map using P and count the number of correctly aligned voxels of the query map. A correctly aligned voxel matches both in terms of geometry and semantics, i.e., an occupied voxel in the aligned query map must also be occupied in the spatially closest voxel in the database map. In addition, the spatial distance of their voxel centers must be smaller than κ and their semantic class labels must match exactly.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena’s method with Schonberger’s technique of determining, based on the image, a depth image of the environment and aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data. Doing so would be obvious so that “The resulting 3D descriptors are robust to missing observations by encoding high-level 3D geometric and semantic information”, to achieve “Accurate camera pose estimation under strong viewpoint changes and illumination/seasonal changes”, and because “robustness against such changes is important, e.g., for AR devices or robots to re-localize robustly in a changing environment”. (See Abstract and Section 1 of Schonberger). Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena in view of Schonberger and further in view of Lee of US 20210078593 A1, published 03/18/2021, hereinafter “Lee”. Regarding claim 4, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Abeywardena additionally teaches: further comprising: determining a GNSS location of the UAV in the environment using a GNSS system; (See at least [0088]: “In an embodiment, the primary navigation system can navigate the UAV 300 along the flight path to the destination. In some examples, the primary navigation system can rely on a satellite or other remote computing device for positional information of the UAV 300, which the system can then use to determine a flight path of the UAV 300. By way of example, the primary navigation system can rely on GPS that can provide the system with the UAV's GPS coordinates. The system can then use the UAV's GPS coordinates to determine the UAV's location with respect to the flight path, and can make adjustments to the flight path and/or the UAV's operation (e.g., speed, flight altitude, etc.) as necessary.”) Abeywardena and Schonberger in combination do not explicitly teach: determining a refined location of the UAV in the environment based on the GNSS location of the UAV in the environment and the determined location of the UAV in the environment; and controlling the UAV to navigate in the environment based on the refined location of the UAV in the environment. Lee teaches: determining a refined location of the UAV in the environment based on the GNSS location of the UAV in the environment and the determined location of the UAV in the environment; and (See at least [0109]: “FIG. 13 shows an example of AV 100 navigating in environment 190. AV system 120 uses one or more types of navigational information for determining a location of the AV 100 in the environment 190. One type of navigational information is a semantic map of the environment (e.g., a map with annotated features of the roadway). For example, the locations of lane markings, street signs, landmarks, or other distinct features of the roadway are included in the semantic map. One or more sensors in the AV system 120 (e.g., cameras, LiDAR, RADAR) detect features of the roadway proximate to AV 100 and then the AV system 120 compares the detected features with the features in the semantic map to determine a location of the AV 100. Another type of navigational information is GPS coordinates. In some embodiments, the AV system 120 uses a combination of different types of navigational information to determine a more precise location of the AV 100” & [0111]: “In some embodiments, the AV system 120 uses GPS to determine approximate coordinates of the AV 100 (e.g., latitude and longitude). The AV system 120 also obtains data from one or more additional sensors (e.g., cameras, LiDAR, and/or RADAR). If the AV 100 is in the mapped region 1302, then a more precise position of the AV 100 can be determined based on the data obtained from the additional sensors and the navigational information (e.g., semantic map) available in the mapped region 1302. For example, the additional sensors may detect road features near the AV 100 (e.g., lane markings, signs, landmarks). When the road features correspond to features in the semantic map at the approximate coordinates determined by the GPS data, the more precise position of vehicle can be determined.”) controlling the UAV to navigate in the environment based on the refined location of the UAV in the environment. (See at least [0112]: “While the AV 100 is in the mapped region 1302 (where navigational information is available), the AV 100 can be operated in an autonomous mode where control functions of the vehicle (e.g., steering, throttling, braking, ignition) are automated (e.g., a fully or highly autonomous mode (Level 3, 4, or 5)). While the AV 100 is operating in the autonomous mode, the AV system 120 navigates the vehicle toward a destination based on data from the additional sensors (e.g., data from cameras, LiDAR, and/or RADAR) in combination with data from GPS.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena and Schonberger’s method with Lee’s technique of determining a refined location of the UAV in the environment based on the GNSS location of the UAV in the environment and the determined location of the UAV in the environment and controlling the UAV to navigate in the environment based on the refined location of the UAV in the environment. Doing so would be obvious so “the more precise position of vehicle can be determined” (See [0111] of Lee). Claim(s) 5-6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena in view of Schonberger and further in view of Kerzner of US 20210110137 A1, published 04/15/2021, hereinafter “Kerzner”. Regarding claim 5, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Abeywardena and Schonberger in combination do not explicitly teach: wherein the aligning the depth image of the environment with the reference depth data representative of the environment to determine the location of the UAV in the environment comprises using an iterative closest point (ICP) algorithm. Kerzner teaches: wherein the aligning the depth image of the environment with the reference depth data representative of the environment to determine the location of the UAV in the environment comprises using an iterative closest point (ICP) algorithm. (See at least [0057]: “For example, the drone 102 may use its camera 104 and one or more other sensing devices, such as a depth sensor or one or more other onboard cameras, to detect previously identified visual landmarks” & [0170-0171]: “In some cases, where no landmark (e.g., no landmark in the viewable vicinity of the drone 102) is detected with a sufficient reliability score (or an insufficient number of landmarks have a sufficient reliability score), a refinement algorithm may be executed with positions of detected landmarks that could have moved. For example, a threshold reliability score may be set to 95%. If the drone 102 detects a landmark that has a reliability score of at least 95%, then there is sufficient confidence in the landmark for the drone 102 to determine, for example, its current position and/or pose, and/or to continue navigation using the detected landmarks. However, if the drone 102 fails to detect a landmark with a reliability score of at least 95%, then the drone 102 may trigger the execution of a refinement algorithm (e.g., iterative closest point algorithm) with positions of detected landmarks that could have moved. For example, the drone 102 may detect a chair with a reliability score of 80% and a table with a reliability score of 90%. The drone 102 may provide the positions of the chair and the table to a refinement algorithm which may use the positions as an initial estimate (e.g., an initial estimate of the drone 102's location in the property and/or of the positions of the table and the chair). However, leveraging a sufficiently reliable landmark, when detected, may be preferred over executing a refinement algorithm since such algorithms can be to CPU demanding. The output of the refinement algorithm may be a more accurate location of the drone 102 in the property, and/or more accurate positions of the chair and table.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena and Schonberger’s method with Kerzner’s technique of using an iterative closest point algorithm to determine the current position of a drone. Doing so would be obvious since “The output of the refinement algorithm(s) may be used to improve localization (e.g., to identify a more accurate location of the drone 102 with respect to the property 120)” (See [0173] of Kerzner). Regarding claim 6, Abeywardena, Schonberger, and Kerzner in combination teach all the limitations of claim 5 as discussed above. Kerzner additionally teaches: wherein the ICP algorithm aligns points from the reference depth data with points from the depth image such that the reference semantic labels from the reference depth data correspond to the one or more semantic labels from the semantic image. (See at least [0070-0071]: “In analyzing the captured images, the monitoring server 130 may identify one or more visual landmarks from the captured images and/or the generated 3D environment map. In identifying one or more visual landmarks, the monitoring server 130 may employ segmentation to label areas within the captured images where physical objects are and/or what the physical objects are. Specifically, the monitoring server 130 may employ two dimensional (2D) scene segmentation, which maps each pixel to a different object or surface category, e.g., wall, floor, furniture, etc. For example, the monitoring server 130 may perform segmentation on the subset of images 108, which, as previously explained, may have been subsampled spatially or temporally as the drone 102 moved through the monitored property 120. The resultant segmentation maps would be registered to the 3D environment map used for navigation, and the results fused. In areas where the segmentation results differ, the monitoring server 130 may associate those areas with a lower confidence of the segmentation. The monitoring server 130 may actually calculate a confidence score for each area or may make a determination as to whether an area is acceptable, e.g., due to consistent segmentation results, or is unacceptable, e.g., due to differing segmentation results” & [0080]: “The monitoring server 130 may analyze the captured images and/or the results of the scene segmentation to confirm the existence of the expected features of the visual landmarks.” See also [0170-0171].) NOTE: Claim 6 recites the following intended use limitation: “…such that the reference semantic labels from the reference depth data correspond to the one or more semantic labels from the semantic image”. This limitation is not positively recited, and is instead recited as intended use since it recites an intended result for “the ICP algorithm aligns points from the reference depth data with points from the depth image”. Therefore, the BRI of claim 6 does not require the aforementioned intended use limitation. Claim(s) 9-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena in view of Schonberger and further in view of Lin of US 20200241574 A1, published 07/30/2020, hereinafter “Lin”. Regarding claim 9, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Abeywardena and Schonberger in combination do not explicitly teach: wherein the semantic image and the depth are determined using a machine learning model. Lin teaches: wherein the semantic image and the depth are determined using a machine learning model. (See at least [0052]: “In other words, the feature representation generator 126 is configured to receive an image of the environment 106 from the camera 116, including the object 108, and to utilize a suitable CNN to represent features of the obtained image, including both a semantic segmentation and a depth map thereof. Of course, during training operations of the training manager 122, such images are represented by simulated or actual training images obtained from the training data 124” and [0074]: “Using at least one convolutional neural network, a semantic segmentation and depth map of the image be determined (206). The semantic segmentation may include labeling of image pixels of the image corresponding to the object. Meanwhile, in the depth map, image pixels of the image corresponding to the object are associated with the distance of the object from the robot. In the example implementations, described in detail below, the semantic segmentation and the depth map may be obtained from the image using a single convolutional neural network, or a single set of convolutional networks.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena and Schonberger’s method with Lin’s technique of determining the semantic image and the depth using a machine learning model. Doing so would be obvious to “utilize an efficient, fast, accurate, complete, and widely-applicable algorithm(s) for detecting and identifying objects in a manner that is largely or completely independent of the environments in which the objects may be found” (See [0019] of Lin). Regarding claim 10, Abeywardena, Schonberger, and Lin in combination teach all the limitations of claim 9 as discussed above. Lin additionally teaches: wherein a semantics branch and a depth branch of the machine learning model operate on a commonly generated feature set. (See at least [0050-0051]: “During operation, in order to identify and execute the approach path 110, the approach policy controller 102 initially utilizes a training manager 122. As illustrated in FIG. 1, the training manager 122 may include, or have access to, training data 124. The training data 124 may include simulations of suitable environments representing the environment 106, and/or may include real-world environmental data. In some examples below, the training data 124 may be referenced as ground truth data. For example, the training data 124 may include environments, included objects and their semantic labels, and map information that may be used to compute distances from points within the environments to the objects included therein. As also described in detail below, the training manager 122 is configured to utilize the action space repository 120 and the training data 124 to train at least two machine-learning algorithms for future use in identifying and executing the approach path 110. As shown in FIG. 1, a first example of such machine-learning algorithms includes a feature representation generator 126. In example implementations, the feature representation generator 126 represents at least one convolutional neural network (CNN) that is configured to generate both semantic segmentation data 128 and depth information data 130.”) Claim(s) 11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena in view of Schonberger and Lin and further in view of Guizilini of US 20210004976 A1, published 01/07/2021, hereinafter “Guizilini ‘976”. Regarding claim 11, Abeywardena, Schonberger, and Lin in combination teach all the limitations of claim 9 as discussed above. Lin additionally teaches: wherein the machine learning model has been trained based on ground truth depth data, (See at least [0050]: “During operation, in order to identify and execute the approach path 110, the approach policy controller 102 initially utilizes a training manager 122. As illustrated in FIG. 1, the training manager 122 may include, or have access to, training data 124. The training data 124 may include simulations of suitable environments representing the environment 106, and/or may include real-world environmental data. In some examples below, the training data 124 may be referenced as ground truth data. For example, the training data 124 may include environments, included objects and their semantic labels, and map information that may be used to compute distances from points within the environments to the objects included therein” and [0057]: “Thus, during training, the training manager 122 may be configured to utilize the feature representation generator 126 to generate semantic segmentation and depth map representations of images from within the training data 124. The training manager 122 may then utilize the ground truth data of the training data 124 to compare the calculated feature representations with the ground truth feature representations. Over multiple iterations, and using training techniques such as those referenced below, the training manager 122 may be configured to parameterize the feature representation generator 126 in a manner that minimizes feature representation errors within the environment 106.”) Abeywardena, Schonberger, and Lin in combination do not explicitly teach: wherein the ground truth depth data is based on performance of a structure from motion (SfM) algorithm on images captured by one or more UAVs. Guizilini ‘976 teaches: wherein the ground truth depth data is based on performance of a structure from motion (SfM) algorithm on images(See at least [0076]: “At 740, the training module 230 determines whether the current training stage is the first or second and jumps to adjusting the depth model 260 if at the first stage or generating the additional second-stage supervised loss if at the second stage. In one embodiment, as previously noted, the first stage of training is a self-supervised structure from motion (SfM) training process that accounts for motion of a camera between the training images of a pair to cause the depth model to learn how to infer depths without using annotated training data (i.e., without the depth data). However, because the resulting depth model 260 from solely training on the self-supervised process does not accurately understand scale (i.e., is scale ambiguous), the training module 230 further imposes the second stage to refine the depth model 260. That is, The training module 230 trains the depth model according to the second stage to refine the depth model 260 using second training data that includes the annotations about depth in the individual images. As previously noted, the sparse depth data includes selective dispersed ground truths providing limited supervision over depth estimates of the individual images.”) Although Guizilini ‘976’s invention is directed to vehicles, Guizilini ‘976 recites that “embodiments are not limited to automobiles” and that “the vehicle 100 may be any electronic/robotic device or other form of powered transport that, for example, perceives an environment according to monocular images” (See [0025] of Guizilini ‘976). As such, it would be obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to apply Guizilini ‘976’s structure from motion algorithm to images from a UAV, which provides the benefit of “improved situational awareness of the implementing device (e.g., the vehicle 100), and improved abilities to navigate and perform other functions therefrom” (See [0078] of Guizilini ‘976). One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena, Schonberger, and Lin’s method with Guizilini ‘976’s ground truth depth being based on performance of a structure from motion (SfM) algorithm on images. Doing so would be obvious to “account[s] for motion of a camera between the training images of a pair to cause the depth model to learn how to infer depths without using annotated training data (i.e., without the depth data)” (See [0076] of Guizilini ‘976). Claim(s) 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena in view of Schonberger and Lin and further in view of Atherton of US 20210240195 A1, published 08/05/2021, hereinafter “Atherton”. Regarding claim 12, Abeywardena, Schonberger, and Lin in combination teach all the limitations of claim 9 as discussed above. Lin additionally teaches: wherein the machine learning model has been trained based on ground truth semantic data, (See at least [0050]: “During operation, in order to identify and execute the approach path 110, the approach policy controller 102 initially utilizes a training manager 122. As illustrated in FIG. 1, the training manager 122 may include, or have access to, training data 124. The training data 124 may include simulations of suitable environments representing the environment 106, and/or may include real-world environmental data. In some examples below, the training data 124 may be referenced as ground truth data. For example, the training data 124 may include environments, included objects and their semantic labels, and map information that may be used to compute distances from points within the environments to the objects included therein” and [0057]: “Thus, during training, the training manager 122 may be configured to utilize the feature representation generator 126 to generate semantic segmentation and depth map representations of images from within the training data 124. The training manager 122 may then utilize the ground truth data of the training data 124 to compare the calculated feature representations with the ground truth feature representations. Over multiple iterations, and using training techniques such as those referenced below, the training manager 122 may be configured to parameterize the feature representation generator 126 in a manner that minimizes feature representation errors within the environment 106.”) Abeywardena, Schonberger, and Lin in combination do not explicitly teach: wherein the ground truth semantic data is based on operator labeling of images captured by one or more UAVs. Atherton teaches: wherein the ground truth semantic data is based on operator labeling of images captured by one or more UAVs. (See at least [0018]: “According to embodiments of the present disclosure, an annotated reference map of an environment in which the autonomous ground vehicle is operating can first be constructed. For example, images of the environment in which the autonomous ground vehicle is operating can be obtained (e.g., by aircraft, unmanned aerial vehicle, ground vehicles, etc.), and these images can be used to generate a two-dimensional reconstruction of the area (e.g., using photogrammetry, etc.)” & [0055]: “In creating the training datasets, images captured by autonomous ground vehicles during operation can be hand annotated to generate pairwise datasets containing the front camera image of the scene and the corresponding binary edge map of the scene. Alternatively, synthetic datasets can be created based on simulated operation of an autonomous ground vehicle. During simulated operation, a line segment detection algorithm can be performed on a bird's-eye camera image that is available in simulation. The depth of the lines can be given by the depth channel of the bird's-eye image and can be ray-casted back to the front camera optical frame to get the binary edge image of the line features in the environment. These pairwise camera and binary images can then be used to train a network or system such as a deep learning model.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena, Schonberger, and Lin’s method with Atherton’s ground truth semantic data being based on operator labeling of images captured by one or more UAVs. Doing so would be obvious so “Training image datasets can be created and then utilized to train a network or system (e.g., a deep learning model), which can be used to detect the environmental features as described herein” (See [0054] of Atherton). Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena in view of Schonberger and Lin and further in view of Wofk of US 20220343521 A1, filed 06/30/2022, hereinafter “Wofk”. Regarding claim 13, Abeywardena, Schonberger, and Lin in combination teach all the limitations of claim 9 as discussed above. Abeywardena, Schonberger, and Lin in combination do not explicitly teach: wherein the machine learning model has been trained using a scale invariant loss. Wofk teaches: wherein the machine learning model has been trained using a scale invariant loss. (See at least [0046]: “The monocular depth estimator circuitry 125 performs monocular depth estimation, as described in connection with FIG. 1. For example, the monocular depth estimator circuitry 125 predicts depth from a monocular image (e.g., single RGB image 105). In the example of FIG. 2, the monocular depth estimator circuitry 125 includes a pretrained model that produces a dense depth map up to a specified scale (e.g., depth estimation model 265). In some examples, a depth estimator can be selected (e.g., DPT-Hybrid, etc.) and a transformer-based model trained on a large meta-dataset using scale- and shift-invariant losses.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena, Schonberger, and Lin’s method with Wofk’s technique of training the machine learning model using a scale invariant loss for training of the depth branch. Doing so would have been obvious since DPT-Hybrid, which is a depth estimator, “is already known to generalize well after having been trained on a massive mixed dataset with scale- and shift-invariant loss functions” (See [0107] of Wofk). Claim(s) 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena in view of Schonberger, Lin, and Wofk, and further in view of Guizilini of US 20210090280 A1, published 03/25/2021, hereinafter “Guizilini ‘280”. Regarding claim 14, Abeywardena, Schonberger, Lin, and Wofk in combination teach all the limitations of claim 13 as discussed above. Wofk additionally teaches: further comprising applying a scale factor to the depth image, (See at least [0040]: “The monocular depth estimator circuitry 125 predicts depth from a monocular image. For example, the monocular depth estimator circuitry 125 can include a pretrained model that takes in a single RGB image (e.g., RGB image 105) and produces a dense depth map up to a specified scale” & [0043]: “The scale aligner circuitry 150 can be used to perform dense (local) scale alignment given that global alignment may not adequately resolve metric scale in all regions of a depth map. As such, a learning-based approach can be used for determining dense (per-pixel) scale factors that are applied to globally aligned depth estimates. In the example of FIG. 1, a ScaleMapLearner (SML) network 165 can be trained (e.g., using an open-source machine learning framework such as MiDaS-small, etc.) to realign individual values in an input depth map to improve metric accuracy. In some examples, the SML network 165 can receive an input of two concatenated data channels, such as the globally aligned depth prediction 155 and/or a scaffolding for a dense scale map (e.g., dense scale map scaffolding 160 based on a metric sparse depth 145 output by the visual-inertial odometry sensor circuitry 130), as described in more detail in connection with FIG. 2. As such, the SML network 165 regresses the dense scale residual map 170 (e.g., dense scale map image 190 regressed by the SML network 165) and a resulting scale map can be generated and applied to the input depth, yielding an example final metric dense depth output 175 (e.g., shown visually using the final depth map output 192).”) Abeywardena, Schonberger, Lin, and Wofk in combination do not explicitly teach: wherein the scale factor comprises a ratio of an altitude of the UAV above ground level relative to a median of a monocular depth map, However, Wofk teaches recovering metric scale for each pixel in a depth map, where a depth value of each pixel is “a distance relative to the camera” (i.e., altitude of the UAV above ground level) (See at least [0040] & [0058]). Wofk further teaches an “SML network 165 learns per-pixel scale factors by which to multiply input depth estimates” (See at least [0113]). Since the scale is being recovered for each pixel, which would include any pixel that is a median of the depth map, the teachings of Wofk render obvious a scale factor that is based on an altitude of the UAV and a median of a monocular depth map, which provides the benefit of “improv[ing] metric accuracy” (See [0043] of Wofk). One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena, Schonberger, and Lin’s method with the teachings of Wofk discussed above. Doing so would be obvious to “improve metric accuracy” (See [0043] of Wofk). Abeywardena, Schonberger, Lin, and Wofk in combination do not explicitly teach: wherein the monocular depth map is based on the reference depth data. Guizilini ‘280 teaches: wherein the monocular depth map is based on the reference depth data. (See at least [0048]: “The semantic model 290 is, in one embodiment, a machine learning algorithm such as a convolutional neural network (CNN) or CNN-based deep neural network that accepts the monocular image 250 as an electronic input and generates the semantic features therefrom. In one or more aspects, the semantic model 290 is a Feature Pyramid Network (FPN) with a ResNet backbone. Accordingly, the semantic model 290 generally performs the process of semantic segmentation on the monocular image 250 to identify the components and boundaries of the components represented therein” & [0050]: “In any case, the semantic model 290 generates the semantic features according to the components (e.g., objects, surfaces, etc.) within the image 250, which intrinsically define boundaries between different aspects of the image 250 by, for example, associating individual pixels with respective components in the image 250. This distinction between boundaries of the different components provides knowledge about the locations of discontinuities (i.e., regions of changing depth) within the image 250, which the depth model 260 may otherwise experience difficulties in identifying. Consequently, injecting the semantic features into the depth model 260 provides for guiding determinations of the depth features with additional knowledge about the discontinuities, thereby avoiding the difficulties and improving prediction of depths from the monocular image 250.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena, Schonberger, Lin, and Wofk’s method with Guizilini ‘280’s monocular depth map being based on reference depth data. Doing so would be obvious “for guiding determinations of the depth features with additional knowledge about the discontinuities, thereby avoiding the difficulties and improving prediction of depths from the monocular image 250” (See [0051] of Guizilini ‘280). Regarding claim 15, Abeywardena, Schonberger, Lin, and Wofk in combination teach all the limitations of claim 13 as discussed above. Wofk additionally teaches: further comprising applying a scale factor to the depth image, (See at least [0040]: “The monocular depth estimator circuitry 125 predicts depth from a monocular image. For example, the monocular depth estimator circuitry 125 can include a pretrained model that takes in a single RGB image (e.g., RGB image 105) and produces a dense depth map up to a specified scale” & [0043]: “The scale aligner circuitry 150 can be used to perform dense (local) scale alignment given that global alignment may not adequately resolve metric scale in all regions of a depth map. As such, a learning-based approach can be used for determining dense (per-pixel) scale factors that are applied to globally aligned depth estimates. In the example of FIG. 1, a ScaleMapLearner (SML) network 165 can be trained (e.g., using an open-source machine learning framework such as MiDaS-small, etc.) to realign individual values in an input depth map to improve metric accuracy. In some examples, the SML network 165 can receive an input of two concatenated data channels, such as the globally aligned depth prediction 155 and/or a scaffolding for a dense scale map (e.g., dense scale map scaffolding 160 based on a metric sparse depth 145 output by the visual-inertial odometry sensor circuitry 130), as described in more detail in connection with FIG. 2. As such, the SML network 165 regresses the dense scale residual map 170 (e.g., dense scale map image 190 regressed by the SML network 165) and a resulting scale map can be generated and applied to the input depth, yielding an example final metric dense depth output 175 (e.g., shown visually using the final depth map output 192).”) Abeywardena, Schonberger, Lin, and Wofk in combination do not explicitly teach: wherein the scale factor comprises a ratio of an altitude of the UAV above ground level relative to an above ground level estimate from a monocular depth map, However, Wofk teaches recovering metric scale for each pixel in a depth map, where a depth value of each pixel is “a distance relative to the camera” (i.e., altitude of the UAV above ground level) (See at least [0040] & [0058]). Wofk further teaches an “SML network 165 learns per-pixel scale factors by which to multiply input depth estimates” (See at least [0113]). Since the scale is being recovered for each pixel, which would include any pixel above ground in the depth map, the teachings of Wofk render obvious a scale factor that is based on an altitude of the UAV and an above ground level estimate from a monocular depth map, which provides the benefit of “improv[ing] metric accuracy” (See [0043] of Wofk). One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena, Schonberger, and Lin’s method with the teachings of Wofk discussed above. Doing so would be obvious to “improve metric accuracy” (See [0043] of Wofk). Abeywardena, Schonberger, Lin, and Wofk in combination do not explicitly teach: wherein the monocular depth map is based on the reference depth data. Guizilini ‘280 teaches: wherein the monocular depth map is based on the reference depth data. (See at least [0048]: “The semantic model 290 is, in one embodiment, a machine learning algorithm such as a convolutional neural network (CNN) or CNN-based deep neural network that accepts the monocular image 250 as an electronic input and generates the semantic features therefrom. In one or more aspects, the semantic model 290 is a Feature Pyramid Network (FPN) with a ResNet backbone. Accordingly, the semantic model 290 generally performs the process of semantic segmentation on the monocular image 250 to identify the components and boundaries of the components represented therein” & [0050]: “In any case, the semantic model 290 generates the semantic features according to the components (e.g., objects, surfaces, etc.) within the image 250, which intrinsically define boundaries between different aspects of the image 250 by, for example, associating individual pixels with respective components in the image 250. This distinction between boundaries of the different components provides knowledge about the locations of discontinuities (i.e., regions of changing depth) within the image 250, which the depth model 260 may otherwise experience difficulties in identifying. Consequently, injecting the semantic features into the depth model 260 provides for guiding determinations of the depth features with additional knowledge about the discontinuities, thereby avoiding the difficulties and improving prediction of depths from the monocular image 250.”) One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena, Schonberger, Lin, and Wofk’s method with Guizilini ‘280’s monocular depth map being based on reference depth data. Doing so would be obvious “for guiding determinations of the depth features with additional knowledge about the discontinuities, thereby avoiding the difficulties and improving prediction of depths from the monocular image 250” (See [0051] of Guizilini ‘280). Claim(s) 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Abeywardena in view of Schonberger and further in view of Kojima of US 20230406536 A1, filed 01/13/2022, hereinafter “Kojima”. Regarding claim 18, Abeywardena and Schonberger in combination teach all the limitations of claim 1 as discussed above. Abeywardena and Schonberger in combination do not explicitly teach: further comprising applying a Kalman filter to the determined location of the UAV in the environment to control navigation of the UAV in the environment. Kojima teaches: further comprising applying a Kalman filter to the determined location of the UAV in the environment to control navigation of the UAV in the environment. (See at least [0077]: “In the position control system 100 for the aircraft 1 in FIG. 9, the second smoothing processing unit 43 is omitted, and based on the position difference between the estimated relative position and the target relative position estimated by the second Kalman filter 46, the velocity difference between the estimated relative velocity and the target relative velocity, and the attitude correction acceleration, the control quantity C′ is derived” & [0050]: “The second Kalman filter 46 performs an estimation based on the relative position (X, Y) and outputs an estimated relative position (X, Y) after the estimation. Specifically, the relative position (X, Y) calculated by the guidance calculation unit 34 is input to the second Kalman filter 46. Upon the input of the relative position (X, Y), the second Kalman filter 46 calculates the estimated relative position (X, Y) by estimating the change of the relative position (X, Y) over time. The second Kalman filter 46 outputs the calculated estimated relative position (X, Y) to the second smoothing processing unit 43.” See also [0076], which recites that the control quantity C’ is used to adjust operation of the aircraft.) NOTE: Claim 18 recites the following intended use limitation: “…to control navigation of the UAV in the environment”. This limitation is not positively recited, and is instead recited as intended use since it recites an intended result for “applying a Kalman filter to the determined location of the UAV in the environment”. Therefore, the BRI of claim 18 does not require the aforementioned intended use limitation. One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to combine Abeywardena and Schonberger’s method with Kojima’s technique of applying a Kalman filter to the determined location of the UAV in the environment to control navigation of the UAV in the environment. Doing so would be obvious since “With this configuration, the flight control of the aircraft 1 can be performed using the position difference between the target relative position and the smoothed relative position, and the velocity difference between the target relative velocity and the estimated relative velocity. This allows the aircraft 1 to accurately perform the target point following hovering with respect to the target landing point 2” (See [0087] of Kojima). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US 20180158197 A1 is directed to facilitate autonomous navigation by the UAV by tracking objects in an environment and extracting semantic cues regarding the tracked objects. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Nikki Molina whose telephone number is (571) 272-5180. The examiner can normally be reached Monday - Thursday and alternate Fridays, 7:30-4:30 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aniss Chad, can be reached on (571) 270-3832. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /NIKKI MARIE M MOLINA/Examiner, Art Unit 3662
Read full office action

Prosecution Timeline

May 06, 2025
Application Filed
Aug 10, 2026
Non-Final Rejection mailed — §103, §112, §DOUBLEPATENT (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12746914
METHOD AND APPARATUS FOR COLLISION AVOIDANCE OR IMPACT FORCE REDUCTION
2y 11m to grant Granted Sep 29, 2026
Patent 12728869
INFORMATION PROCESSING THAT EXCLUDES PERSONAL INFORMATION FROM META DATA
3y 7m to grant Granted Sep 08, 2026
Patent 12722638
METHOD OF CONTROLLING VEHICLE FOR ONE-PEDAL DRIVING ASSISTANCE
3y 9m to grant Granted Sep 01, 2026
Patent 12709176
ELECTRIC VEHICLE CHARGEABLE BY WIND ENERGY
1y 9m to grant Granted Aug 18, 2026
Patent 12687853
REMOTE SUPPORT APPARATUS
2y 2m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
82%
With Interview (+4.1%)
2y 7m (~1y 2m remaining)
Median Time to Grant
Low
PTA Risk
Based on 111 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month