Prosecution Insights
Last updated: September 17, 2026
Application No. 18/420,865

SPARSE FEATURE ENCODING OF MULTIMODAL DATA TO BUILD COMMONSENSE KNOWLEDGE ONTOLOGY SUPPORTING DEDUCTIVE REASONING SYSTEM

Non-Final OA §101§103
Filed
Jan 24, 2024
Priority
Sep 06, 2023 — provisional 63/580,725
Examiner
RHO, YONG DOO
Art Unit
Tech Center
Assignee
Through Sensing LLC
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
12 currently pending
Career history
4
Total Applications
across all art units

Statute-Specific Performance

§101
32.7%
-7.3% vs TC avg
§103
55.8%
+15.8% vs TC avg
§102
3.9%
-36.1% vs TC avg
§112
7.7%
-32.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 0 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS) submitted on 12/27/2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Status of Claims The present application is being examined under the claims filed on 1/24/2024. Claims 1-19 are rejected. Claims 1-19 are pending. Specification The specification filed on 1/24/2024 is acceptable for examination purposes. Drawings The drawings filed on 1/24/2024 are acceptable for examination purposes. Claim Objections Claim 12 is objected to because of the following informalities: In claim 12, lines 6-7, “each of the plurality of the plurality of salient features” should read “each of the plurality of salient features” Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding Claim 1, Step 1: Claim 1 is a method claim. Therefore, Claims 1-11 are directed to a process. Step 2A Prong 1: extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy (mental process – extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy may be performed manually by a user with the aid of pen and paper by observing/analyzing the one or more objects at one or more levels of spatial resolution or structural hierarchy and using a judgement to extract a plurality of salient features. See MPEP 2106.04(a)(2)(III)(C).) creating data structures that associate one or more essential characteristics with each of the plurality of salient features (mental process – creating data structures that associate one or more essential characteristics with each of the plurality of salient features may be performed manually by a user with the aid of pen and paper by observing/analyzing one or more essential characteristics with each of the plurality of salient features and using a judgement to create data structures associating one or more essential characteristics. See MPEP 2106.04(a)(2)(III)(C).) mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy recursively to capture the part- whole relationships of the objects and one or more object components (mental process – mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy recursively to capture the part- whole relationships of the objects and one or more object components may be performed manually by a user with the aid of pen and paper by observing/analyzing the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and using a judgement to map the spatial relationships among the one or more objects and capture the part-whole relationships of the objects and one or more object components. See MPEP 2106.04(a)(2)(III)(C).) building an ontology that encompasses the typical relationships among the classes of objects in a plurality of datasets (mental process – building an ontology that encompasses the typical relationships among the classes of objects in a plurality of datasets may be performed manually by a user with the aid of pen and paper by observing/analyzing the relationships among the classes of objects in a plurality of datasets and using a judgement to build an ontology that encompasses the typical relationships among the classes of objects in the datasets. See MPEP 2106.04(a)(2)(III)(C).) establishing one or more axioms that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples (mental process – establishing one or more axioms that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples may be performed manually by a user with the aid of pen and paper by observing/analyzing the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples and using a judgement to create one or more axioms that encode the relationships. See MPEP 2106.04(a)(2)(III)(C).) applying a deductive reasoning tool to the axioms of the ontology (mental process – applying a deductive reasoning tool to the axioms of the ontology may be performed manually by a user with the aid of pen and paper by applying a logical deduction and reasoning to the axioms of the ontology. See MPEP 2106.04(a)(2)(III)(C).) searching graph-based structures of the ontology (mental process – searching graph-based structures of the ontology may be performed manually by a user with the aid of pen and paper by observing/analyzing the relationships and the axioms of the ontology and using a judgement to search graph-based structures of the ontology. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: acquiring multimodal sensory data including one or more objects (Adding insignificant extra-solution activity to the judicial exception. See MPEP 2106.05(g).) outputting identifying information and relationship information on at least one of the one or more objects (Adding insignificant extra-solution activity to the judicial exception. See MPEP 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: acquiring multimodal sensory data including one or more objects (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) outputting identifying information and relationship information on at least one of the one or more objects (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2-11. The additional limitations of the dependent claims are addressed below. Regarding Claim 2, Step 2A Prong 1: answering the one or more queries by the deductive reasoning tool based on the outputted identifying information and relationship information (mental process – answering the one or more queries by the deductive reasoning tool based on the outputted identifying information and relationship information may be performed manually by a user with the aid of pen and paper by observing/analyzing the outputted identifying information and relationship information and using a judgement to make deductions and answer the one or more queries. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: receiving one or more queries (Adding insignificant extra-solution activity to the judicial exception. See MPEP 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: receiving one or more queries (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) Regarding Claim 3, Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 3 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the multimodal sensory data includes one or more of visible light images, video, sound recordings, radar signals, infrared and ultraviolet images, and physical measurements (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the multimodal sensory data includes one or more of visible light images, video, sound recordings, radar signals, infrared and ultraviolet images, and physical measurements (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 4, Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 4 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the plurality of salient features of the one or more objects are extracted via one or more of computer vision, image processing, or statistical methods (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the plurality of salient features of the one or more objects are extracted via one or more of computer vision, image processing, or statistical methods (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Regarding Claim 5, Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 5 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the one or more essential characteristics include at least one of color and material composition (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the one or more essential characteristics include at least one of color and material composition (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 6, Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 6 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the spatial relationship among the one or more objects is mapped utilizing a mathematical topology principle and the part-whole relationships of the objects are captured based on region adjacency graphs or scene graphs (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the spatial relationship among the one or more objects is mapped utilizing a mathematical topology principle and the part-whole relationships of the objects are captured based on region adjacency graphs or scene graphs (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 7, Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 7 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the deductive reasoning tool is an automated theorem prover (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the deductive reasoning tool is an automated theorem prover (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 8, Step 2A Prong 1: See the rejection of Claim 2 above, which Claim 8 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the one or more queries are addressed through a natural language interface (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) wherein the answer is determined by at least an artificial intelligence or machine learning device (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the one or more queries are addressed through a natural language interface (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) wherein the answer is determined by at least an artificial intelligence or machine learning device (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Regarding Claim 9, Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 9 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: applying the method to a mathematical word problem or other problem that requires visualization (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: applying the method to a mathematical word problem or other problem that requires visualization (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 10, Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 10 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: applying the method to an alternative Al method (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) constraining the alternative Al method to reduce or prevent data hallucination, bias, irrelevant conclusions, and/or Al misbehavior (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: applying the method to an alternative Al method (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) constraining the alternative Al method to reduce or prevent data hallucination, bias, irrelevant conclusions, and/or Al misbehavior (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 11, Step 2A Prong 1: See the rejection of Claim 10 above, which Claim 11 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the alternative Al method is a large language model or a small language model (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the alternative Al method is a large language model or a small language model (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 12, Step 1: Claim 12 is a system claim. Therefore, Claims 12-19 are directed to a machine. Step 2A Prong 1: [an extraction module] that extracts a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy (mental process – extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy may be performed manually by a user with the aid of pen and paper by observing/analyzing the one or more objects at one or more levels of spatial resolution or structural hierarchy and using a judgement to extract a plurality of salient features. See MPEP 2106.04(a)(2)(III)(C).) [a data structure module that associates one or more essential characteristics with each of the plurality of the plurality of salient features], creates data structures that associate one or more essential characteristics with each of the plurality of salient features (mental process – creating data structures that associate one or more essential characteristics with each of the plurality of salient features may be performed manually by a user with the aid of pen and paper by observing/analyzing one or more essential characteristics with each of the plurality of salient features and using a judgement to create data structures associating one or more essential characteristics. See MPEP 2106.04(a)(2)(III)(C).) maps the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy recursively to capture the part- whole relationships of the objects and one or more object components (mental process – mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy recursively to capture the part- whole relationships of the objects and one or more object components may be performed manually by a user with the aid of pen and paper by observing/analyzing the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and using a judgement to map the spatial relationships among the one or more objects and capture the part-whole relationships of the objects and one or more object components. See MPEP 2106.04(a)(2)(III)(C).) [an ontology module] that builds an ontology that encompasses the typical relationships among the classes of objects in a plurality of datasets (mental process – building an ontology that encompasses the typical relationships among the classes of objects in a plurality of datasets may be performed manually by a user with the aid of pen and paper by observing/analyzing the relationships among the classes of objects in a plurality of datasets and using a judgement to build an ontology that encompasses the typical relationships among the classes of objects in the datasets. See MPEP 2106.04(a)(2)(III)(C).) establishes one or more axioms that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples (mental process – establishing one or more axioms that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples may be performed manually by a user with the aid of pen and paper by observing/analyzing the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples and using a judgement to create one or more axioms that encode the relationships. See MPEP 2106.04(a)(2)(III)(C).) [a deductive reasoning tool] that is applied to the axioms of the ontology and searches graph- based structures of the ontology [to output identifying information and relationship information on at least one of the one or more objects] (mental process – applying a deductive reasoning tool to the axioms of the ontology and searching graph-based structures of the ontology may be performed manually by a user with the aid of pen and paper by applying a logical deduction/reasoning to the axioms of the ontology, and observing/analyzing the relationships and the axioms of the ontology, and using a judgement to search graph-based structures of the ontology. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: one or more sensors that capture multimodal sensory data (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) an extraction module (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) a data structure module that associates one or more essential characteristics with each of the plurality of the plurality of salient features (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) an ontology module (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) a deductive reasoning tool (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) to output identifying information and relationship information on at least one of the one or more objects (Adding insignificant extra-solution activity to the judicial exception. See MPEP 2106.05(g).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. The additional elements recite generic computer elements and programs at a high-level of generality to perform the judicial exception as well as recitation of generic computer functionality such as one or more sensors, an extraction module, a data structure module, an ontology module and a deductive reasoning tool. Additional Elements: one or more sensors that capture multimodal sensory data (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) an extraction module (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) a data structure module that associates one or more essential characteristics with each of the plurality of the plurality of salient features (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) an ontology module (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) a deductive reasoning tool (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) to output identifying information and relationship information on at least one of the one or more objects (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer) For the reasons above, Claim 12 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 13-19. The additional limitations of the dependent claims are addressed below. Regarding Claim 13, Step 2A Prong 1: [wherein the deductive reasoning tool further receives one or more queries] and answers the one or more queries based on the output identifying information and relationship information (mental process – answering the one or more queries by the deductive reasoning tool based on the output identifying information and relationship information may be performed manually by a user with the aid of pen and paper by observing/analyzing the output identifying information and relationship information and using a judgement to make deductions and answer the one or more queries. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the deductive reasoning tool further receives one or more queries (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the deductive reasoning tool further receives one or more queries (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Regarding Claim 14, Step 2A Prong 1: See the rejection of Claim 13 above, which Claim 14 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the multimodal sensory data includes one or more of visible light images, video, sound recordings, radar signals, infrared and ultraviolet images, and physical measurements (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the multimodal sensory data includes one or more of visible light images, video, sound recordings, radar signals, infrared and ultraviolet images, and physical measurements (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 15, Step 2A Prong 1: See the rejection of Claim 14 above, which Claim 15 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the plurality of salient features of the one or more objects are extracted via one or more of computer vision, image processing, or statistical methods (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the plurality of salient features of the one or more objects are extracted via one or more of computer vision, image processing, or statistical methods (merely using a computer as a tool to perform an abstract idea. See MPEP 2106.05(f).) Regarding Claim 16, Step 2A Prong 1: See the rejection of Claim 12 above, which Claim 16 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the one or more sensors are integrated into one of an unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the one or more sensors are integrated into one of an unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 17, Step 2A Prong 1: See the rejection of Claim 16 above, which Claim 17 depends on. Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is configured to navigate autonomously based on the output identifying information and relationship information (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is configured to navigate autonomously based on the output identifying information and relationship information (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 18, Step 2A Prong 1: [wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is configured to] use a sorting or search algorithm for the features in the ontology module for visual reasoning in order to obtain an understanding of geography of a region or environment (mental process – using a sorting or search algorithm for the features in the ontology module for visual reasoning in order to obtain an understanding of geography of a region or environment may be performed manually by a user with the aid of pen and paper by observing/analyzing the features in the ontology module for visual reasoning and using a sorting or search algorithm to understand geography of a region or environment. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is configured to (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is configured to (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Regarding Claim 19, Step 2A Prong 1: [wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is further configured to] determine the location of a perceived object or class of objects in geographical coordinates (mental process – determining the location of a perceived object or class of objects in geographical coordinates may be performed manually by a user with the aid of pen and paper by observing/analyzing the location of a perceived object or class of objects in geographical coordinates. See MPEP 2106.04(a)(2)(III)(C).) Step 2A Prong 2: The judicial exceptions are not integrated into a practical application. Additional Elements: wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is further configured to (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Step 2B: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. Additional Elements: wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is further configured to (merely reciting the words "apply it" (or an equivalent) with the judicial exception. See MPEP 2106.05(f).) Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 4-9, 12-13 and 15-19 are rejected under 35 U.S.C. 103 as being unpatentable over Golestan Irani (US 20200089251 A1), in view of Rosinol et al. (“3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans”) (hereinafter Rosinol), and further in view of Lassoued et al. (US 20200082016 A1) (hereinafter Lassoued). Regarding Claim 1, Golestan Irani teaches: “A method for answering queries by applying deductive reasoning using a data structure based on knowledge derived from sensory data, comprising:” (preamble) “acquiring multimodal sensory data including one or more objects” (Golestan Irani, Paragraph [0027], “In an example embodiment, the cameras 112, LiDAR units 114 and SAR units 116 are located at the front, rear, left side and right side of the vehicle 105 to capture data about the environment in front, rear, left side and right side of the vehicle 105. The cameras 112, LiDAR units 114 and SAR units 116 are mounted or otherwise located to have different fields of view (FOVs) or coverage areas to capture data about the environment surrounding the vehicle 105.”; Examiner’s note: acquiring multimodal sensory data including one or more objects (i.e., the cameras 112, LiDAR units 114 and SAR units 116 capturing data about the environment surrounding the vehicle 105) is taught.) “creating data structures that associate one or more essential characteristics with each of the plurality of salient features” (Golestan Irani, Paragraph [0070], “The output of SHD-ASC module 312 is a semantic point cloud map (e.g. a point cloud map that has been semantically labeled). The semantic point cloud map may comprise a point cloud map with an enhanced data layer that defines the semantic labels for various points. Alternatively, the enhanced data layer may be stored separate from the point cloud map. Each coordinate may have more than one label, and different labels for a particular coordinate may be generated in the same or different sessions. If a particular coordinate of the semantic point cloud map already has a label then a new label will be added to it. The SHD-ASC 312 stores the association between semantic label(s) and the matching coordinate(s) of the point cloud map in memory of the computer vision system 300.”; Examiner’s note: creating data structures that associate one or more essential characteristics with each of the plurality of salient features (i.e., creating the output of SHD-ASC module 312 which is a semantic point cloud map including semantic label(s) and the matching coordinate(s)) is taught.) “building an ontology that encompasses the typical relationships among the classes of objects in a plurality of datasets” (Golestan Irani, Paragraph [0044], “An AD ontology, also sometimes known as a knowledge graph, is used to associate the semantic data with the sensor data. The process of associating semantic data with sensor data may comprise fusing or combining the semantic data with the sensor data using techniques described below and herein […] The predefined AD ontology is used by the method and system of the present disclosure to generate a semantic point cloud map. The present disclosure uses a sensor fusion-based model to associate a set of data points of a raw point cloud map with a label that is semantically linked to the set of data points of the raw point cloud map through the sematic data (which is structured data) generated from soft data (which is unstructured data) received from a human. The sensor fusion-based model uses sensor fusion algorithms/techniques/tools such as ontologies to fuse the sensor data originating from different modalities.”; Examiner’s note: building an ontology that encompasses the typical relationships among the classes of objects in a plurality of datasets (i.e., the sensor fusion-based model using sensor fusion algorithms/techniques/tools to fuse the sensor data from different modalities) is taught.) “outputting identifying information and relationship information on at least one of the one or more objects” (Golestan Irani, Paragraph [0082], “The user interface screen 800 includes a plurality of user interface (UI) buttons (or boxes) 805, 815, 825, 835, 845 and 855 as well as a none/cancel button 860. Each of the UI buttons 805-855 includes potential attribute-values pairs based the respective likelihoods of match. The user need only touch the matching UI button 805-855 corresponding to a set of attribute-value pairs to select the intended attribute-value pairs if shown or touch the none/cancel button 860 if the intended attribute-value pairs are not shown.”; Examiner’s note: outputting identifying information and relationship information on at least one of the one or more objects (i.e., the user interface screen 800 including a plurality of UI buttons displaying potential attribute-values pairs based the respective likelihoods of match) is taught.) Golestan Irani does not explicitly teach: “extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy” “mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy recursively to capture the part-whole relationships of the objects and one or more object components” “establishing one or more axioms that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples” “applying a deductive reasoning tool to the axioms of the ontology” “searching graph-based structures of the ontology” Rosinol teaches: “extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy” (Rosinol, Fig. 1, PNG media_image1.png 556 702 media_image1.png Greyscale ; Rosinol, Section I, “We present a unified representation for actionable spatial perception: 3D Dynamic Scene Graphs (DSGs, Fig. 1). A DSG, introduced in Section III, is a layered directed graph where nodes represent spatial concepts (e.g., objects, rooms, agents) and edges represent pairwise spatiotemporal relations. The graph is layered, in that nodes are grouped into layers that correspond to different levels of abstraction of the scene (i.e., a DSG is a hierarchical representation). Our choice of nodes and edges in the DSG also captures places and their connectivity, hence providing a strict generalization of the notion of topological maps [85, 86] and making DSGs an actionable representation for navigation and planning.”; Examiner’s note: extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy (i.e., a layered directed graph where nodes represent spatial concepts (e.g., objects, rooms, agents) and edges represent pairwise spatiotemporal relations. Nodes are grouped into layers corresponding to different levels of the scene (i.e., a DSG is a hierarchical representation)) is taught.) “mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy recursively to capture the part-whole relationships of the objects and one or more object components” (Rosinol, Fig. 1 and Section III, “[…] a DSG is a layered directed graph where nodes represent spatial concepts (e.g., objects, rooms, agents) and edges represent pairwise spatio-temporal relations (e.g., “agent A is in room B at time t”) […] spatial concepts are semantic concepts that are spatially grounded (in other words, each node in our DSG includes spatial coordinates and shape or bounding-box information as attributes) […] A DSG is a layered graph, i.e., nodes are grouped into layers that correspond to different levels of abstraction […] Edges between objects describe relations, such as co-visibility, relative size, distance, or contact (“the cup is on the desk”).”; Examiner’s note: mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy (i.e., grouping nodes into layers corresponding to different levels of abstraction and mapping the edges representing pairwise spatio-temporal relations) recursively to capture the part-whole relationships of the objects and one or more object components (i.e., edges between objects describe relations, such as co-visibility, relative size, distance, or contact) is taught. See Fig. 1 in the above limitation of claim 1.) “searching graph-based structures of the ontology” (Rosinol, Section III and Section VI, “[…] a DSG is a layered directed graph where nodes represent spatial concepts (e.g., objects, rooms, agents) and edges represent pairwise spatio-temporal relations (e.g., “agent A is in room B at time t”) […] DSGs also provide a powerful tool for high-level planning queries. For instance, the (connected) subgraph of places and objects in a DSG can be used to issue the robot a high-level command (e.g., object search [38]), and the robot can directly infer the closest place in the DSG it has to reach to complete the task, and can plan a feasible path to that place.”; Examiner’s note: searching graph-based structures of the ontology (i.e., high-level planning queries (e.g., object search) in the DSG, a layered directed graph with nodes and edges) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the method and system for generating a semantic point cloud map in Golestan Irani, and the 3D dynamic scene graphs as taught in Rosinol. Golestan Irani teaches acquiring multimodal sensory data, creating data structures that associate one or more essential characteristics with each of the plurality of salient features, building an ontology that encompasses the relationships among the classes of objects in a plurality of datasets and outputting identifying information and relationship information. Rosinol teaches extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy, mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and searching graph-based structures of the ontology. One of ordinary skill would have motivation to combine Golestan Irani and Rosinol to “provide[] a representation for hierarchical planning and fast collision checking, provide[] an interpretable abstraction of the scene and enable[] data compression” (Rosinol, Section I). Lassoued teaches: “establishing one or more axioms that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples” (Lassoued, Paragraph [0014], “Various embodiments described herein provide a logic-based relationship graph extraction operation for extracting a graph of relationships between concepts from text data based on a user query (e.g., input query) according to a domain ontology including a set of logical rules (axioms) for logical reasoning.”; Examiner’s note: establishing one or more axioms (i.e., a set of logical rules (axioms)) that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples (i.e., extracting a graph relationships between concepts from text data based on a user query according to a domain ontology) is taught.) “applying a deductive reasoning tool to the axioms of the ontology” (Lassoued, Paragraph [0078], “[…] a theorem prover may be used to enable inferring new facts or relations between a subset of facts in the domain of logical rules. The theorem proving operation may be used to expand the user input query using the domain ontology and logical rules to request additional concepts and relationship, which may be combined to infer relationship statement relevant to the user.”; Examiner’s note: applying a deductive reasoning tool (i.e., using a theorem prover) to the axioms of the ontology (i.e., logical rules in the domain ontology) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the method and system for generating a semantic point cloud map in Golestan Irani, the 3D dynamic scene graphs in Rosinol and logic-based relationship graph expansion and extraction as taught in Lassoued. Golestan Irani teaches acquiring multimodal sensory data, creating data structures that associate one or more essential characteristics with each of the plurality of salient features, building an ontology that encompasses the relationships among the classes of objects in a plurality of datasets and outputting identifying information and relationship information. Rosinol teaches extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy, mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and searching graph-based structures of the ontology. Lassoued teaches establishing one or more axioms that encode the relationships of individual datasets/classes of objects from multiple data examples and applying a deductive reasoning tool to the axioms of the ontology. One of ordinary skill would have motivation to combine Golestan Irani, Rosinol and Lassoued to “expand[] the search space of a user input query, using logical rules and a domain ontology, to retrieve relevant results” (Lassoued, Paragraph [0014]). Regarding Claim 2, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 1, further comprising” (preamble) “receiving one or more queries” (Golestan Irani, Paragraph [0061], “The SD-EXT module 310 parses the semantic data by running multiple queries in the AD ontology 330, extracts a list of attribute values from the semantic data and determines a class corresponding to each attribute value to generate a list of attribute-value pairs. The subject and predicate terms are queried on the AD ontology 330 as if asking the AD ontology 330 “What is <word in subject or predicate>?” The responses can be considered as filters that will be applied to sensor data.”; Examiner’s note: receiving one or more queries (i.e., multiple queries in the AD ontology 330) is taught.) “answering the one or more queries by the deductive reasoning tool based on the outputted identifying information and relationship information” (Golestan Irani, Paragraph [0060], “The SD-VAL module 308 determines whether the semantic data is valid by comparing the words to the AD ontology 330. The SD-VAL module 308 determines whether the semantic data is valid by querying the AD ontology 330 on all the words included in the subject and predicate terms of the RDF triplet. If each word in the subject and predicate terms of the RDF triplet corresponds to an instance defined in AD ontology 330, the semantic data is valid”; Examiner’s note: answering the one or more queries (i.e., determining whether the semantic data is valid by querying the AD ontology) by the deductive reasoning tool (i.e., the SD-VAL module) based on the outputted identifying information and relationship information (i.e., each word in the subject and predicate terms of the RDF triplet corresponding to an instance defined in AD ontology) is taught.) The reasons of obviousness have been noted in the rejection of Claim 1 above and applicable herein. Regarding Claim 4, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 2,” (preamble) “wherein the plurality of salient features of the one or more objects are extracted via one or more of computer vision, image processing, or statistical methods” (Golestan Irani, Paragraphs [0041], [0046], [0070] and [0089], “The software modules 168 include a computer vision module 172, which in combination with the EM wave based sensors 110, provide a computer vision system 300 (FIG. 3), and other modules 176. The computer vision module 172 includes a semantic point cloud map generating (SPCMG) module 174 […] The modules of the SPCMG module 174 includes instructions that execute or run on one or more processing units or processor system 102, such as a CPU or GPU, or a combination thereof. The modules of the SPCMG module 174 comprise a point cloud map acquisition (PC-ACQ) module 302, a hard data acquisition (HD-ACQ) module 304, a soft data acquisition (SD-ACQ) module 306, soft data validation (SD-VAL) module 308, a soft data extraction (SD-EXT) module 310 and a soft data/hard data association (SHD-ASC) 312 […] The output of SHD-ASC module 312 is a semantic point cloud map (e.g. a point cloud map that has been semantically labeled). The semantic point cloud map may comprise a point cloud map with an enhanced data layer that defines the semantic labels for various points. Alternatively, the enhanced data layer may be stored separate from the point cloud map. Each coordinate may have more than one label, and different labels for a particular coordinate may be generated in the same or different sessions. If a particular coordinate of the semantic point cloud map already has a label then a new label will be added to it. The SHD-ASC 312 stores the association between semantic label(s) and the matching coordinate(s) of the point cloud map in memory of the computer vision system 300 […] one or more dedicated digital signal processors (DSPs), graphical processing units (GPU), or image processors may be used to perform some of the described operations.”; Examiner’s note: wherein the plurality of salient features of one or more objects (i.e., association between semantic label(s) and the matching coordinate(s) of the point cloud map) are extracted via one or more of computer vision (i.e., the computer vision system 300), image processing (i.e., DSPs, GPU, or image processors), or statistical methods (i.e., the SPCMG module including instructions that execute or run on one or more processing units or processor system 102) is taught.) The reasons of obviousness have been noted in the rejection of Claim 2 above and applicable herein. Regarding Claim 5, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 2,” (preamble) “wherein the one or more essential characteristics include at least one of color and material composition” (Golestan Irani, Paragraph [0069], “For example, referring to the example of FIG. 7B, when sensor data query passes the attribute-value pair (“color”, “red”) to the AD ontology 330, the sensor data query unit 702 may receive “Image” or “Camera” as the corresponding data source. This means only the red parts of the camera image need to be used for point cloud map coordinate matching.”; Golestan Irani, Fig. 7B, PNG media_image2.png 362 542 media_image2.png Greyscale ; Examiner’s note: wherein the one or more essential characteristics include at least one of color and material composition (i.e., attribute-value pairs, such as Color: Red, Type: Building, Label: School) is taught.) The reasons of obviousness have been noted in the rejection of Claim 2 above and applicable herein. Regarding Claim 6, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 2,” (preamble) “wherein the spatial relationship among the one or more objects is mapped utilizing a mathematical topology principle and the part-whole relationships of the objects are captured based on region adjacency graphs or scene graphs” (Golestan Irani, Paragraph [0070], “The filtered point cloud map coordinates output by the sensor data query unit 702 is received as input by a sensor/point cloud map (PCM) data association unit 704, and are matched with the primary point cloud map using a point cloud matching technique, such as the iterative closest point (ICP) algorithm which can be employed to minimize the difference between two clouds of points. Lastly, the particular semantic label(s) are associated with the matching points in the point cloud map for semantic data generated for the particular soft data observation. The output of SHD-ASC module 312 is a semantic point cloud map (e.g. a point cloud map that has been semantically labeled). The semantic point cloud map may comprise a point cloud map with an enhanced data layer that defines the semantic labels for various points.”; Examiner’s note: wherein the spatial relationship among the one or more objects (i.e., a semantic point cloud map teaches the spatial relationship, such as distance and adjacency, among the one or more objects) is mapped utilizing a mathematical topology principle (i.e., a point cloud matching technique, such as the iterative closest point (ICP) algorithm) and the part-whole relationships of the objects are captured based on region adjacency graphs (i.e., two clouds of points with the minimal distance) or scene graphs (i.e., the semantic point cloud map teaches 3D scene graphs) is taught.) The reasons of obviousness have been noted in the rejection of Claim 2 above and applicable herein. Regarding Claim 7, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 2,” (preamble) “wherein the deductive reasoning tool is an automated theorem prover” (Lassoued, Paragraph [0078], “[…] a theorem prover may be used to enable inferring new facts or relations between a subset of facts in the domain of logical rules. The theorem proving operation may be used to expand the user input query using the domain ontology and logical rules to request additional concepts and relationship, which may be combined to infer relationship statement relevant to the user.”; Examiner’s note: wherein the deductive reasoning tool is an automated theorem prover (i.e., a theorem prover) is taught.) The reasons of obviousness have been noted in the rejection of Claim 2 above and applicable herein. Regarding Claim 8, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 2,” (preamble) “wherein the one or more queries are addressed through a natural language interface” (Golestan Irani, Paragraph [0050], “The SD-ACQ module 306 receives speech or voice inputs via a human-machine interface device (HMD) 502 of the vehicle control system 115. The HMD 502 for receiving the speech or voice inputs may comprise one or more microphones 140 with the SD-ACQ module 306 of the computer vision system 300 receiving speech or voice inputs received by the one or more microphones 140.”; Examiner’s note: wherein the one or more queries are addressed through a natural language interface (i.e., receiving speech or voice inputs via a human-machine interface device (HMD) 502) is taught.) “wherein the answer is determined by at least an artificial intelligence or machine learning device” (Golestan Irani, Paragraph [0076], “At operation 408, the RDF decomposition module 506 of the SD-ACQ module 306 performs text decomposition upon the generated text (e.g. the soft data observation) to decompose the generated text into semantic data comprising keywords, such as an RDF triplet in the form of <subject, predicate, object> using NLP techniques that determine which part of the text is a subject, a predicate, and an object, discarding other words.”; Examiner’s note: wherein the answer is determined by at least an artificial intelligence or machine learning device (i.e., the RDF decomposition module decomposing the generated text into semantic data comprising keywords, such as an RDF triplet in the form of <subject, predicate, object> using NLP techniques) is taught.) The reasons of obviousness have been noted in the rejection of Claim 2 above and applicable herein. Regarding Claim 9, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 8,” (preamble) “applying the method to a mathematical word problem or other problem that requires visualization” (Golestan Irani, Paragraph [0072], “The method 400 is performed when a sematic point cloud map learning mode of the computer vision system 300 of the vehicle 105 is activated (or engaged) by an operator, such as a driver. The sematic point cloud map learning mode may be activated through interaction with a human-machine interface device (HMD) 502 (FIG. 5A) of the vehicle control system 115, such as voice activation via a pre-defined keyword combination or other user interaction, such as touch activation via a GUI of the computer vision system 300 displayed on the touchscreen 136.”; Examiner’s note: applying the method to a mathematical word problem or other problem that requires visualization (i.e., the method 400 being performed when a semantic point cloud map learning mode is activated through interaction with a human-machine interface device (HMD) 502, such as voice activation or touch activation on the touchscreen 136) is taught.) The reasons of obviousness have been noted in the rejection of Claim 8 above and applicable herein. Regarding Claim 12, Golestan Irani teaches: “A system for answering queries by applying deductive reasoning using a data structure based on knowledge derived from sensory data, comprising;” (preamble) “one or more sensors that capture multimodal sensory data” (Golestan Irani, Paragraph [0027], “In an example embodiment, the cameras 112, LiDAR units 114 and SAR units 116 are located at the front, rear, left side and right side of the vehicle 105 to capture data about the environment in front, rear, left side and right side of the vehicle 105. The cameras 112, LiDAR units 114 and SAR units 116 are mounted or otherwise located to have different fields of view (FOVs) or coverage areas to capture data about the environment surrounding the vehicle 105.”; Examiner’s note: one or more sensors that capture multimodal sensory data (i.e., the cameras 112, LiDAR units 114 and SAR units 116 capturing data about the environment surrounding the vehicle 105) is taught.) “a data structure module that associates one or more essential characteristics with each of the plurality of the plurality of salient features, creates data structures that associate one or more essential characteristics with each of the plurality of salient features” (Golestan Irani, Paragraph [0070], “The output of SHD-ASC module 312 is a semantic point cloud map (e.g. a point cloud map that has been semantically labeled). The semantic point cloud map may comprise a point cloud map with an enhanced data layer that defines the semantic labels for various points. Alternatively, the enhanced data layer may be stored separate from the point cloud map. Each coordinate may have more than one label, and different labels for a particular coordinate may be generated in the same or different sessions. If a particular coordinate of the semantic point cloud map already has a label then a new label will be added to it. The SHD-ASC 312 stores the association between semantic label(s) and the matching coordinate(s) of the point cloud map in memory of the computer vision system 300.”; Examiner’s note: a data structure module (i.e., SHD-ASC module 312) that associates one or more essential characteristics with each of the plurality of the plurality of salient features (i.e., a semantic point cloud map including semantic label(s) and the matching coordinate(s)), creates data structures that associate one or more essential characteristics with each of the plurality of salient features (i.e., creating the output of SHD-ASC module 312 which is a semantic point cloud map including semantic label(s) and the matching coordinate(s)) is taught.) “an ontology module that builds an ontology that encompasses the typical relationships among the classes of objects in a plurality of datasets” (Golestan Irani, Paragraph [0044], “An AD ontology, also sometimes known as a knowledge graph, is used to associate the semantic data with the sensor data. The process of associating semantic data with sensor data may comprise fusing or combining the semantic data with the sensor data using techniques described below and herein […] The predefined AD ontology is used by the method and system of the present disclosure to generate a semantic point cloud map. The present disclosure uses a sensor fusion-based model to associate a set of data points of a raw point cloud map with a label that is semantically linked to the set of data points of the raw point cloud map through the sematic data (which is structured data) generated from soft data (which is unstructured data) received from a human. The sensor fusion-based model uses sensor fusion algorithms/techniques/tools such as ontologies to fuse the sensor data originating from different modalities.”; Examiner’s note: an ontology module that builds an ontology that encompasses the typical relationships among the classes of objects in a plurality of datasets (i.e., the sensor fusion-based model using sensor fusion algorithms/techniques/tools to fuse the sensor data from different modalities) is taught.) “to output identifying information and relationship information on at least one of the one or more objects” (Golestan Irani, Paragraph [0082], “The user interface screen 800 includes a plurality of user interface (UI) buttons (or boxes) 805, 815, 825, 835, 845 and 855 as well as a none/cancel button 860. Each of the UI buttons 805-855 includes potential attribute-values pairs based the respective likelihoods of match. The user need only touch the matching UI button 805-855 corresponding to a set of attribute-value pairs to select the intended attribute-value pairs if shown or touch the none/cancel button 860 if the intended attribute-value pairs are not shown.”; Examiner’s note: outputting identifying information and relationship information on at least one of the one or more objects (i.e., the user interface screen 800 including a plurality of UI buttons displaying potential attribute-values pairs based the respective likelihoods of match) is taught.) Golestan Irani does not explicitly teach: “an extraction module that extracts a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy” “a data structure module that maps the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy recursively to capture the part-whole relationships of the objects and one or more object components” “an ontology module that establishes one or more axioms that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples” “a deductive reasoning tool that is applied to the axioms of the ontology” “a deductive reasoning tool that searches graph-based structures of the ontology” Rosinol teaches: “an extraction module that extracts a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy” (Rosinol, Fig. 1, PNG media_image1.png 556 702 media_image1.png Greyscale ; Rosinol, Section I, “We present a unified representation for actionable spatial perception: 3D Dynamic Scene Graphs (DSGs, Fig. 1). A DSG, introduced in Section III, is a layered directed graph where nodes represent spatial concepts (e.g., objects, rooms, agents) and edges represent pairwise spatiotemporal relations. The graph is layered, in that nodes are grouped into layers that correspond to different levels of abstraction of the scene (i.e., a DSG is a hierarchical representation). Our choice of nodes and edges in the DSG also captures places and their connectivity, hence providing a strict generalization of the notion of topological maps [85, 86] and making DSGs an actionable representation for navigation and planning.”; Examiner’s note: an extraction module (i.e., a Spatial Perception eNgine (SPIN)) that extracts a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy (i.e., a layered directed graph where nodes represent spatial concepts (e.g., objects, rooms, agents) and edges represent pairwise spatiotemporal relations. Nodes are grouped into layers corresponding to different levels of the scene (i.e., a DSG is a hierarchical representation)) is taught.) “a data structure module that maps the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy recursively to capture the part-whole relationships of the objects and one or more object components” (Rosinol, Fig. 1 and Section III, “[…] a DSG is a layered directed graph where nodes represent spatial concepts (e.g., objects, rooms, agents) and edges represent pairwise spatio-temporal relations (e.g., “agent A is in room B at time t”) […] spatial concepts are semantic concepts that are spatially grounded (in other words, each node in our DSG includes spatial coordinates and shape or bounding-box information as attributes) […] A DSG is a layered graph, i.e., nodes are grouped into layers that correspond to different levels of abstraction […] Edges between objects describe relations, such as co-visibility, relative size, distance, or contact (“the cup is on the desk”).”; Examiner’s note: a data structure module (i.e., a Spatial Perception eNgine (SPIN)) that maps the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy (i.e., grouping nodes into layers corresponding to different levels of abstraction and mapping the edges representing pairwise spatio-temporal relations) recursively to capture the part-whole relationships of the objects and one or more object components (i.e., edges between objects describe relations, such as co-visibility, relative size, distance, or contact) is taught. See Fig. 1 in the above limitation of claim 12.) “a deductive reasoning tool that searches graph-based structures of the ontology” (Rosinol, Section III and Section VI, “[…] a DSG is a layered directed graph where nodes represent spatial concepts (e.g., objects, rooms, agents) and edges represent pairwise spatio-temporal relations (e.g., “agent A is in room B at time t”) […] DSGs also provide a powerful tool for high-level planning queries. For instance, the (connected) subgraph of places and objects in a DSG can be used to issue the robot a high-level command (e.g., object search [38]), and the robot can directly infer the closest place in the DSG it has to reach to complete the task, and can plan a feasible path to that place.”; Examiner’s note: a deductive reasoning tool that searches graph-based structures of the ontology (i.e., a powerful tool for high-level planning queries (e.g., object search) in the DSG, a layered directed graph with nodes and edges) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the method and system for generating a semantic point cloud map in Golestan Irani, and the 3D dynamic scene graphs as taught in Rosinol. Golestan Irani teaches acquiring multimodal sensory data, creating data structures that associate one or more essential characteristics with each of the plurality of salient features, building an ontology that encompasses the relationships among the classes of objects in a plurality of datasets and outputting identifying information and relationship information. Rosinol teaches extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy, mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and searching graph-based structures of the ontology. One of ordinary skill would have motivation to combine Golestan Irani and Rosinol to “provide[] a representation for hierarchical planning and fast collision checking, provide[] an interpretable abstraction of the scene and enable[] data compression” (Rosinol, Section I). Lassoued teaches: “an ontology module that establishes one or more axioms that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples” (Lassoued, Paragraph [0014], “Various embodiments described herein provide a logic-based relationship graph extraction operation for extracting a graph of relationships between concepts from text data based on a user query (e.g., input query) according to a domain ontology including a set of logical rules (axioms) for logical reasoning.”; Examiner’s note: an ontology module (i.e., a domain ontology) that establishes one or more axioms (i.e., a set of logical rules (axioms)) that encode the specific relationships of individual datasets and the general relationships of classes of objects from multiple data examples (i.e., extracting a graph relationships between concepts from text data based on a user query according to a domain ontology) is taught.) “a deductive reasoning tool that is applied to the axioms of the ontology” (Lassoued, Paragraph [0078], “[…] a theorem prover may be used to enable inferring new facts or relations between a subset of facts in the domain of logical rules. The theorem proving operation may be used to expand the user input query using the domain ontology and logical rules to request additional concepts and relationship, which may be combined to infer relationship statement relevant to the user.”; Examiner’s note: a deductive reasoning tool that is applied (i.e., using a theorem prover) to the axioms of the ontology (i.e., logical rules in the domain ontology) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the method and system for generating a semantic point cloud map in Golestan Irani, the 3D dynamic scene graphs in Rosinol and logic-based relationship graph expansion and extraction as taught in Lassoued. Golestan Irani teaches acquiring multimodal sensory data, creating data structures that associate one or more essential characteristics with each of the plurality of salient features, building an ontology that encompasses the relationships among the classes of objects in a plurality of datasets and outputting identifying information and relationship information. Rosinol teaches extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy, mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and searching graph-based structures of the ontology. Lassoued teaches establishing one or more axioms that encode the relationships of individual datasets/classes of objects from multiple data examples and applying a deductive reasoning tool to the axioms of the ontology. One of ordinary skill would have motivation to combine Golestan Irani, Rosinol and Lassoued to “expand[] the search space of a user input query, using logical rules and a domain ontology, to retrieve relevant results” (Lassoued, Paragraph [0014]). Regarding Claim 13, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The system of claim 12,” (preamble) “wherein the deductive reasoning tool further receives one or more queries” (Golestan Irani, Paragraph [0061], “The SD-EXT module 310 parses the semantic data by running multiple queries in the AD ontology 330, extracts a list of attribute values from the semantic data and determines a class corresponding to each attribute value to generate a list of attribute-value pairs. The subject and predicate terms are queried on the AD ontology 330 as if asking the AD ontology 330 “What is <word in subject or predicate>?” The responses can be considered as filters that will be applied to sensor data.”; Examiner’s note: wherein the deductive reasoning tool (i.e., the SD-EXT module 310) further receives one or more queries (i.e., multiple queries in the AD ontology 330) is taught.) “wherein the deductive reasoning tool answers the one or more queries based on the output identifying information and relationship information” (Golestan Irani, Paragraph [0060], “The SD-VAL module 308 determines whether the semantic data is valid by comparing the words to the AD ontology 330. The SD-VAL module 308 determines whether the semantic data is valid by querying the AD ontology 330 on all the words included in the subject and predicate terms of the RDF triplet. If each word in the subject and predicate terms of the RDF triplet corresponds to an instance defined in AD ontology 330, the semantic data is valid”; Examiner’s note: wherein the deductive reasoning tool (i.e., the SD-VAL module 308) answers the one or more queries (i.e., determining whether the semantic data is valid by querying the AD ontology) based on the output identifying information and relationship information (i.e., each word in the subject and predicate terms of the RDF triplet corresponding to an instance defined in AD ontology) is taught.) The reasons of obviousness have been noted in the rejection of Claim 12 above and applicable herein. Regarding Claim 15, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The system of claim 14,” (preamble) “wherein the plurality of salient features of the one or more objects are extracted via one or more of computer vision, image processing, or statistical methods” (Golestan Irani, Paragraphs [0041], [0046], [0070] and [0089], “The software modules 168 include a computer vision module 172, which in combination with the EM wave based sensors 110, provide a computer vision system 300 (FIG. 3), and other modules 176. The computer vision module 172 includes a semantic point cloud map generating (SPCMG) module 174 […] The modules of the SPCMG module 174 includes instructions that execute or run on one or more processing units or processor system 102, such as a CPU or GPU, or a combination thereof. The modules of the SPCMG module 174 comprise a point cloud map acquisition (PC-ACQ) module 302, a hard data acquisition (HD-ACQ) module 304, a soft data acquisition (SD-ACQ) module 306, soft data validation (SD-VAL) module 308, a soft data extraction (SD-EXT) module 310 and a soft data/hard data association (SHD-ASC) 312 […] The output of SHD-ASC module 312 is a semantic point cloud map (e.g. a point cloud map that has been semantically labeled). The semantic point cloud map may comprise a point cloud map with an enhanced data layer that defines the semantic labels for various points. Alternatively, the enhanced data layer may be stored separate from the point cloud map. Each coordinate may have more than one label, and different labels for a particular coordinate may be generated in the same or different sessions. If a particular coordinate of the semantic point cloud map already has a label then a new label will be added to it. The SHD-ASC 312 stores the association between semantic label(s) and the matching coordinate(s) of the point cloud map in memory of the computer vision system 300 […] one or more dedicated digital signal processors (DSPs), graphical processing units (GPU), or image processors may be used to perform some of the described operations.”; Examiner’s note: wherein the plurality of salient features of one or more objects (i.e., association between semantic label(s) and the matching coordinate(s) of the point cloud map) are extracted via one or more of computer vision (i.e., the computer vision system 300), image processing (i.e., DSPs, GPU, or image processors), or statistical methods (i.e., the SPCMG module including instructions that execute or run on one or more processing units or processor system 102) is taught.) The reasons of obviousness have been noted in the rejection of Claim 14 above and applicable herein. Regarding Claim 16, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The system of claim 12,” (preamble) “wherein the one or more sensors are integrated into one of an unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle” (Golestan Irani, Paragraph [0025], “For convenience, the present disclosure describes example embodiments of methods and systems with reference to a motor vehicle, such as a car, truck, bus, boat or ship, submarine, aircraft, warehouse equipment, construction equipment, tractor or other farm equipment. The teachings of the present disclosure are not limited to any particular type of vehicle, and may be applied to vehicles that do not carry passengers as well as vehicles that do carry passengers. The teachings of the present disclosure may also be implemented in mobile robot vehicles including, but not limited to, autonomous vacuum cleaners, rovers, lawn mowers, unmanned aerial vehicle (UAV), and other objects.”; Examiner’s note: wherein the one or more sensors are integrated into one of an unmanned ground vehicle (i.e., vehicles that do not carry passengers), unmanned aerial vehicle (i.e., unmanned aerial vehicle (UAV)), or unmanned underwater vehicle (i.e., submarine that do not carry passengers) is taught.) The reasons of obviousness have been noted in the rejection of Claim 12 above and applicable herein. Regarding Claim 17, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The system of claim 16,” (preamble) “wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is configured to navigate autonomously based on the output identifying information and relationship information” (Golestan Irani, Paragraphs [0009], [0025] and [0026], “The semantic point cloud map generated by the present disclosure provides a higher-level dimension than merely 3D spatial coordinates, and enhances an autonomous vehicle's understanding of the current environment, which leads to better vehicle localization in cases where usual low-level features are of little assistance (e.g., indoor areas such as covered parking lots). The present disclosure also allows customization of a point cloud map based on the semantic information including semantic labels which may be helpful in adapting the autonomous vehicle for a particular user population, such as a particular demographics which may vary based on localization or consumer/brand preferences. Lastly, the semantic point cloud map improves the understanding of the environment, which can further be used by other modules of autonomous vehicles, such as perception or planning […] For convenience, the present disclosure describes example embodiments of methods and systems with reference to a motor vehicle, such as a car, truck, bus, boat or ship, submarine, aircraft, warehouse equipment, construction equipment, tractor or other farm equipment. The teachings of the present disclosure are not limited to any particular type of vehicle, and may be applied to vehicles that do not carry passengers as well as vehicles that do carry passengers. The teachings of the present disclosure may also be implemented in mobile robot vehicles including, but not limited to, autonomous vacuum cleaners, rovers, lawn mowers, unmanned aerial vehicle (UAV), and other objects […] The vehicle control system 115 can in various embodiments allow the vehicle 105 to be operable in one or more of a fully-autonomous, semi-autonomous or fully user-controlled mode.”; Examiner’s note: wherein the unmanned ground vehicle (i.e., vehicles that do not carry passengers), unmanned aerial vehicle (i.e., unmanned aerial vehicle (UAV)), or unmanned underwater vehicle (i.e., submarine that do not carry passengers) is configured to navigate autonomously (i.e., the vehicle control system 115 allowing the vehicle 105 to be operable in a fully-autonomous or semi-autonomous mode) based on the output identifying information and relationship information (i.e., understanding of the current environment leading to better vehicle localization and semantic information including semantic labels) is taught.) The reasons of obviousness have been noted in the rejection of Claim 16 above and applicable herein. Regarding Claim 18, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The system of claim 17,” (preamble) “wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is configured to use a sorting or search algorithm for the features in the ontology module for visual reasoning in order to obtain an understanding of geography of a region or environment” (Golestan Irani, Paragraphs [0009], [0025], [0069] and [0070], “The semantic point cloud map generated by the present disclosure provides a higher-level dimension than merely 3D spatial coordinates, and enhances an autonomous vehicle's understanding of the current environment, which leads to better vehicle localization in cases where usual low-level features are of little assistance (e.g., indoor areas such as covered parking lots) […] Lastly, the semantic point cloud map improves the understanding of the environment, which can further be used by other modules of autonomous vehicles, such as perception or planning […] For convenience, the present disclosure describes example embodiments of methods and systems with reference to a motor vehicle, such as a car, truck, bus, boat or ship, submarine, aircraft, warehouse equipment, construction equipment, tractor or other farm equipment. The teachings of the present disclosure are not limited to any particular type of vehicle, and may be applied to vehicles that do not carry passengers as well as vehicles that do carry passengers. The teachings of the present disclosure may also be implemented in mobile robot vehicles including, but not limited to, autonomous vacuum cleaners, rovers, lawn mowers, unmanned aerial vehicle (UAV), and other objects […] The sensor data query unit 702 receives a data source and scope identification and applies this data as criteria to filter the input sensor data (i.e., LiDAR data) to localize the coordinate(s) to be associated with the label(s) […] The filtered point cloud map coordinates output by the sensor data query unit 702 is received as input by a sensor/point cloud map (PCM) data association unit 704, and are matched with the primary point cloud map using a point cloud matching technique, such as the iterative closest point (ICP) algorithm which can be employed to minimize the difference between two clouds of points.”; Examiner’s note: wherein the unmanned ground vehicle (i.e., vehicles that do not carry passengers), unmanned aerial vehicle (i.e., unmanned aerial vehicle (UAV)), or unmanned underwater vehicle (i.e., submarine that do not carry passengers) is configured to use a sorting (i.e., filtering the input sensor data) or search algorithm for the features in the ontology module for visual reasoning (i.e., the iterative closest point (ICP) algorithm to minimize the difference between two clouds of points) in order to obtain an understanding of geography of a region or environment (i.e., semantic point cloud map improving the understanding of the environment leading to better vehicle localization) is taught.) The reasons of obviousness have been noted in the rejection of Claim 17 above and applicable herein. Regarding Claim 19, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The system of claim 18,” (preamble) “wherein the unmanned ground vehicle, unmanned aerial vehicle, or unmanned underwater vehicle is further configured to determine the location of a perceived object or class of objects in geographical coordinates” (Golestan Irani, Paragraphs [0009], [0025], [0069] and [0070], “The semantic point cloud map generated by the present disclosure provides a higher-level dimension than merely 3D spatial coordinates, and enhances an autonomous vehicle's understanding of the current environment, which leads to better vehicle localization in cases where usual low-level features are of little assistance (e.g., indoor areas such as covered parking lots) […] Lastly, the semantic point cloud map improves the understanding of the environment, which can further be used by other modules of autonomous vehicles, such as perception or planning […] For convenience, the present disclosure describes example embodiments of methods and systems with reference to a motor vehicle, such as a car, truck, bus, boat or ship, submarine, aircraft, warehouse equipment, construction equipment, tractor or other farm equipment. The teachings of the present disclosure are not limited to any particular type of vehicle, and may be applied to vehicles that do not carry passengers as well as vehicles that do carry passengers. The teachings of the present disclosure may also be implemented in mobile robot vehicles including, but not limited to, autonomous vacuum cleaners, rovers, lawn mowers, unmanned aerial vehicle (UAV), and other objects […] The sensor data query unit 702 receives a data source and scope identification and applies this data as criteria to filter the input sensor data (i.e., LiDAR data) to localize the coordinate(s) to be associated with the label(s) […] The filtered point cloud map coordinates output by the sensor data query unit 702 is received as input by a sensor/point cloud map (PCM) data association unit 704, and are matched with the primary point cloud map using a point cloud matching technique, such as the iterative closest point (ICP) algorithm which can be employed to minimize the difference between two clouds of points.”; Examiner’s note: wherein the unmanned ground vehicle (i.e., vehicles that do not carry passengers), unmanned aerial vehicle (i.e., unmanned aerial vehicle (UAV)), or unmanned underwater vehicle (i.e., submarine that do not carry passengers) is further configured to determine the location of a perceived object or class of objects in geographical coordinates (i.e., receiving a data source, scoping identification and filtering input sensor data to localize the coordinates to be associated with the labels) is taught.) The reasons of obviousness have been noted in the rejection of Claim 18 above and applicable herein. Claims 3 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Golestan Irani, in view of Rosinol, and further in view of Lassoued as applied in claim 1, and further in view of Hang et al. (JP 2018530181 A) (hereinafter Hang). Regarding Claim 3, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 2,” (preamble) wherein the multimodal sensory data includes one or more of visible light images, video, sound recordings, radar signals, and physical measurements (Golestan Irani, Paragraphs [0027], [0028] and [0050], “EM wave based sensors 110 may for example include digital cameras 112 that provide a computer vision system, light detection and ranging (LiDAR) units 114, and radar units such as synthetic aperture radar (SAR) units 116 […] Vehicle sensors 111 can include inertial measurement unit (IMU) 118, an electronic compass 119, and other vehicle sensors 120 such as a speedometer, a tachometer, wheel traction sensor, transmission gear sensor, throttle and brake position sensors, and steering angle sensor […] The HMD 502 for receiving the speech or voice inputs may comprise one or more microphones 140 with the SD-ACQ module 306 of the computer vision system 300 receiving speech or voice inputs received by the one or more microphones 140.”; Examiner’s note: wherein multimodal sensory data includes one or more visible light images, video (i.e., digital cameras 112 provides visible light images and video), sound recordings (i.e., speech or voice inputs received by the one or more microphones 140), radar signals (i.e., EM wave based sensors including LiDAR units 114 and radar units such as SAR units 116), and physical measurements (i.e., vehicle sensors such as IMU 118, an electronic compass 119, other vehicle sensors 120 such as a speedometer, a tachometer, wheel traction sensor, transmission gear sensor, etc.) is taught.) The combination of Golestan Irani, Rosinol and Lassoued does not explicitly teach: wherein the multimodal sensory data includes one or more of infrared and ultraviolet images Hang teaches: wherein the multimodal sensory data includes one or more of infrared and ultraviolet images (Hang, Page 5, Lines 34-35, “an electromagnetic radiation sensor configured to generate one or more ultraviolet (UV) data and infrared (IR) data associated with the image”; Examiner’s note: wherein the multimodal sensory data includes one or more of infrared and ultraviolet images (i.e., an electromagnetic radiation sensor generating one or more ultraviolet (UV) data and infrared (IR) data associated with the image) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the method and system for generating a semantic point cloud map in Golestan Irani, the 3D dynamic scene graphs in Rosinol, logic-based relationship graph expansion and extraction in Lassoued and auto white balance using infrared and/or ultraviolet signals as taught in Hang. Golestan Irani teaches acquiring multimodal sensory data, creating data structures that associate one or more essential characteristics with each of the plurality of salient features, building an ontology that encompasses the relationships among the classes of objects in a plurality of datasets, outputting identifying information and relationship information and the multimodal sensory data including one or more of visible light images, video, sound recordings, radar signals, and physical measurements. Rosinol teaches extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy, mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and searching graph-based structures of the ontology. Lassoued teaches establishing one or more axioms that encode the relationships of individual datasets/classes of objects from multiple data examples and applying a deductive reasoning tool to the axioms of the ontology. Hang teaches the multimodal sensory data including one or more of infrared and ultraviolet images. One of ordinary skill would have motivation to combine Golestan Irani, Rosinol, Lassoued and Hang to “improve auto white balance (AWB) with ultraviolet (UV) data 106 and/or infrared (IR) data 108” (Hang, Page 7, Lines 33-34). Regarding Claim 14, The combination of Golestan Irani, Rosinol, Lassoued and Hang teaches: “The system of claim 13,” (preamble) wherein the multimodal sensory data includes one or more of visible light images, video, sound recordings, radar signals, infrared and ultraviolet images, and physical measurements (Golestan Irani, Paragraphs [0027], [0028] and [0050], “EM wave based sensors 110 may for example include digital cameras 112 that provide a computer vision system, light detection and ranging (LiDAR) units 114, and radar units such as synthetic aperture radar (SAR) units 116 […] Vehicle sensors 111 can include inertial measurement unit (IMU) 118, an electronic compass 119, and other vehicle sensors 120 such as a speedometer, a tachometer, wheel traction sensor, transmission gear sensor, throttle and brake position sensors, and steering angle sensor […] The HMD 502 for receiving the speech or voice inputs may comprise one or more microphones 140 with the SD-ACQ module 306 of the computer vision system 300 receiving speech or voice inputs received by the one or more microphones 140.”; Hang, Page 5, Lines 34-35, “an electromagnetic radiation sensor configured to generate one or more ultraviolet (UV) data and infrared (IR) data associated with the image”; Examiner’s note: wherein multimodal sensory data includes one or more visible light images, video (i.e., digital cameras 112 provides visible light images and video), sound recordings (i.e., speech or voice inputs received by the one or more microphones 140), radar signals (i.e., EM wave based sensors including LiDAR units 114 and radar units such as SAR units 116), infrared and ultraviolet images (i.e., an electromagnetic radiation sensor generating one or more ultraviolet (UV) data and infrared (IR) data associated with the image), and physical measurements (i.e., vehicle sensors such as IMU 118, an electronic compass 119, other vehicle sensors 120 such as a speedometer, a tachometer, wheel traction sensor, transmission gear sensor, etc.) is taught.) The reasons of obviousness have been noted in the rejection of Claim 13 above and applicable herein. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Golestan Irani, in view of Rosinol, and further in view of Lassoued as applied in claim 1, and further in view of Steiner et al. (US 20180012133 A1) (hereinafter Steiner). Regarding Claim 10, The combination of Golestan Irani, Rosinol and Lassoued teaches: “The method of claim 8,” (preamble) The combination of Golestan Irani, Rosinol and Lassoued does not explicitly teach: “applying the method to an alternative Al method” “constraining the alternative Al method to reduce or prevent data hallucination, bias, irrelevant conclusions, and/or Al misbehavior” Steiner teaches: “applying the method to an alternative Al method” (Steiner, Paragraph [0053], “[…] the inventive method respectively the inventive system can be used to supervise AI-based systems with are only acting in an informational space with no direct mechanical actors like mechanical extremities, such like trading platforms or computers in economic areas or other informational entities, like chat-bot or other AI driven software based decision engines.”; Examiner’s note: applying the method to an alternative AI method (i.e., inventive method being used to supervise AI-based systems, like chat-bot or other AI driven software based decision engines) is taught.) “constraining the alternative Al method to reduce or prevent data hallucination, bias, irrelevant conclusions, and/or Al misbehavior” (Steiner, Paragraphs [0020]-[0021], [0031]-[0033], [0036]-[0037] and [0056], “e) detect deviations of said path (which is related to the behavior of said AI-based system); f) in the case of a deviation of the behavior of said AI-based system generate an information about the deviating behavior […] a) creating an empty behavioral profile of AI behavior; b) in a learning phase gathering of external accessible data of the AI behavior of a AI-based system, these might include mechanical accessible data like position, direction, speed, velocity and the like; c) training the AI behavioral profile using said gathered external accessible data of the AI-based system […] f) if newly captured AI behavior is substantial (explicit defined or more than 20% relative) different to already captured AI behavior than: g) inform a supervising authority, these supervising authority may be a third party related to said AI-based system by a contract […] detecting of biased decision-making in relation to a larger set of similar supervised AI-based system”; Examiner’s note: constraining the alternative AI method (i.e., creating an behavioral profile of AI behavior and training the behavioral profile using the gathered accessible data of the AI-based system teach the alternative AI method is constrained with the AI behavioral profile of AI behavior) to reduce or prevent data hallucination (i.e., detecting deviations of said path teaches reducing/preventing data being deviated from the path), bias (i.e., detecting of biased decision-making), irrelevant conclusions (i.e., deviation of the behavior of said AI-based system generating an information about the deviating behavior), and/or AI misbehavior (i.e., newly captured AI behavior different to already captured AI behavior) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the method and system for generating a semantic point cloud map in Golestan Irani, the 3D dynamic scene graphs in Rosinol, logic-based relationship graph expansion and extraction in Lassoued and the method and system for behavior control of AI-based systems as taught in Steiner. Golestan Irani teaches acquiring multimodal sensory data, creating data structures that associate one or more essential characteristics with each of the plurality of salient features, building an ontology that encompasses the relationships among the classes of objects in a plurality of datasets, outputting identifying information and relationship information and the multimodal sensory data including one or more of visible light images, video, sound recordings, radar signals, and physical measurements. Rosinol teaches extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy, mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and searching graph-based structures of the ontology. Lassoued teaches establishing one or more axioms that encode the relationships of individual datasets/classes of objects from multiple data examples and applying a deductive reasoning tool to the axioms of the ontology. Steiner teaches applying the method to an alternative Al method and constraining the alternative Al method to reduce or prevent data hallucination, bias, irrelevant conclusions, and/or Al misbehavior. One of ordinary skill would have motivation to combine Golestan Irani, Rosinol, Lassoued and Steiner so that “each deviation of the usual behavior of said AI-based system is detectable, particularly unexpected or unwanted behavior, by an independent supervising authority” (Steiner, Paragraph [0022]). Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Golestan Irani, in view of Rosinol, and further in view of Lassoued, and further in view of Steiner as applied in claim 10, and further in view of Wang et al. (CN 116595130 A) (hereinafter Wang). Regarding Claim 11, The combination of Golestan Irani, Rosinol, Lassoued and Steiner teaches: “The method of claim 10,” (preamble) The combination of Golestan Irani, Rosinol, Lassoued and Steiner does not explicitly teach: “wherein the alternative Al method is a large language model or a small language model” Wang teaches: “wherein the alternative Al method is a large language model or a small language model” (Wang, Page 3, Lines 26-31, “[…] a large language model and a small language model are obtained, wherein the model scale of the large language model is larger than the model scale of the small language model; Pre-training; based on a variety of natural language tasks, multi-task training is performed on the pre-trained large language model; the large language model after multi-task training is used as a teacher model, and the small language model after pre-training is used as a student model.”; Examiner’s note: wherein the alternative AI method is a large language model (i.e., the large language model used as a teacher model) or a small language model (i.e., the small language model used as a student model) is taught.) It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to combine the method and system for generating a semantic point cloud map in Golestan Irani, the 3D dynamic scene graphs in Rosinol, logic-based relationship graph expansion and extraction in Lassoued, the method and system for behavior control of AI-based systems in Steiner and the language material expanding method and device under multiple tasks based on small language model as taught in Wang. Golestan Irani teaches acquiring multimodal sensory data, creating data structures that associate one or more essential characteristics with each of the plurality of salient features, building an ontology that encompasses the relationships among the classes of objects in a plurality of datasets, outputting identifying information and relationship information and the multimodal sensory data including one or more of visible light images, video, sound recordings, radar signals, and physical measurements. Rosinol teaches extracting a plurality of salient features of the one or more objects at one or more levels of spatial resolution or structural hierarchy, mapping the spatial relationships among the one or more objects in visual scenes at multiple levels of spatial resolution or structural hierarchy and searching graph-based structures of the ontology. Lassoued teaches establishing one or more axioms that encode the relationships of individual datasets/classes of objects from multiple data examples and applying a deductive reasoning tool to the axioms of the ontology. Steiner teaches applying the method to an alternative Al method and constraining the alternative Al method to reduce or prevent data hallucination, bias, irrelevant conclusions, and/or Al misbehavior. Wang teaches the alternative Al method is a large language model or a small language model. One of ordinary skill would have motivation to combine Golestan Irani, Rosinol, Lassoued, Steiner and Wang to “improve[] the quality of the data generated by data enhancement and provide[] solutions for a variety of natural language tasks” (Wang, Page 2, Lines 15-16). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to YONG D RHO whose telephone number is (571)270-0194. The examiner can normally be reached 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at 5712705871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /YONG DOO RHO/Examiner, Art Unit 2147 /VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147
Read full office action

Prosecution Timeline

Jan 24, 2024
Application Filed
Sep 10, 2026
Non-Final Rejection mailed — §101, §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month