Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
No amendments or cancellations.
Claims 1-20 are pending.
Response to Arguments/Remarks
The Amendments to the Specification/abstract (04/13/2026) corrects the objection. The Objection is withdrawn.
Examiner thanks the applicant for the Interview summary. Note that in the interview the Examiner suggested some “recommendations to move prosecution forward.” No amendments were made.
Applicant argues:
Hu discloses using natural language processing to identify a navigational goal and the constraints of the goal. Hu relies on a predefined semantic map and grounds a goal phrase to a location in the predefined semantic map. The device plans a collision-free path using the semantic map and the goal phrase for the robot to follow (See Hu at Abstract). As the Office Action concedes, Hu does not disclose poses (See Office Action at page 11-12). Therefore, Hu fails to disclose "determining a target pose of the mobile device in the mapped environment corresponding to the embedding, wherein determining the target pose of the mobile device in the mapped environment corresponding to the embedding comprises: identifying a threshold similarity of the embedding to one of a plurality of embeddings of a multi-modal node graph; wherein the multi-modal node graph includes representations of a plurality of poses of the mobile device in the mapped environment and the plurality of embeddings correspond to the plurality of poses of the mobile device in the mapped environment; and moving the mobile device to a location and an orientation of the target pose in the mapped environment," as required by claim 1.
Samarasekera fails to cure the deficiencies of Hu. Samarasekera discloses a multi-modal and multi-sensor platform to uses poses to generate corrected navigation data. "The navigation and mapping subsystem 126 can utilize a combination of data obtained from an IMU and video data to estimate an initial navigation path for the platform, and then use the relative 6 degrees of freedom (6DOF) poses that are estimated by the 6DOF pose estimation module to generate a 3D map based on data obtained from, e.g., a scanning lidar sensor. The 3D LIDAR features can then10
be further exploited to improve the navigation path and the map previously generated based on, e.g., the IMU and video data." (See Samarasekera at Paragraph [0037]). While Samarasekera disclose poses, Samarasekera fails to disclose target poses, and moving the mobile device to a location and an orientation of the target pose in the mapped environment, as required by claim 1. Samarasekera uses poses during navigation to improve the navigation path. However, in the Specification as filed, the mobile device determines a target pose and causes navigation of the mobile device to a location and an orientation of the target pose in the mapped environment (see Specification at paragraph [0032]). Therefore, Samarasekera fails to disclose "determining a target pose of the mobile device in the mapped environment corresponding to the embedding, wherein determining the target pose of the mobile device in the mapped environment corresponding to the embedding comprises: identifying a threshold similarity of the embedding to one of a plurality of embeddings of a multi-modal node graph; wherein the multi-modal node graph includes representations of a plurality of poses of the mobile device in the mapped environment and the plurality of embeddings correspond to the plurality of poses of the mobile device in the mapped environment; and moving the mobile device to a location and an orientation of the target pose in the mapped environment", as required by claim 1. Accordingly, Applicant respectfully requests withdrawal of the rejections of claim 1. Applicant also respectfully requests withdrawal of claims 15 and 16 (and their respective dependent claims) for similar reasons.
Examiner respectfully disagrees. The art of record does teach the limitations as shown below. Examiner mentioned in the interview summary a list of more ART that also cover these limitations. Examiner has also found some additional art. (see 103 below and additional ART in the conclusion).
Note that under a broadest reasonable interpretation (BRI), words of the claim must be given their plain meaning, unless such meaning is inconsistent with the specification. The plain meaning of a term means the ordinary and customary meaning given to the term by those of ordinary skill in the art at the relevant time. The ordinary and customary meaning of a term may be evidenced by a variety of sources, including the words of the claims themselves, the specification, drawings, and prior art. However, the best source for determining the meaning of a claim term is the specification - the greatest clarity is obtained when the specification serves as a glossary for the claim terms. The words of the claim must be given their plain meaning unless the plain meaning is inconsistent with the specification. 2111.01 (I). See also In re Marosi, 710 F.2d 799, 802, 218 USPQ 289, 292 (Fed. Cir. 1983) ("'[C]laims are not to be read in a vacuum, and limitations therein are to be interpreted in light of the specification in giving them their ‘broadest reasonable interpretation.'"2111.01 (II)
With respect to the interpretation of claim terms, MPEP 2111 states:
The Patent and Trademark Office ("PTO") determines the scope of claims in patent applications not solely on the basis of the claim language, but upon giving claims their broadest reasonable construction "in light of the specification as it would be interpreted by one of ordinary skill in the art." In re Am. Acad. of Sci. Tech. Ctr., 367 F.3d 1359, 1364[, 70 USPQ2d 1827, 1830] (Fed. Cir. 2004). Indeed, the rules of the PTO require that application claims must "conform to the invention as set forth in the remainder of the specification and the terms and phrases used in the claims must find clear support or antecedent basis in the description so that the meaning of the terms in the claims may be ascertainable by reference to the description." 37 CFR 1.75(d)(1).
The words of the claim must be given their plain meaning unless the plain meaning is inconsistent with the specification In re Zletz, 893 F.2d 319, 13 USPQ2d 1320 (Fed. Cir. 1989).
"Though understanding the claim language may be aided by explanations contained in the written description, it is important not to import into a claim limitations that are not part of the claim. For example, a particular embodiment appearing in the written description may not be read into a claim when the claim language is broader than the embodiment." Superguide Corp. v. DirecTV Enterprises, Inc., 358 F.3d 870, 875, 69 USPQ2d 1865, 1868 (Fed. Cir. 2004).(see MPEP 2111.01).
During patent examination, the pending claims must be "given their broadest reasonable interpretation consistent with the specification." The broadest reasonable interpretation does not mean the broadest possible interpretation. Rather, the meaning given to a claim term must be consistent with the ordinary and customary meaning of the term (unless the term has been given a special definition in the specification), and must be consistent with the use of the claim term in the specification and drawings. Further, the broadest reasonable interpretation of the claims must be consistent with the interpretation that those skilled in the art would reach. In re Cortright, 165 F.3d 1353, 1359, 49 USPQ2d 1464, 1468 (Fed. Cir. 1999) (see PMEP 2111).
Accordingly, the claims herein will be interpreted in accordance with the MPEP 2111.
Claim Rejections - 35 USC§ 103
In the event the determination of the status of the application as subject to AIA 35
U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
Determining the scope and contents of the prior art.
Ascertaining the differences between the prior art and the claims at issue.
Resolving the level of ordinary skill in the pertinent art.
Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 5, 6, 12, 15, 16, and 20 are rejected under 35 U.S.C. 103 as being
unpatentable over HU (Z. Hu, J. Pan, T. Fan, R. Yang and D. Manocha, "Safe Navigation With Human Instructions in Complex Scenes," in IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 753-760, April 2019) in view of SAMARASEKERA (US 20150269438 A1).
Regarding claim 1:
HU discloses:
at a mobile device in communication with one or more input devices: [See at least HU, Pg 754 Section I ("Our approach contains four steps: pre-processing, phrase classification, goal and constraint grounding, and motion planning. First, the pre-processing step takes the command sentence as the input and divides it into phrases according to conjunctions and commas. Next, we assume that each phrase has one of three labels: goal, constraint, or uninformative phrase. For example, the phrase "go to the restaurant" is a goal phrase, "keep away from people" is a constraint phrase, and the phrase "you know" is uninformative. We train a Long Short-Term Memory (LSTM) network [19] to classify phrases into those three types. The LSTM is suitable for our task because the classification output does not depend only on individual words but also on the meaning of the entire phrase. After recognizing the constraint and goal phrases, we need to ground them with the physical world. In particular, we ground the goal phrase by computing the similarity between the noun extracted from the goal phrase and the location name in the predefined semantic map, and use the most similar location as the goal configuration for the robot navigation.")];
receiving, via the one or more input devices, an input including a
description associated with an object in a mapped environment; [See at least HU, Pg 754-755, Section Ill, "Our approach contains four steps: pre-processing, phrase classification, goal and constraint grounding, and motion planning. First, the pre-processing step takes the command sentence as the input and divides it into phrases according to conjunctions and commas. Next, we assume that each phrase has one of three labels: goal, constraint, or uninformative phrase. For example, the phrase "go to the restaurant" is a goal phrase, "keep away from people" is a constraint phrase, and the phrase "you know" is uninformative. We train a Long Short-Term Memory (LSTM) network [19] to classify phrases into those three types. The LSTM is suitable for our task because the classification output does not depend only on individual words but also on the meaning of the entire phrase. After recognizing the constraint and goal phrases, we need to ground them with the physical world. In particular, we ground the goal phrase by computing the similarity between the noun extracted from the goal phrase and the location name in the predefined semantic map, and use the most similar location as the goal configuration for the robot navigation."; Pg 4, Section V, "For goal grounding, we compute the goal location by extracting the noun from the goal phrase and then computing the similarity between this noun and the location names in the predefined semantic map. The similarity is computed as the cosine similarity between the embedded vectors of two given word items. The embedding is accomplished using the Word2Vec embedding network [22], [23] that can convert a word into a vector based on the Wiki corpus. We use the 2D coordinate of the location that has the highest similarity with the noun in the goal phrase as the goal location.")].
generating, using a text encoder of a multi-modal model, an embedding corresponding to the input; [See at least HU, Pg 755, Section Ill, "We train a Long Short-Term Memory (LSTM) network [19] to classify phrases into those three types. The LSTM is suitable for our task because the classification output does not depend only on individual words but also on the meaning of the entire phrase."; Pg 755, Section IV ( "More specifically, as shown in Fig. 3, our classification subsystem contains four layers: the embedding layer, the bi-directional LSTM layer, the attention layer, and the output layer. The embedding layer transforms a phrase S with length T : S = {x1 , x2 ,..., xT } into a list of real-valued vectors E = {e1 , e2 ,..., eT }. We transform a word xi into the embedding vector through an embedding matrix We : ei = We vi, where vi is a vector of value 1 at index ei and 0 otherwise. The dimension of this embedding layer is 25. Next, we extract the representation for the sequential data E using LSTM, which provides an elegant way of modeling sequential data. However, in standard LSTM, the information encoded in the inputs can only flow in one direction and the future hidden unit cannot affect the current unit. To overcome this drawback, we employ a bi-directional LSTM layer [18], which can be trained using all the available inputs from two directions to improve the prediction performance.”)];
of the mobile device in the mapped environment corresponding to the embedding, wherein determining the target pose of the mobile device in the mapped environment corresponding to the embedding comprises: [See at least HU, Pgs. 754, Section I ("we present an algorithm to enable a mobile robot to understand the navigation goal and trajectory constraints from natural language human instructions and sensor measurements, in order to generate a high quality navigation trajectory following user's requirements. We first use a Long Short-Term Memory (LSTM) recursive arXiv:1809.04280v1 [cs.RO] 12 Sep 2018 1 neural network to parse and interpret the commands and generate the navigation goal and a set of constraint phrases. The navigation goal is in form of a location name in the semantic map. The constraint phrase describes the spatial relationship between the robot and a target object, e.g., in "keep away from the desk", "desk" is the target object and "keep away from" is the spatial relation."; Pg 756, Section V, "For goal grounding, we compute the goal location by extracting the noun from the goal phrase and then computing the similarity between this noun and the location names in the predefined semantic map. The similarity is computed as the cosine similarity between the embedded vectors of two given word items. The embedding is accomplished using the Word2Vec embedding network [22], [23] that can convert a word into a vector based on the Wiki corpus. We use the 2D coordinate of the location that has the highest similarity with the noun in the goal phrase as the goal location.")];
identifying a threshold similarity of the embedding to one of a plurality of embeddings of a multi-modal node graph; [See at least HU, Pg 756, Section V ("For goal grounding, we compute the goal location by extracting the noun from the goal phrase and then computing the similarity between this noun and the location names in the predefined semantic map. The similarity is computed as the cosine similarity between the embedded vectors of two given word items. The embedding is accomplished using the Word2Vec embedding network [22], [23] that can convert a word into a vector based on the Wiki corpus. We use the 2D coordinate of the location that has the highest similarity with the noun in the goal phrase as the goal location."); Pg 757, Section V ("Once we label the objects in the current scene, we compute the similarity between the nouns that occur in the constraint phrase and those object labels discovered by the scene understanding. Again, we use the Word2Vec embedding network [22], [23] to embed all words in the vector space and compute the cosine similarity between word vectors of the constraint object and objects detected in the instance segmentation. If the similarity is larger than a predefined threshold, we add this constraint to the motion planner.")];
wherein the multi-modal node graph includes the plurality of embeddings correspond to the plurality of poses of the mobile device in the mapped environment; and [See at least HU, Pg 756-757, Section V ("The robot then starts to follow the globally planned trajectory, but it needs to per-form local planning to adapt to the dynamically added constraints. In particular, the local planning algorithm will maintain a costmap in the robot’s local coordinate. The costmap contains both static obstacles and constraints. Static obstacles such as walls, rooms, and buildings can be directly added into the costmap by transforming a subset of the global map. The constraints can also be conveniently modeled in the costmap. Given the location (x0 , y0) of a constraint object computed via constraint grounding, we can mark the cells inside the region {(x, y) : (x − x0 )2 + (y − y0 )2 ≤ a2 } as the obstacle cells, where a is a parameter indicating the influencing radius of the constraint. Finally, we perform a smooth inflation operation over all the obstacle cells in the grid map to enable the robot to keep a safe distance from the obstacle. Samples of the computed costmap are shown in Fig. 6. ")];
moving the mobile device to a location and an orientation of the target pose in the mapped environment. [See at least HU, Pg 753, Abstract, "In this letter, we present a robotic navigation algorithm with natural language interfaces that enables a robot to safely walk through a changing environment with moving persons by following human instructions such as “go to the restaurant and keep away from people.” We first classify human instructions into three types: goal, constraints, and uninformative phrases. Next, we provide grounding in a dynamic manner for the extracted goal and constraint items along with the navigation process to deal with target objects that are too far away for sensor observation and the appearance of moving obstacles such as humans. In particular, for a goal phrase (e.g., “go to the restaurant”), we ground it to a location in a predefined semantic map and treat it as a goal for a global motion planner, which plans a collision-free path in the workspace for the robot to follow. For a constraint phrase (e.g., “keep away from people”), we dynamically add the corresponding constraint into a local planner by adjusting the values of a local costmap according to the results returned by the object detection module. The updated costmap is then used to compute a local collision avoidance control for the safe navigation of the robot. By combining natural language processing, motion planning, and computer vision, our developed system can successfully follow natural language navigation instructions to achieve navigation tasks in both simulated and real-world scenarios.“); Pg 754-755, Section Ill ("To add the target object in a constraint into the planner’s costmap dynamically, we use an object detection module to look for the object mentioned in the constraint phrase and then output the object’s 3D location in the robot’s local coordinate system. To deal with moving objects (like “people”) or occluded objects, we perform the grounding and navigation in a dynamic manner. That is, whenever a constraint object occurs in the camera frame, we add it to the local planner’s costmap. Finally, our costmap motion planner will compute a suitable trajectory based on the costmap as it is updated in real-time according to the grounded constraint result. To improve the navigation performance in crowd scenarios, we further use the planning result to guide a reinforcement learning based local collision avoidance approach developed in our previous work [17]. An overview of our proposed navigation system is shown in Fig. 2.")].
HU does not disclose, but SAMARASEKERA teaches:
determining a target pose [See at least SAMARASEKERA, ¶ 0037 (discusses “pose estimation module” and how the embodiment operates ")];
representations of a plurality of poses of the mobile device in the mapped environment and poses [See at least SAMARASEKERA, ¶ 0037; 0042 (expands on "FIGS. 13A and 13B show examples of visualization output resulting from the operations of an embodiment of the live analytics subsystem”)];
an orientation of the target pose [see at least SAMARASEKERA, ¶ 0037;¶
0042)].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance within HU to include the tracking of poses (3d space) for generating an accurate map within SAMARASEKERA to yield a more effective robot navigation system that can account for the orientation of the robot rather than merely the position.
Regarding claim 5:
HU in view of SAMARASEKERA discloses the limitations within claim 1 and HU
does not disclose, but SAMARASEKERA further teaches:
a pose of the plurality of poses is represented using at least three values representing a location and an orientation of the mobile device in the mapped environment. [see at least SAMARASEKERA, ¶ 0003 (discusses “robot navigation technology”); 0037].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance within HU to include the tracking of poses (3d space) for generating an accurate map within SAMARASEKERA to yield a more effective robot navigation system that can account for the orientation of the robot rather than merely the position.
Regarding claim 6:
HU in view of SAMARASEKERA discloses the limitations within claim 5 and HU
further discloses:
the location is represented using at least two coordinates of a two dimensional coordinate system of the mobile device in the mapped environment and [see at least HU, Pg 757, Section VI, "For both global and local planning, we assume that the robot can accurately localize itself in a predefined global semantic map. We use a state-of-the-art 2D localization technique [21] for the robot localization. The global semantic map is constructed using SLAM (simultaneous localization and mapping) algorithm. The resulting map consists of grid cells which may be one of three types: free space, obstacle, and no information.")];
HU does not disclose, but SAMARASEKERA teaches:
the orientation is represented by yaw of the mobile device in the mapped environment. [see at least SAMARASEKERA, ¶ 0039 (discusses the “multi-modal geo-spatial data integration module” ); Note SAMARASEKERA to yield a more effective robot navigation system that can account for the orientation of the robot rather than merely the position].
Regarding claim 12:
HU in view of SAMARASEKERA discloses the limitations within claim 1 and HU
further discloses:
in accordance with identifying less than the threshold similarity of the embedding to the plurality of embeddings of the multi-modal node graph, forgoing moving the mobile device to the location and the orientation of the target pose in the environment. [see at least HU, Pg 756, Section V ("Once we label the objects in the current scene, we compute the similarity between the nouns that occur in the constraint phrase and those object labels…”)].
Regarding claim 15:
With regards to claim 15, this claim is the non-transitory computer readable storage medium claim to method claim 1 and is substantially similar to claim 1 and is therefore rejected using the same references and rationale.
Regarding claim 16:
With regards to claim 16, this claim is the mobile device claim to method claim 1 and is substantially similar to claim 1 and is therefore rejected using the same references and rationale.
Regarding claim 20:
With regards to claim 20, this claim is substantially similar to claim 5 and is therefore rejected using the same references and rationale.
Claims 2, 3, 7-9, 11, 14, 17, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over HU (Z. Hu, J. Pan, T. Fan, R. Yang and D. Manocha, "Safe Navigation With Human Instructions in Complex Scenes," in IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 753-760, April 2019) in view of SAMARASEKERA (US 20150269438 A1) in further view of ZHANG C. (US 20230118864 A1).
Regarding claim 2:
HU in view of SAMARASEKERA discloses the limitations within claim 1 and HU further discloses:
the multi-modal model generates embeddings in a shared dimensionality space for one [see at least HU, Pg 755, Section IV (discusses phase classification and "To understand the roles that different instruction phrases play in the motion planning, we propose an LSTM-based method to ground phrases into different types.”); Pg 756, Section V ( discusses "goal grounding” and how to use to generate embeddings)].
HU does not disclose, but ZHANG C. teaches:
or more language inputs or one or more image inputs. [see at least ZHANG C., ¶ 0190 (discusses detailed image inputs); 0207 (more on image inputs); 0244 (more on images)].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance with the tracking of 3d poses within HU in view of SAMARASEKERA to include using images to create vector graph embeddings of vector nodes for mapping within ZHANG C. to effectively yield a robot capable of mapping the given environment utilizing onboard images.
Regarding claim 3:
HU in view of SAMARASEKERA in further view of ZHANG C. discloses the limitations within claim 2 and HU further discloses:
a second respective embedding of the embeddings output by the multi-modal model for a second respective input including a text description of the respective object [see at least HU, Pg 754-755, Section Ill and IV (discusses input and embedding layer)];
are generated in the shared dimensionality space and have a similarity above the threshold similarity. [see at least HU, Pg 755, Section IV; Pg 756, Section V)].
HU does not disclose, but ZHANG C. teaches:
a first respective embedding of the embeddings output by the multi-modal model for a first respective input including an image of a respective object and [see at least ZHANG C., ¶ 0190 (discusses embedding); 0207 ("In step 207 a graph embedding is generated by translating the generated scene graph into a vector that can be used for subsequent comparisons (e.g. for image recognition). Optionally, the graph embedding is generated by a graph embedding network."); 0244 (more on embedding)].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance with the tracking of 3d poses within HU in view of SAMARASEKERA to include using images to create vector graph embeddings of vector nodes for mapping within ZHANG C. to effectively yield a robot capable of mapping the given environment utilizing onboard images.
Regarding claim 7:
HU in view of SAMARASEKERA discloses the limitations within claim 1 and HU
does not disclose, but ZHANG C. teaches:
the one or more input devices include one or more image sensors, the method further comprising: receiving, via the one or more image sensors, an image input; and [see at least ZHANG C., ¶ 0035 (discusses image and image identifying)];
generating, using an image encoder of the multi-modal model, an embedding corresponding to the image input. [see at least ZHANG C., ¶ 0190; 0207; 0244].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance with the tracking of 3d poses within HU in view of SAMARASEKERA to include using images to create vector graph embeddings of vector nodes for mapping within ZHANG C. to effectively yield a robot capable of mapping the given environment utilizing onboard images.
Regarding claim 8:
HU in view of SAMARASEKERA discloses the limitations within claim 1 and HU
further discloses:
the mapped environment is mapped using the multi-modal model; [see at least HU, Pg 755, Section VI; Pg 758-759, Section VII (discusses mapping)];
HU does not disclose, but ZHANG C. teaches:
mapping the environment includes generating a plurality of nodes of the multi-modal node graph; and [see at least ZANG C., ¶ 0022 (discussing images, mapping and how to use them in grafting); 0209 ( "The encoder 702 is configured to map node and edge features”)];
generating the plurality of nodes of the multi-modal node graph includes generating, using an image encoder of the multi-modal model, a respective embedding corresponding to a respective pose corresponding to a respective image. [see at least ZHANG C., ¶ 0002; 0022; 0173 (discusses encoders); 0208 (more embedding and encoders): ¶ 0209)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance with the tracking of 3d poses within HU in view of SAMARASEKERA to include using images to create vector graph embeddings of vector nodes for mapping within ZHANG C. to effectively yield a robot capable of mapping the given environment utilizing onboard images.
Regarding claim 9:
HU in view of SAMARASEKERA in further view of ZANG C. discloses the limitations within claim 8 and HU does not disclose, but ZHANG C. teaches:
generating the plurality of nodes of the multi-modal node graph includes: [see at least ZHANG C., ¶ 0022; 0209];
receiving a first embedding output by the image encoder of the multi-modal model at a first pose; [see at least ZHANG C., ¶ 0002; ¶ 0209)];
in accordance with a determination that the first pose has less than a threshold similarity with the plurality of poses at the plurality of nodes of the multi-modal node graph, adding a new node to the multi-modal node graph with the first embedding and the first pose; and [see at least ZHANG C., ¶ 0097; 0290 ("In an example there is a method comprising modifying a scene graph (e.g. generated by step 206) to add additional nodes (associated with an additional object) and edges connecting the node to other nodes of the graph. In this way, the scene graph can be modified/augmented to include an object that was not present in the image obtained in step 201. Adding edges between nodes controls the position of the additional object within the scene. Additionally or alternatively the scene graph generated by step 206 is modified before generating a graph embedding (e.g. in step 207) to remove nodes from the graph representation. Removing nodes may be advantageous where, for example, the image data obtained in step 201 is known to include a temporary object that would likely not be present in the reference set. Consequently, by removing this node the associated object will not be taken into account when performing similarity comparisons with graph embeddings in the reference set. Optionally the modified/augmented is generated based on the output of step 206 (generating a scene graph) and provided to the input of step 207 (generating a graph embedding).")];
in accordance with a determination that the first embedding has a threshold similarity with a second embedding at a node of the multi-modal node graph corresponding to the first pose, forgoing adding the new node to the multi-modal node graph with the first embedding and the first pose. [see at least ZHANG C., ¶ 0097; ¶ 0289; ¶ 0290)].
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance with the tracking of 3d poses within HU in view of SAMARASEKERA to include using images to create vector graph embeddings of vector nodes with threshold comparisons for determining meaningful changes for mapping within ZHANG C. to effectively yield a robot capable of mapping the given environment utilizing collected images.
Regarding claim 11:
HU in view of SAMARASEKERA in further view of ZANG C. discloses the limitations within claim 8 and HU further discloses:
while moving the mobile device to the location and the orientation of the target pose in the mapped environment: [see at least HU, Abstract; Pg 754-755, Section Ill].
HU does not disclose, but SAMARASEKERA teaches:
orientation of the target pose [see at least SAMARASEKERA, ¶ 0037]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance within HU to include the tracking of poses (3d space) for generating an accurate map within SAMARASEKERA to yield a more effective robot navigation system that can account for the orientation of the robot rather than merely the position.
HU in view of SAMARASEKERA does not disclose, but ZHANG C. teaches:
receiving a first embedding output by the image encoder of the multi-modal model at a first pose; [see at least ZHANG C., ¶ 0190, "Consequently, given N D-dimension local pixel descriptors as input, and K automatically found cluster centres, the output of 'NetVLAD' at each pixel is a KxD matrix based on the descriptor distance to each cluster center in the feature space. In 'NetVLAD' all pixel matrices are simply aggregated into a single matrix. This matrix is then flattened and used as the image embedding for image retrieval."; ¶ 0207; 0244]);
in accordance with a determination that the first pose has a threshold similarity with a second pose at a respective node of the plurality of nodes of the multi-modal node graph and the first embedding has less than a threshold similarity with a second embedding at the respective node, updating the respective node to include the first embedding at the first pose; and [see at least ZHANG C., ¶ 0097; 0289 ("Representing the scene as graph has further advantages. For example, representing the scene as a graph allows easier manipulation such that nodes (i.e. objects) can be added or removed from the scene graph representation of the input image."; 0290 ]
in accordance with a determination that the first embedding has a threshold similarity with a second embedding at the respective node of the plurality of nodes of the multi-modal node graph corresponding to the first pose, forgoing updating the respective node. [see at least ZHANG C., ¶ 0097, ""; ¶ 0289, ""; ¶ 0290, "")
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance with the tracking of 3d poses within HU in view of SAMARASEKERA to include using images to create vector graph embeddings of vector nodes with threshold comparisons for determining meaningful changes for mapping within ZHANG C. to effectively yield a robot capable of mapping the given environment utilizing onboard images.
Regarding claim 14:
HU in view of SAMARASEKERA in further view of ZANG C. discloses the limitations within claim 8 and HU further discloses:
the one or more input devices [see at least HU, Pg 754-755, Section Ill, ( discusses/suggests that there are more than one input devices in order to fulfil the goals stated.)]
HU does not teach, but SAMARASEKERA discloses:
includes an image capture device, [see at least SAMARASEKERA, abstract (indicated the use of image data): ¶ 0030 (discusses use of image sensors to gather the image data)];
a motion sensor, or an odometry sensor. [see at least SAMARASEKERA, ¶
0030)].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance within HU to include the tracking of poses (3d space) using an image sensor and IMU for generating an accurate map within SAMARASEKERA to yield a more effective robot navigation system that can account for the orientation of the robot rather than merely the position.
Regarding claim 17:
With regards to claim 17, this claim is substantially similar to claim 2 and is therefore rejected using the same references and rationale.
Regarding claim 18:
With regards to claim 18, this claim is substantially similar to claim 3 and is therefore rejected using the same references and rationale.
Claims 4 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over HU (Z. Hu, J. Pan, T. Fan, R. Yang and D. Manocha, "Safe Navigation With Human Instructions in Complex Scenes," in IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 753-760, April 2019) in view of SAMARASEKERA (US 20150269438 A1) in
further view of RADFORD (A Radford, J. Kim, C. Hallacy, A Ramesh, et al., "Learning Transferable Visual Models From Natural Language Supervision.", International Conference on Machine Learning, 2021).
Regarding claim 4:
HU in view of SAMARASEKERA discloses the limitations within claim 1 and HU does not disclose, but RADFORD teaches:
the multi-modal model includes a contrastive language image pre-training model. [see at least RADFORD, Pg 3, Figure 2; Pg 2 Col 2 - Pg 3 Col 1 (discusses models and learning images using " natural language supervision at large scale” and crating new datasets from image and text using a pairing function (thus contrastive; Showing another article (Hestness et al., 2017; Kaplan et al., 2020) to support the model formation function)].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify, with a reasonable expectation of success, the model for motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance with the tracking of 3d poses within HU in view of SAMARASEKERA to utilize a contrastive language image pre-training (CLIP) model within RADFORD to yield a more efficient planning model that scales better when working on a training set of millions of images as anticipated by RADFORD.
Regarding claim 19:
With regards to claim 19, this claim is substantially similar to claim 4 and is therefore rejected using the same references and rationale.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over HU (Z. Hu,
J. Pan, T. Fan, R. Yang and D. Manocha, "Safe Navigation With Human Instructions in Complex Scenes," in IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 753-760, April 2019) in view of SAMARASEKERA (US 20150269438 A1) in further view of ZHANG C. (US 20230118864 A1) in further view of XIE (US 20210097739 A1).
Regarding claim 10:
HU in view of SAMARASEKERA in further view of ZHANG C. (US 20230118864 A1) discloses the limitations within claim 8 and HU does not disclose, but ZHANG C. teaches:
the generating the plurality of nodes of the multi-modal node graph includes: [see at least ZHANG C., ¶ 0022 (discusses place recognition and the steps needed using multiple nodes); 0209 (discusses map node and grafing using “initial node and edge vectors”)]
updating the multi-modal node graph [see at least ZHANG C., ¶ 0097 (further discusses using nodes and other specific data to “generate a graph of the scene”)}; 0290 ("a method comprising modifying a scene graph”)].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with natural language recognition for determining the goal and local planning to examine current positions for collision avoidance with the tracking of 3d poses within HU in view of SAMARASEKERA to include using images to create vector graph embeddings of vector nodes for mapping within ZHANG C. to effectively yield a robot capable of mapping the given environment utilizing onboard cameras.
HU in view of SAMARASEKERA in further view of ZHANG C. does not disclose, but XIE teaches:
identifying one or more loop closures for the multi-modal node graph based on embeddings output by the image encoder of the multi-modal node graph, and [see at least XIE, ¶ 0003 (discusses the use of "Loop closures” within more than one node) ; 0102 (further describes using loop closure between graph nodes); 0145 ( "The intuition is that the relative poses encapsulated in the loop closure pair should be consistent with the odometry information of the outbound and inbound trajectory segments (i.e. the portions of the trajectory between pose nodes of the loop closures.”)];
based on the one or more loop closures. [see at least XIE, ¶ 0034 (discusses "The method may then process the graph to estimate confidence scores for each loop closure. Typically, the confidence score is generated by performing pairwise consistency tests between each loop closure and a set of other loop closures. Conveniently, the method then generates an augmented graph from the initial graph."); 0035 (indicates "Generating the augmented graph may comprise retaining and/or deleting each loop closure based upon the confidence scores.")].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine, with a reasonable expectation of success, the motion planning with images to create vector nodes within HU in view of SAMARASEKERA in further view of ZHANG to include loop closure detection and modification of the pose graph based on loop closures as within XIE to effectively yield an improved pose estimation system that corrects for cumulative odometry error as disclosed in XIE ¶ 0003-0004.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over HU (Z. Hu,
J. Pan, T. Fan, R. Yang and D. Manocha, "Safe Navigation With Human Instructions in Complex Scenes," in IEEE Robotics and Automation Letters, vol. 4, no. 2, pp. 753-760, April 2019) in view of SAMARASEKERA (US 20150269438 A1) in further view of GRZESIAK (KR20210109722A).
Regarding claim 13:
HU in view of SAMARASEKERA discloses the limitations within claim 1 and HU further discloses:
text representation of the description associated with the object in the mapped environment. [see at least HU, Pg 754-755, Section Ill].
HU does not disclose, but GRZESIAK
wherein the one or more input devices include one or more audio sensors, the input includes a voice command, and receiving the input includes: capturing, via the one or more audio sensors, the voice command; and converting the voice command into a text representation [see at least GRZESIAK, ¶ 0001 (indicates use of " a method for controlling the device that performs voice recognition based on a user's utterance status and generates a control command based on the voice recognition”); 0006 (further discussion on use of voice commands)].
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify, with a reasonable expectation of success, the phrase collection and processing within HU to implement voice commands within GRZESIAK to effectively yield a more versatile command system that allows for user voice to be inputted for goal setting.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
R. Giubilato, M. Vayugundla, W. Stürzl, M. J. Schuster, A. Wedler and R. Triebel, "Multi-Modal Loop Closing in Unstructured Planetary Environments with Visually Enriched Submaps," 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 2021, pp. 8758-8765.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOAN T GOODBODY whose telephone number is (571) 270-7952. The examiner can normally be reached on M-TH 7-3 (US Eastern time).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at https://www.uspto.gov/patents/uspto-automated-interview-request-air-form.html.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, RACHID BENDIDI can be reached at (571) 272-4896. The Fax phone number for the organization where this application or proceeding is assigned is (571) 273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspot.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at (866) 217-9197 (toll-free). If you would like assistance from the USPTO Customer Serie Representative or access to the automated information system, call (800) 786-9199 (IN USA OR CANADA) or (571) 272-1000.
/JOAN T GOODBODY/
Primary Examiner, Art Unit 3664
(571) 270-7952