DETAILED ACTION
This communication is in response to the Amendments and Arguments filed on 04/08/2026.
Claims 1-7, 10-15, 17-20, and 22-24 are pending and have been examined.
All previous objections/rejections not mentioned in this Office Action have been withdrawn by the examiner.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed 04/08/2026 have been fully considered.
Applicant asserts on pgs 17-18 that El Dokor does not teach a projection of a direction onto a micomap to determine a POI. The Examiner respectfully disagrees with this assertion. El Dokor teaches identifying a target region from the input data and gesture recognition module, such as a 3D cone, that may be overlapped on the micromap to identify any POIs located within the overlapped area (see (5:64-6:11),(6:20-67),(7:1-17,28-43),(8:20-62),(9:3-7,28-33),(10:11-51)). This reads on the BRI of mapping a projection of the direction onto a map of the environment. Regarding gaze direction specifically, Powderly is now cited to teach the updated claim language.
Hence, Applicant’s arguments are not persuasive and/or are moot.
Applicant’s arguments with respect to claim(s) 18 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Please see the updated mappings citing Panchaksharaiah for further detail.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 11, 13, 14, 17, and 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over El Dokor et al. (U.S. Patent No. 8,818,716), hereinafter El Dokor, in view of Kale et al. (U.S. Patent No. 11804035), hereinafter Kale.
Regarding claim 11, El Dokor teaches
A system comprising (a system (2:54-3:7)):
one or more processors to (multiple processors may be included in each device (4:48-59)):
receive audio data obtained using one or more microphones of a vehicle, the audio data representing speech (the voice recognition module receives an output signal from the microphone in the vehicle that captures voice signals, i.e. receive audio data obtained using one or more microphones of the vehicle…representing speech (3:66-4:10));
receive image data representing a user located within an interior of the vehicle (a camera system captures images of a gesture performed by a user in the vehicle, i.e. receive image data representing a user located within an interior of the vehicle (5:64-6:11));
determine, based at least on the image data, a direction indicated by the user located within the interior of the vehicle (a camera system captures images of a gesture performed by a user in the vehicle, i.e. determining based at least on the image data, and the gesture recognition module analyzes the gesture to determine a direction vector representing the direction of the gesture, i.e. determining…a direction indicated by the user located within the interior of the vehicle (5:64-6:11));
determine, based at least on localizing the vehicle with respect to a map of an environment, a projection of the direction overlayed on the map (the input analysis module receives input data from the gesture recognition module and the location module to identify a target region and micromap with POIs, using the location, orientation, and speed of the vehicle to automatically identify micromaps, i.e. based at least on localizing the vehicle with respect to a map of an environment, where a three-dimensional target region like a cone may be used to retrieve information within the micromapped region by overlapping the target region with the micromapped region, i.e. determine…a projection of the direction overlayed on the map (5:64-6:11),(6:20-67),(7:1-17,28-43),(8:20-62),(9:3-7,28-33),(10:11-51));
determine, based at least on the projection of the direction on the map at least partially overlapping a point of interest (POI) represented by the map, that the user is indicating the POI within the environment (the input analysis module receives input data from the gesture recognition module and the location module to identify a target region to overlay on a micromap with POIs, i.e. based at least the projection of the direction on the map…overlapping…the map, where a POI within the target region is identified, i.e. at least partially overlapping with a POI, where the parameters of the target region may be adjusted until a single POI is returned, i.e. determine…that the user is indicating the POI within the environment (5:64-6:11),(6:20-67),(7:1-17,28-43),(8:20-62),(9:3-33),(10:11-51));
determine, using one or more …models and based at least on the audio data and contextual information associated with the POI, an output associated with the speech (the input analysis module receives input data, including from the voice recognition module, i.e. based at least on the audio data, and the POI search module queries a databases to retrieve information related to POI, i.e. based at least on…contextual information associated with the POI, which can be merged by an artificial intelligence unit, and which can be received by the data output module for output to the user, i.e. determine…an output associated with the speech (6:12-19),(6:31-51),(7:6-17),(8:47-9:27),(9:27-67)); and
cause the vehicle to provide the output associated with the speech (the data output module receives information related to the POI from the search modules, and sends the information to the in-vehicle communications system or display, i.e. causing the vehicle to provide the output, in response to the voice command and gesture, i.e. associated with the speech (5:64-6:19),(7:6-17),(9:51-60)).
While El Dokor provides recognizing voice commands using an algorithm and providing information to a user in response to the command, El Dokor does not specifically teach using a machine learning model to process two types of data to determine an output, and thus does not teach
determine, using one or more machine learning models and based at least on the audio data and contextual information associated with the POI, an output associated with the speech.
Kale, however, teaches determine, using one or more machine learning models and based at least on the audio data and contextual information associated with the POI, an output associated with the speech (the NLG component generates language for a textual or spoken reply, i.e. determining…an output, based on the decisions made by the artificial intelligence framework, i.e. using one or more machine learning models, where the NLU component may parse user inputs to determine the user intent and intent-related parameters, such as a dominant object and a variety of attributes related to that object, such as looking for a dress for a wedding in June in Italy, i.e. based at least on the audio data, where the result of the search may include an item list of candidate products, i.e. based at least on…contextual information associated with the POI, and the output can include the search results, i.e. determine…an output associated with the speech (1:60-2:7),(10:12-27,50-65),(17:23-18:28),(20:2-5,14-36).
El Dokor and Kale are analogous art because they are from a similar field of endeavor in providing responses to users based on multimodal input. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the recognizing voice commands using an algorithm and providing information to a user in response to the command teachings of El Dokor with the use of an artificial intelligence framework to determine a reply and generate a response using the intent and various parameters as taught by Kale. It would have been obvious to combine the references to learn from user intents to enhance user understanding and provide an improved user experience (Kale (1:60-2:7)).
Regarding claim 13, El Dokor in view of Kale teaches claim 11, and Kale further teaches
determine, using one or more second machine learning models and based at least on the audio data, an intent associated with the speech (the artificial intelligence framework includes an NLU component, i.e. one or more second machine learning models, that operates to extract user intent and various intent parameters, i.e. determine an intent associated with the speech (1:60-2:7),(7:63-67),(17:34-45));
append the contextual information to the intent (the extracted data includes the intent and intent-related parameters such as the attributes and attribute values related to a dominant object, i.e. contextual information, and the extracted data is provided to the dialog manager as a formal, machine-readable, structured representation of the query including the user query and further data, i.e. append the contextual information to the intent (17:23-59),(17:60-18:28)); and
apply, as an input to the one or more machine learning models, the contextual information appended to the intent (the NLG component generates language for a textual or spoken reply based on the decisions made by the artificial intelligence framework, i.e. the one or more machine learning models, where the NLU component may parse user inputs to determine the user intent and intent-related parameters, such as a dominant object and a variety of attributes related to that object, such as looking for a dress for a wedding in June in Italy, where the result of the search may include an item list of candidate products, and the output can include the search results and be based on the information from the context manager including the relevant intent and all parameters and related results, i.e. apply as an input…the contextual information appended to the intent (1:60-2:7),(10:12-27,50-65),(17:23-18:28),(20:2-5,14-36)).
Where the motivation to combine is the same as previously presented.
Regarding claim 14, El Dokor in view of Kale teaches claim 11, and El Dokor further teaches
determine, using one or more second …models and based at least on the audio data, at least one of an intent associated with the speech or one or more parameters for one or more slots associated with the intent (the voice recognition module receives an output signal from the microphone in the vehicle that captures voice signals, i.e. audio data, and performs a voice recognition algorithm on the received signal to identify voice commands that provide additional information about the object being identified, i.e. determine an intent associated with the speech, such as a skyscraper in the city the vehicle is driving through, and that the person spoke the word “building” to identify the particular object, i.e. one or more parameters for one or more slots corresponding to the intent (3:66-4:10),(6:12-19),(7:28-43),(8:1-19),(10:11-25));
wherein the determination of the output associated with the speech is based at least on the contextual information and the at least one of the intent or the one or more parameters (the input analysis module receives input data, including from the voice recognition module, i.e. the at least one of the intent or the one or more parameters, and the POI search module queries a databases to retrieve information related to POI, i.e. contextual information, which can be merged by an artificial intelligence unit, and which can be received by the data output module for output to the user, i.e. determination of the output associated with the speech is based at least on (6:12-19),(6:31-51),(7:6-17),(8:47-9:27),(9:27-67)).
Where Kale further teaches determine, using one or more second machine learning models and based at least on the audio data, at least one of an intent associated with the speech or one or more parameters for one or more slots associated with the intent (the NLG component generates language for a textual or spoken reply based on the decisions made by the artificial intelligence framework, i.e. using one or more second machine learning models, where the NLU component may parse user inputs to determine the user intent and intent-related parameters, such as a dominant object and a variety of attributes related to that object, such as looking for a dress for a wedding in June in Italy, i.e. determine…at least one of an intent associated with the speech or one or more parameters for one or more slots associated with the intent (1:60-2:7),(10:12-27,50-65),(17:23-18:28),(20:2-5,14-36));
And where the motivation to combine is the same as previously presented.
Regarding claim 17, El Dokor in view of Kale teaches claim 11, and El Dokor further teaches
the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for the autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for presenting virtual reality (VR) content;
a system for presenting augmented reality (AR) content;
a system for presenting mixed reality (MR) content;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing deep learning operations;
a system implemented using an edge device (the system may be performed by a device in the operating environment, such as an in-vehicle communications system, a wireless mobile communication device, and/or a remote server, i.e. edge devices (2:54-3:13));
a system implemented using a robot;
a system for performing conversational AI operations;
a system for generating synthetic data;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center (the system may be performed by a device in the operating environment, such as an in-vehicle communications system, a wireless mobile communication device, and/or a remote server, and may further access a database on the server or a service operating on a third-party server, such as Yelp or Google, i.e. at least partially in a data center (2:54-3:13),(6:58-51));
or a system implemented at least partially using cloud computing resources (the system may be performed by a device in the operating environment, such as an in-vehicle communications system, a wireless mobile communication device, and/or a remote server, i.e. at least partially using cloud computing resources (2:54-3:13).
Regarding claim 23, El Dokor in view of Kale teaches claim 11, and El Dokor further teaches
store, based at least on the user indicating the POI within the environment, an indication of the POI within a log (multiple POIs may be identified by the POI search module, i.e. store…an indication of the POI within a log based on the target region identified by the input analysis module, i.e. based at least on the user indicating the POI within the environment (8:63-9:27)); and
determine, based at least on the audio data, the contextual information using the indication of the POI from the log (the voice command in combination with the gesture, i.e. based at least on the audio data, may be used to narrow the results of the POI search for a specific term within the target region, where multiple POIs returned from a search may be iteratively narrowed to a single POI, and the single POI information is sent to the data output module, i.e. determine…the contextual information using the indication of the POI from the log (8:63-9:27)).
Where Kale further teaches that the interaction history with the user may be stored to be used during searches, i.e. store…within a log (8:9-23).
And where the motivation to combine is the same as previously presented.
Claim(s) 1, 3-7, 10, 12, 15, and 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over El Dokor, in view of Kale, and further in view of Powderly et al. (U.S. PG Pub No. 2018/0307303), hereinafter Powderly.
Regarding claim 1, El Dokor teaches
A method comprising (method of the present invention (12:45-54)):
determining, using one or more first …models and based at least on audio data obtained using one or more microphones of a vehicle, that speech represented by the audio data indicates an intent associated with a point of interest (POI) located external to the vehicle (the voice recognition module receives an output signal from the microphone in the vehicle that captures voice signals, i.e. audio data obtained using one or more microphones of a vehicle that speech represented by the audio data, and performs a voice recognition algorithm on the received signal to identify voice commands that provide additional information about the object being identified, i.e. determining, using one or more first …models…an intent associated with a point of interest (POI), such as a skyscraper in the city the vehicle is driving through, i.e. a point of interest (POI) located external to the vehicle, and that the person spoke the word “building” to identify the particular object, i.e. one or more words that describe the POI/one or more parameters for one or more slots corresponding to the intent (3:66-4:10),(6:12-19),(7:28-43),(8:1-19),(10:11-25));
determining, based at least on image data obtained using one or more image sensors of the vehicle, a … direction of a user within an interior of the vehicle (a camera system captures images of a gesture performed by a user in the vehicle, i.e. determining based at least on image data obtained using one or more image sensors of the vehicle, and the gesture recognition module analyzes the gesture to determine a direction vector representing the direction of the gesture, i.e. determining…a direction of the user within the interior of the vehicle (5:64-6:11));
mapping, based at least on localizing the vehicle with respect to a map of an environment, a projection of the … direction onto the map of the environment (the input analysis module receives input data from the gesture recognition module and the location module to identify a target region and micromap with POIs, using the location, orientation, and speed of the vehicle to automatically identify micromaps, i.e. based at least on localizing the vehicle with respect to a map of an environment, where a three-dimensional target region like a cone may be used to retrieve information within the micromapped region by overlapping the target region with the micromapped region, i.e. mapping…a projection of the direction onto the map of the environment (5:64-6:11),(6:20-67),(7:1-17,28-43),(8:20-62),(9:3-7,28-33),(10:11-51));
determining, based at least the projection of the … direction on the map at least partially overlapping with a location of the POI as represented by the map, contextual information from at least the map that is associated with the POI (the input analysis module receives input data from the gesture recognition module and the location module to identify a target region to overlay on a micromap with POIs, i.e. based at least the projection of the…direction on the map…overlapping with…the map, where a POI within the target region is identified, i.e. at least partially overlapping with a location of the POI, and a POI search for information related to a POI and stored in the micromap is performed, i.e. contextual information from at least the map that is associated with the POI (5:64-6:11),(6:20-67),(7:1-17,28-43),(8:20-62),(9:3-33),(10:11-51));
determining, based at least on one or more second…models processing first data representative of the intent and second data representative of the contextual information associated with the POI, an output that includes at least a portion of the contextual information (the input analysis module receives input data, including from the voice recognition module, i.e. first data representative of the intent and the one or more words describing the POI, and the POI search module queries a databases to retrieve information related to POI, i.e. second data representative of the contextual information associated with the POI, which can be merged by an artificial intelligence unit, and which can be received by the data output module for output to the user, i.e. determining…based at least on processing first data…and second data…an output that includes at least a portion of the contextual information (6:12-19),(6:31-51),(7:6-17),(8:47-9:27),(9:27-67)); and
causing the vehicle to provide the output associated with the speech (the data output module receives information related to the POI from the search modules, and sends the information to the in-vehicle communications system or display, i.e. causing the vehicle to provide the output, in response to the voice command and gesture, i.e. associated with the speech (5:64-6:19),(7:6-17),(9:51-60)).
While El Dokor provides recognizing voice commands using an algorithm and providing information to a user in response to the command, El Dokor does not specifically teach identifying intent using a machine learning model or using a machine learning model to process two types of data to determine an output, and thus does not teach
one or more first machine learning models…;
determining, based at least on one or more second machine learning models processing first data representative of the intent and the one or more words describing the POI and second data representative of the contextual information associated with the POI, an output that includes at least a portion of the contextual information.
Kale, however, teaches one or more first machine learning models…(the artificial intelligence framework includes an NLU component, i.e. one or more first machine learning models, that operates to extract user intent and various intent parameters (1:60-2:7),(7:63-67),(17:34-45));
determining, based at least on one or more second machine learning models processing first data representative of the intent and the one or more words describing the POI and second data representative of the contextual information associated with the POI, an output that includes at least a portion of the contextual information (the NLG component generates language for a textual or spoken reply, i.e. determining…an output, based on the decisions made by the artificial intelligence framework, i.e. based at least on the one or more second machine learning models, where the NLU component may parse user inputs to determine the user intent and intent-related parameters, such as a dominant object and a variety of attributes related to that object, such as looking for a dress for a wedding in June in Italy, i.e. first data representative of the intent and the one or more words describing the POI, where the result of the search may include an item list of candidate products, i.e. second data representative of the contextual information associated with the POI, and the output can include the search results, i.e. output that includes at least a portion of the contextual information (1:60-2:7),(10:12-27,50-65),(17:23-18:28),(20:2-5,14-36)).
El Dokor and Kale are analogous art because they are from a similar field of endeavor in providing responses to users based on multimodal input. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the recognizing voice commands using an algorithm and providing information to a user in response to the command teachings of El Dokor with the use of an artificial intelligence framework to determine a reply and generate a response using the intent and various parameters as taught by Kale. It would have been obvious to combine the references to learn from user intents to enhance user understanding and provide an improved user experience (Kale (1:60-2:7)).
While El Dokor in view of Kale provides determining a direction of a user gesture based on images, El Dokor in view of Kale does not specifically teach determining a gaze direction, and thus does not teach
determining, based at least on image data obtained using one or more image sensors of the vehicle, a gaze direction of a user…;
mapping,…, a projection of the gaze direction onto the map of the environment;
determining, based at least the projection of the gaze direction on the map at least partially overlapping with a location of the POI as represented by the map, contextual information from at least the map that is associated with the POI.
Powderly, however, teaches determining, based at least on image data obtained using one or more image sensors…, a gaze direction of a user… (the wearable system can include eye tracking cameras oriented toward the eyes of the user to triangulate eye vectors to track where the user is looking [0120],[0123]);
mapping,…, a projection of the gaze direction onto the map of the environment (the wearable system can use cone casting to determine which objects are along the direction of the user’s eye pose, i.e. mapping,…, a projection of the gaze direction, where the user’s FOR may be part of a world map, i.e. onto the map of the environment [0111],[0113-4],[0120],[0123],[0126],[0138]);
determining, based at least the projection of the gaze direction on the map at least partially overlapping with a location of the POI as represented by the map, …information … associated with the POI (the wearable system can use cone casting to determine which objects are along the direction of the user’s eye pose, where the user’s FOR may be part of a world map, i.e. based at least the projection of the gaze direction on the map at least partially overlapping with a location of the POI as represented by the map, where information regarding the target object is identified, such as the location, i.e. information … associated with the POI [0111],[0113-4],[0120-1],[0123],[0126],[0138]).
El Dokor, Kale, and Powderly are analogous art because they are from a similar field of endeavor in providing responses to users based on multimodal input. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the determining a direction of a user gesture based on images teachings of El Dokor, as modified by Kale, with using eye tracking with cone projection and map information to identify an object of focus for the user as taught by Powderly. It would have been obvious to combine the references to enable the use of multimodal inputs to reduce the degree of specificity required in an input command and to reduce error rate associated with an imprecise command (Powderly [0052]).
Regarding claim 3, El Dokor in view of Kale and Powderly teaches claim 1, and El Dokor further teaches
the contextual information includes at least an identifier associated with a landmark as represented by the map (the POI information, i.e. contextual information, may include a name for the POI, such as a skyscraper, i.e. at least an identifier associated with a landmark (6:52-64),(8:47-62),(10:11-25)).
Regarding claim 4, El Dokor in view of Kale and Powderly teaches claim 1, and El Dokor further teaches
determining at least one of a geographic area associated with the user or a time period (location sensors are physical sensors of an in-vehicle system, and the output data used by the location module to determine the current location and orientation of the vehicle that has a driver or passengers, i.e. user, where the location, orientation, and speed are related to a micromap the vehicle is likely to enter, i.e. determining at least one of a geographic area associated with the user/environment (3:14-24,53-65),(6:20-30),(6:65-7:5)), wherein the determining of the output associated with the speech is further based at least on the at least one of the geographic area or the time period (the input analysis module receives input data, including from the voice recognition module, i.e. associated with the speech, and the POI or micromap search module queries a databases to retrieve information related to POI within the target region for output, i.e. determining of the output…is further based at least on the at least one of the geographic area (6:12-19,31-51),(7:6-17)).
Regarding claim 5, El Dokor in view of Kale and Powderly teaches claim 1, and Powderly further teaches
receiving second image data representing an image depicting the environment (FOV cameras take images of the field of view of the user, i.e. an image depicting the environment Fig. 9, 22A-B,[0113]); and
determining, based at least on the second image data, additional information associated with the POI, wherein the second data further represents the additional information (object recognizers may crawl through collected points based on the user’s FOV camera, i.e. based at least on the second image data, to recognize one or more objects and convey information to the system, such as orientation and position in relation to various objects and other surroundings, as well as other parameters, such as identifying a red chair, i.e. determining…additional information associated with the POI wherein the second data further represents the additional information Fig. 9, 22A-B,[0113],[0121]).
Where the motivation to combine is the same as previously presented.
Regarding claim 6, El Dokor in view of Kale and Powderly teaches claim 1, and El Dokor further teaches
determining, based at least on the audio data, the one or more or more words are associated with one or more parameters for one or more slots associated with the intent, wherein the first data further represents the one or more words (the voice recognition module receives an output signal from the microphone in the vehicle that captures voice signals and performs a voice recognition algorithm on the received signal, i.e. determining based at least on the audio data, to identify voice commands, i.e. intent, that provide additional information about the object being identified, i.e. one or more parameters for one or more slots associated with the intent, and outputs the words as a character string that can be used by the search modules in a query to obtain more accurate results, such as ignoring information for non-building objects when the user says “building”, where the input analysis module receives input data, including from the voice recognition module, for querying a database to retrieve information, i.e. one or more words are associated with one or more parameters for one or more slots associated with the intent wherein the first data further represents the one or more words (2:42-52),(3:66-4:10),(6:12-19,31-51),(7:6-17), (8:1-12),(8:47-9:27),(9:27-67)).
Regarding claim 7, El Dokor in view of Kale and Powderly teaches claim 1, and Powderly further teaches
determining that the user includes the gaze direction for at threshold period of time (the user’s eyes can be tracked to determine if the user’s gaze at an object is for longer than a threshold time, where cone casting techniques are used to determine objects along the direction of the user’s eye pose [0125-6]),
wherein the determining the contextual information from the at least the map is further based at least on the user including the gaze direction for the threshold period of time (the user’s eyes can be tracked to determine if the user’s gaze at an object is for longer than a threshold time, at which point the object is selected as the user input or target object with which the user would like to interact, such as a POI, i.e. determining…information…is further based at least on the user including the gaze direction for the threshold period of time [0125-6],[0148],[0200-2]).
Where El Dokor teaches that the POI has associated information stored within the micromap and received once the POI is identified (8:59-9:60).
And where the motivation to combine is the same as previously presented.
Regarding claim 10, El Dokor in view of Kale and Powderly teaches claim 1, and El Dokor further teaches
the output associated with the speech comprises at least one of:
second audio data representing one or more second words that provide the contextual information (the data output module speaks out information related to the point of interest, such as the name of, i.e. second audio data representing one or more second words that provide, the restaurant or building, i.e. contextual information (2:42-53),(6:12-19,41-51),(7:6-17),(8:1-12),(9:50-67));
or content data representing one or more images depicting content associated with the contextual information (the data output module may display information, such as reviews and photos of the restaurant, or information about the building, i.e. content data representing one or more images depicting content associated with the contextual information (2:42-53),(6:12-19,41-51),(7:6-17),(8:1-12),(9:50-67),(10:1-25)).
Regarding claim 12, El Dokor in view of Kale teaches claim 11.
While El Dokor in view of Kale provides determining a direction of a user gesture based on images, El Dokor in view of Kale does not specifically teach determining POIs from images external to the vehicle, and thus does not teach
determine, based at least on second image data representing one or more images depicting the environment exterior the vehicle, that the one or more images depict the POI; and
determining that the POI represented by the map includes a same POI as the POI depicted by the one or more images,
wherein the determination that the user is indicating the POI within the environment is based at least on the POI represented by the map including the same POI as the POI depicted by the one or more images.
Powderly, however, teaches determine, based at least on second image data representing one or more images depicting the environment exterior the vehicle, that the one or more images depict the POI (object recognizers may crawl through collected points based on the user’s FOV camera, i.e. based at least on second image data representing one or more images depicting the environment, to recognize one or more objects and convey information to the system, such as orientation and position in relation to various objects and other surroundings, as well as other parameters, such as identifying a red chair, i.e. determine…that the one or more images depict the POI Fig. 9, 22A-B,[0113],[0121]); and
determining that the POI represented by the map includes a same POI as the POI depicted by the one or more images (objects within the FOV can be recognized using a map database [0113]),
wherein the determination that the user is indicating the POI within the environment is based at least on the POI represented by the map including the same POI as the POI depicted by the one or more images (a user’s gaze on an object within the FOV, which is populated with objects identified by object recognizers, i.e. based at least on the POI represented by the map including the same POI as the POI depicted by the one or more images, may be identified as the object related to the user input, i.e. the determination that the user is indicating the POI within the environment [0113],[0126]).
El Dokor, Kale, and Powderly are analogous art because they are from a similar field of endeavor in providing responses to users based on multimodal input. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the determining a direction of a user gesture based on images teachings of El Dokor, as modified by Kale, with using eye tracking with cone projection, object recognition from FOV cameras, and map information to identify an object of focus for the user as taught by Powderly. It would have been obvious to combine the references to enable the use of multimodal inputs to reduce the degree of specificity required in an input command and to reduce error rate associated with an imprecise command (Powderly [0052]).
Regarding claim 15, El Dokor in view of Kale teaches claim 11, and El Dokor further teaches
determining, based at least on the projection of the direction at least partially overlapping the POI represented by the map, a … confidence … associated with the POI (the input analysis module receives input data from the gesture recognition module and the location module to identify a target region to overlay on a micromap with POIs, i.e. based at least the projection of the direction on the map…overlapping…the map, where a POI of multiple POIs within the target region is identified, i.e. at least partially overlapping with a POI, where there may be uncertainty as to which POI the user was attempting to identify, such as two buildings in close proximity, i.e. determining…a confidence associated with the POI (5:64-6:11),(6:20-67),(7:1-17,28-43),(8:20-62),(9:3-33),(10:11-51));
determining, based at least on a second projection of the direction at least partially overlapping a second POI represented by a second map, a second confidence … associated with the second POI (the input analysis module receives input data from the gesture recognition module and the location module to identify a target region to overlay on retrieved micromaps with POIs, where target regions can be adjusted, i.e. based at least on a second projection of the direction on the map…overlapping…a second map, where a POI of multiple POIs within the target region is identified, i.e. at least partially overlapping with a second POI, where there may be uncertainty as to which POI the user was attempting to identify, such as two buildings in close proximity, i.e. determining…a second confidence associated with the second POI (5:64-6:11),(6:20-67),(6:65-7:17),(7:28-43),(8:20-62),(9:3-33),(10:11-51).
While El Dokor in view of Kale provides identifying multiple POIs, El Dokor in view of Kale does not specifically teach determining two confidence scores and determining the POI based on the confidence scores, and thus does not teach
determining…a first confidence score associated with the POI;
determining…a second confidence score associated with the second POI;
determining, based at least on the first confidence score and the second confidence score, that the user is indicating the POI within the environment.
Powderly, however, teaches determining…a first confidence score associated with the POI (each candidate object may be associated with a confidence score [0173]);
determining…a second confidence score associated with the second POI (each candidate object may be associated with a confidence score [0173]);
determining, based at least on the first confidence score and the second confidence score, that the user is indicating the POI within the environment (the candidate object with the highest confidence score is selected by the system as the target object [0173]).
El Dokor, Kale, and Powderly are analogous art because they are from a similar field of endeavor in providing responses to users based on multimodal input. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the identifying multiple POIs teachings of El Dokor, as modified by Kale, with calculating confidence scores for each candidate object and choosing the object with the highest score as the target object as taught by Powderly. It would have been obvious to combine the references to enable the use of multimodal inputs to reduce the degree of specificity required in an input command and to reduce error rate associated with an imprecise command (Powderly [0052]).
Regarding claim 22, El Dokor in view of Kale and Powderly teaches claim 1, and El Dokor further teaches
at least one of the determining the gaze direction of the user or the determining the contextual information using the map is further based at least on the intent associated with the speech (the voice recognition module receives an output signal from the microphone in the vehicle that captures voice signals and performs a voice recognition algorithm on the received signal to identify voice commands that provide additional information about the object being identified, i.e. the intent associated with the speech, and outputs the words as a character string that can be used by the search modules in a query to obtain more accurate results, such as ignoring information for non-building objects when the user says “building”, when the POI information gathered is stored in a micromap, i.e. determining the contextual information using the map is further based at least on the intent associated with the speech (2:42-52),(3:66-4:10),(6:5-19,41-64),(8:1-12)).
Claim(s) 2 is/are rejected under 35 U.S.C. 103 as being unpatentable over El Dokor, in view of Kale, in view of Powderly, and further in view of Zhao et al. (U.S. PG Pub No. 2021/0248376), hereinafter Zhao.
Regarding claim 2, El Dokor in view of Kale and Powderly teaches claim 1.
While El Dokor in view of Kale and Powderly provides the use of vectors, El Dokor in view of Kale and Powderly does not specifically teach that the first data and second data represent vectors, and thus does not teach
the first data represents one or more first vectors corresponding to at least the intent and the one or more words that describe the POI and the second data represents one or more second vectors corresponding to at least the contextual information associated with the POI.
Zhao, however, teaches the first data represents one or more first vectors corresponding to at least the intent and the one or more words that describe the POI and the second data represents one or more second vectors corresponding to at least the contextual information associated with the POI (the system extracts a query vector, generates textual-context vectors, i.e. first data represents one or more first vectors corresponding to at least the intent and the one or more words, as well as generates candidate-response vectors representing candidate responses to the question, i.e. the second data represents one or more second vectors corresponding to at least the contextual information [0023-5]).
Where Kale teaches that the information is related to the intent, parameters, and results (1:60-2:7),(10:12-27,50-65), and El Dokor teaches that all the information is related specifically to a POI (2:24-41).
El Dokor, Kale, Powderly, and Zhao are analogous art because they are from a similar field of endeavor in providing responses to users based on multimodal input. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the use of vectors teachings of El Dokor, as modified by Kale and Powderly, with the use of vectors to describe information processed by the system as taught by Zhao. It would have been obvious to combine the references to enable analyzing both audio and visual cues to provide answers to user questions with accuracy and multiple contextual modes for questions (Zhao [0021]).
Claim(s) 18-20 and 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Panchaksharaiah et al. (U.S. PG Pub No. 2024/0073518), hereinafter Panchaksharaiah, in view of El Dokor, and further in view of Kale.
Regarding claim 18, Panchaksharaiah teaches
One or more processors comprising ([0026]):
processing circuitry to ([0026]):
determine, based at least on audio data obtained using one or more microphones … located within an environment, at least an intent associated with speech and one or more parameters for one or more slots corresponding to the intent (the user utters a query made to a digital assistant, such as through a microphone, i.e. based at least on audio data obtained using one or more microphones located within an environment, where the voice query is processed to determine an intent, i.e. determine…at least an intent associated with speech, as well as classifying portions of the query, or a query type that references qualities of an object or filtering results with portions of the query that are ambiguous, such as buying a doll of a particular size, i.e. one or more parameters for one or more slots corresponding to the intent [0013],[0020],[0037-40],[0044]);
determine, based at least on first image data obtained using one or more internal image sensors …, a direction indicated by a user …(a camera is activated to capture, i.e. based at least on first image data obtained using one or more internal image sensors, the gesture made by the user that is directed at a particular object, i.e. determine…a direction indicated by a user [0013-4],[0019]);
determine that the direction indicated by the user is associated with a point of interest (POI) that is not represented by a map (a camera is activated to capture the target object, i.e. a point of interest (POI) that is not represented by a map, that is in the direction of the user’s pointing gesture, i.e. determine that the direction indicated by the user is associated with a point of interest (POI) [0013-4],[0018-9]) ;
based at least on the POI not being represented by the map, determine, based at least on projecting the direction indicated by the user to one or more images represented by second image data obtained using one or more external image sensors …, contextual information associated with the POI (a camera is activated to capture the target object, i.e. based at least on the POI not being represented by the map, that is in the direction of the user’s pointing gesture, i.e. based at least on projecting the direction indicated by the user to one or more images represented by second image data obtained using one or more external image sensors, visual models can extract contextual information from analyzing the image, such as determine that the user is pointing to an oil bottle with a POMPEIAN label, i.e. determine…contextual information associated with the POI [0013-4],[0018-21],[0031],[0036]).
While Panchaksharaiah provides determining information about a target object from a voice query and images, Panchaksharaiah does not specifically teach a vehicle or providing an output to the user in response to the query, and thus does not teach
of a vehicle;
determine, based at least on one or more machine learning models processing first data representative of the intent and the one or more parameters and second data representative of the contextual information, an output associated with the speech that includes at least a portion of the contextual information; and
cause the vehicle to provide the output associated with the speech.
El Dokor, however, teaches of a vehicle (the voice recognition module receives an output signal from the microphone in the vehicle that captures voice signals, i.e. receive audio data obtained using one or more microphones of a vehicle (3:66-4:10), a camera system captures images of a gesture performed by a user in the vehicle, i.e. determining based at least on image data obtained using one or more image sensors of the vehicle, and the gesture recognition module analyzes the gesture to determine a direction vector representing the direction of the gesture, i.e. determine…a direction indicated by a user located within an interior of the vehicle (5:64-6:11));
determine, based at least on one or more … models processing first data representative of the intent and the one or more parameters and second data representative of the contextual information, an output associated with the speech that includes at least a portion of the contextual information (the input analysis module receives input data, including from the voice recognition module, i.e. first data representative of the intent and the one or more words describing the POI, and the POI search module queries a databases to retrieve information related to POI, i.e. second data representative of the contextual information associated with the POI, which can be merged by an artificial intelligence unit, and which can be received by the data output module for output to the user, i.e. determining…based at least on processing first data…and second data…an output that includes at least a portion of the contextual information (6:12-19),(6:31-51),(7:6-17),(8:47-9:27),(9:27-67)); and
cause the vehicle to provide the output associated with the speech (the data output module receives information related to the POI from the search modules, and sends the information to the in-vehicle communications system or display, i.e. causing the vehicle to provide the output, in response to the voice command and gesture, i.e. associated with the speech (5:64-6:19),(7:6-17),(9:51-60)).
Panchaksharaiah and El Dokor are analogous art because they are from a similar field of endeavor in using gestures to identify target objects for a query. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the determining information about a target object from a voice query and images teachings of Panchaksharaiah with the use of the resulting information to provide the user with a response to the query as taught by El Dokor. It would have been obvious to combine the references to allow a user to rapidly identify and retrieve information about a POI near the vehicle without having to navigate a user interface by manipulating a touchscreen or physical buttons (El Dokor (2:24-41)).
While Panchaksharaiah in view of El Dokor provides recognizing voice commands using an algorithm and providing information to a user in response to the command, Panchaksharaiah in view of El Dokor does not specifically teach identifying using a machine learning model to process the two types of data to determine an output, and thus does not teach
determine, based at least on one or more machine learning models processing first data representative of the intent and the one or more parameters and second data representative of the contextual information, an output associated with the speech that includes at least a portion of the contextual information.
Kale, however, teaches determine, based at least on one or more machine learning models processing first data representative of the intent and the one or more parameters and second data representative of the contextual information, an output associated with the speech that includes at least a portion of the contextual information (the NLG component generates language for a textual or spoken reply, i.e. determine…an output, based on the decisions made by the artificial intelligence framework, i.e. based at least on the one or more machine learning models, where the NLU component may parse user inputs to determine the user intent and intent-related parameters, such as a dominant object and a variety of attributes related to that object, such as looking for a dress for a wedding in June in Italy, i.e. processing the first data representative of the intent and the one or more parameters, where the result of the search may include an item list of candidate products, i.e. processing…the second data representative of the contextual information, and the output can include the search results, i.e. output associated with the speech that includes at least a portion of the contextual information (1:60-2:7),(10:12-27,50-65),(17:23-18:28),(20:2-5,14-36)).
Panchaksharaiah, El Dokor, and Kale are analogous art because they are from a similar field of endeavor in providing responses to users based on multimodal input. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the recognizing voice commands using an algorithm and providing information to a user in response to the command teachings of Panchaksharaiah, as modified by El Dokor, with the use of an artificial intelligence framework to determine a reply and generate a response using the intent and various parameters as taught by Kale. It would have been obvious to combine the references to learn from user intents to enhance user understanding and provide an improved user experience (Kale (1:60-2:7)).
Regarding claim 19, Panchaksharaiah in view of El Dokor and Kale teaches claim 18, and Panchaksharaiah further teaches
determine, based at least on the first image data, that the speech is associated with the user of one or more users located within the interior of the vehicle (the camera may capture initial frames in a standard mode, i.e. based at least on the first image data, to identify the user as the source of the voice query, i.e. determine…that the speech is associated with the user of one or more users [0018],[0032],[0044]),
wherein the determination of the direction indicated by the user is based at least on the speech being associated with the user (subsequent frames are captured to extract details about the target, such as a user’s pointing gesture being directed to the target object, i.e. wherein the determination of the directed indicated by the user, where multiple users may be looking at different objects, but only one user issues a query and is identified as the source of the voice query, i.e. based at least on the speech being associated with the user [0018-9],[0044]).
Where El Dokor specifically teaches the vehicle (3:66-4:10),(5:64-6:11).
And where the motivation to combine is the same as previously presented.
Regarding claim 20, Panchaksharaiah in view of El Dokor and Kale teaches claim 18, and El Dokor further teaches
the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for the autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for presenting virtual reality (VR) content;
a system for presenting augmented reality (AR) content;
a system for presenting mixed reality (MR) content;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing deep learning operations;
a system implemented using an edge device (the system may be performed by a device in the operating environment, such as an in-vehicle communications system, a wireless mobile communication device, and/or a remote server, i.e. edge devices (2:54-3:13));
a system implemented using a robot;
a system for performing conversational AI operations;
a system for generating synthetic data;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center (the system may be performed by a device in the operating environment, such as an in-vehicle communications system, a wireless mobile communication device, and/or a remote server, and may further access a database on the server or a service operating on a third-party server, such as Yelp or Google, i.e. at least partially in a data center (2:54-3:13),(6:58-51));
or a system implemented at least partially using cloud computing resources (the system may be performed by a device in the operating environment, such as an in-vehicle communications system, a wireless mobile communication device, and/or a remote server, i.e. at least partially using cloud computing resources (2:54-3:13).
Where the motivation to combine is the same as previously presented.
Regarding claim 24, Panchaksharaiah in view of El Dokor and Kale teaches claim 18, and Panchaksharaiah further teaches
determining, based at least on processing the second image data using one or more second machine learning models, that the one or more images represent the POI and the contextual information associated with the POI (a camera is activated to capture the target object that is in the direction of the user’s pointing gesture, i.e. second image data…one or more images, and machine learning visual models can extract, i.e. based at least on processing the second image data using one or more second machine learning models, contextual information from analyzing the image, such as determining that the user is pointing to an oil bottle with a POMPEIAN label, i.e. determine…the contextual information associated with the POI, where one of multiple objects can be identified as the target object, i.e. determine…that the one or more images represent the POI [0013-4],[0018-21],[0031],[0036],[0044]);
determining, based at least on projecting the direction to the one or more images, that the direction is towards the POI within the environment (a camera is activated to capture the target object, i.e. a point of interest (POI) that is not represented by a map, that is in the direction of the user’s pointing gesture, i.e. determine that the direction indicated by the user is associated with a point of interest (POI) [0013-4],[0018-9]); and
determining the contextual information based at least on the direction being towards the POI (a camera is activated to capture the target object that is in the direction of the user’s pointing gesture, and machine learning visual models can extract contextual information from analyzing the image, such as determining that the user is pointing to an oil bottle with a POMPEIAN label, i.e. determine…determining the contextual information based at least on the direction being towards the POI [0013-4],[0018-21],[0031],[0036],[0044]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICOLE A K SCHMIEDER whose telephone number is (571)270-1474. The examiner can normally be reached 8:00 - 5:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571) 272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NICOLE A K SCHMIEDER/Primary Examiner, Art Unit 2659