DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation - 35 USC § 101
The limitations “provide one or more navigational commands to control navigation to a destination within a graphical user interface associated with the 3D model responsive to the verbal input.” Are considered a practical application of providing a navigational aid to a user via a display.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1, 4, 9, 10, 11, 12, 15, 20-23 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Li et al. (US 2023/0128422)(Hereinafter referred to as Li).
Regarding claim 1, Li teaches A system (FIG. 1 illustrates an example network environment 100 associated with an assistant system. Network environment 100 includes a client system 130, an assistant system 140, a social-networking system 160, and a third-party system 170 connected to each other by a network 110. See paragraph [0051])( FIG. 2 illustrates an example architecture 200 of the assistant system 140. In particular embodiments, the assistant system 140 may assist a user to obtain information or services. The assistant system 140 may enable the user to interact with the assistant system 140 via user inputs of various modalities ( e.g., audio, voice, text, vision, image, video, gesture, motion, activity, location, orientation) in stateful and multi-tum conversations to receive assistance from the assistant system 140. See paragraph [0071]) comprising:
one or more processors (The assistant system 140 may enable the user to interact with the assistant system 140 via user inputs of various modalities ( e.g., audio, voice, text, vision, image, video, gesture, motion, activity, location, orientation) in stateful and multi-tum conversations to receive assistance from the assistant system 140. See paragraph [0071])(It is clear there are processors and memory as the assistant system in taking inputs and executing instructions based on the inputs )( User inputs provided by a user may be associated with particular assistant-related tasks, and may include, for example, user requests (e.g., verbal requests for information or performance of an action), user interactions with the assistant application 136 associated with the assistant system 140 ( e.g., selection of UI elements via touch or gesture), or any other type of suitable user input that may be detected and understood by the assistant system 140 (e.g., user movements detected by the client device 130 of the user). See paragraph [0071]); and
memory containing instructions to control the one or more processors (The assistant system 140 may enable the user to interact with the assistant system 140 via user inputs of various modalities ( e.g., audio, voice, text, vision, image, video, gesture, motion, activity, location, orientation) in stateful and multi-tum conversations to receive assistance from the assistant system 140. See paragraph [0071])(It is clear there are processors and memory as the assistant system in taking inputs and executing instructions based on the inputs )( User inputs provided by a user may be associated with particular assistant-related tasks, and may include, for example, user requests (e.g., verbal requests for information or performance of an action), user interactions with the assistant application 136 associated with the assistant system 140 ( e.g., selection of UI elements via touch or gesture), or any other type of suitable user input that may be detected and understood by the assistant system 140 (e.g., user movements detected by the client device 130 of the user). See paragraph [0071])to:
receive a 3D digital model representing a physical environment (Augmented reality can be defined as a system that incorporates three basic features: a combination of real and virtual worlds, real-time interaction, and accurate 3D registration of virtual and real objects. The overlaid sensory information can be constructive (i.e. additive to the natural environment), or destructive (i.e. masking of the natural environment). This experience is seamlessly interwoven with the physical world such that it is perceived as an immersive aspect of the real environment. See paragraph [0005]);
receive a first user input, the first user input including an first verbal input to control navigation to a destination within the 3D model (In particular embodiments, a client system may implement voice commands within an augmented reality (AR), virtual reality (VR), mixed reality (MR), or extended reality (XR) environment via a voice SDK, which allows XR applications to easily integrate voice commands on XR devices (e.g., the client system). In particular embodiments, XR may be one or more of a combination of AR, VR, or MR. As an example and not by way of limitation, XR may include elements of AR or VR or MR. In particular embodiments, the client system may implement voice commands in combination with gestures within XR environments via a voice SDK. As XR content becomes more immersive, voice integration may make the content more engaging to interact with various applications. That is, people typically interact with the real world using their voice and hands. Currently, navigating through XR content may be cumbersome as the user may need to navigate through unfamiliar nested menus. As an example and not by way oflimitation, a user may need to click through various virtual menus by ray-casting from a controller or inputting text through a virtual keyboard to navigate to a desired destination. To improve upon the user experience, the XR system may implement voice commands, combined with one or more other modalities ( e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0009])( The solution presented by the embodiments disclosed herein to address this challenge may be the XR system may implement voice commands, combined with one or more other modalities ( e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/ search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0011])( As an example and not by way of limitation, the XR display device may render a navigation menu to access information and/or other parts of the application associated with the XR environment the user is located. In particular embodiments, the XR display device may render, for one or more displays of the XR display device, information corresponding to frequently asked questions responsive to executing the first task. Although this disclosure describes rendering one or more elements in a particular manner, this disclosure contemplates rendering one or more elements in any suitable manner. See paragraph [0172]);
translate the first verbal input into a first text query using a first machine learning model (In particular embodiments, the user input may comprise speech input. The speech input may be received at the ASR module 208 for extracting the text transcription from the speech input. The ASR module 208 may use statistical models to determine the most likely sequences of words that correspond to a given portion of speech received by the assistant system 140 as audio input. The models may include one or more of hidden Markov models, neural networks, deep learning models, or any combination thereof. The received audio input may be encoded into digital data at a particular sampling rate ( e.g., 16, 44.1, or 96 kHz) and with a particular number of bits representing each sample ( e.g., 8, 16, of 24 bits). See paragraph [0134]);
analyze, by a second machine learning model, the first text query to determine a desired navigation (the text transcription from the ASR module 208 may be sent to the NLU module 210. The NLU module 210 may process the text transcription and extract the user intention (i.e., intents) and parse the slots or parsing result based on the linguistic ontology. In particular embodiments, the intents and slots from the NLU module 210 and/or the events and contexts from the context engine 220 may be sent to the entity resolution module 212. See paragraph [0137])( As an example and not by way of limitation, the user input may comprise "direct me to my next meeting." The assistant system 140 may use a calendar agent to retrieve the location of the next meeting. The assistant system 140 may then use a navigation agent to direct the user to the next meeting. See paragraph [0126])( FIG. 7 illustrates an example flow diagram of processing an audio input. In particular embodiments, a client system 702 ( e.g., client system 130) may be embodied as a XR display device, such as an AR headset or a VR headset. The NLU model may match the text to intent and extract entities for the slots. In particular embodiments, the NLU model 714 may send the results back to the application 706. In particular embodiments, the NLU model 714 may send the intents, slots, and entities to a dialog manager 716. The dialog manager 716 may resolve the intents and slots received from the NLU model. In particular embodiments, the assistant system 704 may perform a task or send instructions to execute a task to the client system 702. The dialog manager 716 may send instructions to the application 706 to perform a task based on the received intents, slots, and entities from the NLU model. In particular embodiments, the assistant system 704 may generate a result from processing the intents, slots, and entities from the NLU model 714 and render a response through the text-to-speech module 718. See paragraph [0176] ); and
provide one or more navigational commands to control navigation to a destination within a graphical user interface associated with the 3D model responsive to the verbal input ( As an example and not by way of limitation, the user input may comprise "direct me to my next meeting." The assistant system 140 may use a calendar agent to retrieve the location of the next meeting. The assistant system 140 may then use a navigation agent to direct the user to the next meeting. See paragraph [0126])( A XR system may use a language model to identify phrases from an audio input and compare that to a list of phrases associated with the application. The XR system may identify a corresponding gesture performed and compare that to a list associated with a combination of phrases and gestures. When a particular phrase and gesture combination is detected, then the XR system may present a corresponding output, such as a resulting wizard spell. The XR system may perform object recognition on real-world objects and allow the user to interact with the object (e.g., see a real-world menu, CV recognizes it, then point at an item on the menu and ask the assistant to tell you about it). See paragraph [0154])( the XR system may implement voice commands, combined with one or more other modalities (e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0152]).
Regarding claim 4, Li teaches The system of claim 1, wherein the first verbal response is generated based on context from the user (Para 0175, FIGS. 6A-6B illustrate an example flow diagram of processing an audio input. the process 600 may start at step 602, where a client system receives an utterance from a user. At step 604, the client system may use ASR and a natural language model to process the utterance. The client system may process the utterance to generate one or more intents and one or more entities. the client system may present an audio output to the user to respond to the user utterance. See paragraph [0175]) (A XR system may use a language model to identify phrases from an audio input and compare that to a list of phrases associated with the application. The XR system may identify a corresponding gesture performed and compare that to a list associated with a combination of phrases and gestures. When a particular phrase and gesture combination is detected, then the XR system may present a corresponding output, such as a resulting wizard spell. The XR system may perform object recognition on real-world objects and allow the user to interact with the object (e.g., see a real-world menu, CV recognizes it, then point at an item on the menu and ask the assistant to tell you about it) See paragraph [0155]).
Regarding claim 9, Lit teaches The system of claim 1, wherein the memory containing instructions to further control the one or more processors to: receive a second user input, the second first user input including a second verbal input to request information of an aspect of the physical environment (the one or more client systems may receive, by the one or more microphones of the one or more client systems, a second audio input comprising a second voice command of a second plurality of voice commands associated with a second application. In particular embodiments, the second plurality of voice commands may comprise a third set of commands executable by a second customized on-device natural language understanding (NLU) model installed on the one or more client systems and a fourth set of commands executable by the server-side assistant system) See paragraph [0026]);
translate the second verbal input into a second text query using the first machine learning model (The speech input may be received at the ASR module 208 for extracting the text transcription from the speech input. The ASR module 208 may use statistical models to determine the most likely sequences of words that correspond to a given portion of speech received by the assistant system 140 as audio input. The models may include one or more of hidden Markov models, neural networks, deep learning models, or any combination thereof. See paragraph [0134])(the ASR module 208 may comprise one or more of a grapheme-to-phoneme (G2P) model, a pronunciation learning model, a personalized acoustic model, a personalized language model (PLM), or an end-pointing model. In particular embodiments, the grapheme-to-phoneme (G2P) model may be used to determine a user's grapheme-to phoneme style (i.e., what it may sound like when a particular user speaks a particular word). In particular embodiments, the personalized acoustic model may be a model of the relationship between audio signals and the sounds of phonetic units in the language. See Paragraph [0135]);
analyze, by the second machine learning model, the second text query to determine an inquiry result ((The speech input may be received at the ASR module 208 for extracting the text transcription from the speech input. The ASR module 208 may use statistical models to determine the most likely sequences of words that correspond to a given portion of speech received by the assistant system 140 as audio input. The models may include one or more of hidden Markov models, neural networks, deep learning models, or any combination thereof. See paragraph [0134]) (the ASR module 208 may comprise one or more of a grapheme-to-phoneme (G2P) model, a pronunciation learning model, a personalized acoustic model, a personalized language model (PLM), or an end-pointing model. In particular embodiments, the grapheme-to-phoneme (G2P) model may be used to determine a user's grapheme-to-phoneme style (i.e., what it may sound like when a particular user speaks a particular word). In particular embodiments, the personalized acoustic model may be a model of the relationship between audio signals and the sounds of phonetic units in the language). See paragraph [0135]); and
provide a second verbal response based on the inquiry result (( As an example and not by way of limitation, the user input may comprise "direct me to my next meeting." The assistant system 140 may use a calendar agent to retrieve the location of the next meeting. The assistant system 140 may then use a navigation agent to direct the user to the next meeting. See paragraph [0126])(A XR system may use a language model to identify phrases from an audio input and compare that to a list of phrases associated with the application. The XR system may identify a corresponding gesture performed and compare that to a list associated with a combination of phrases and gestures. When a particular phrase and gesture combination is detected, then the XR system may present a corresponding output, such as a resulting wizard spell. The XR system may perform object recognition on real-world objects and allow the user to interact with the object (e.g., see a real-world menu, CV recognizes it, then point at an item on the menu and ask the assistant to tell you about it). See paragraph [0154]) (the XR system may implement voice commands, combined with one or more other modalities (e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/search, voice FAQ, and voice driven gameplay & experiences). See paragraph [0152]).
Regarding claim 10, Li teaches the system of claim 9, wherein analyze the second text query to determine the inquiry result is based on data from external data sources (agents 228 a/b may provide several functionalities for the assistant system 140 including, for example, native template generation, task specific business logic, and querying external APls; Para 0126, The third-party agents may be associated with third-party providers that provide content objects and/or services hosted by the third-party system 170). See paragraph [0093]).
Regarding claim 11, Li teaches the system of claim 1, wherein the first verbal input is received from a chat session (An assistant system can provide information or services on behalf of a user based on a combination of user input, location awareness, and the ability to access information from a variety of online sources (such as weather conditions, traffic congestion, news, stock prices, user schedules, retail prices, etc.). The user input may include text (e.g., online chat), especially in an instant messaging application or other applications, voice, images, motion, or a combination of them. See paragraph [0003]) ( FIG. 2 illustrates an example architecture 200 of the assistant system 140. In particular embodiments, the assistant system 140 may assist a user to obtain information or services. The assistant system 140 may enable the user to interact with the assistant system 140 via user inputs of various modalities ( e.g., audio, voice, text, vision, image, video, gesture, motion, activity, location, orientation) in stateful and multi-tum conversations to receive assistance from the assistant system 140. See paragraph [0071]).
Regarding claim 12, Li teaches A non-transitory computer-readable medium comprising executable instructions, the executable instructions being executable by one or more processors to perform a method (FIG. 1 illustrates an example network environment 100 associated with an assistant system. Network environment 100 includes a client system 130, an assistant system 140, a social-networking system 160, and a third-party system 170 connected to each other by a network 110. See paragraph [0051])( FIG. 2 illustrates an example architecture 200 of the assistant system 140. In particular embodiments, the assistant system 140 may assist a user to obtain information or services. The assistant system 140 may enable the user to interact with the assistant system 140 via user inputs of various modalities ( e.g., audio, voice, text, vision, image, video, gesture, motion, activity, location, orientation) in stateful and multi-tum conversations to receive assistance from the assistant system 140. See paragraph [0071]) (The assistant system 140 may enable the user to interact with the assistant system 140 via user inputs of various modalities ( e.g., audio, voice, text, vision, image, video, gesture, motion, activity, location, orientation) in stateful and multi-tum conversations to receive assistance from the assistant system 140. See paragraph [0071])(It is clear there are processors and memory as the assistant system in taking inputs and executing instructions based on the inputs )( User inputs provided by a user may be associated with particular assistant-related tasks, and may include, for example, user requests (e.g., verbal requests for information or performance of an action), user interactions with the assistant application 136 associated with the assistant system 140 ( e.g., selection of UI elements via touch or gesture), or any other type of suitable user input that may be detected and understood by the assistant system 140 (e.g., user movements detected by the client device 130 of the user). See paragraph [0071]),
the method comprising:
receiving a 3D digital model representing a physical environment (Augmented reality can be defined as a system that incorporates three basic features: a combination of real and virtual worlds, real-time interaction, and accurate 3D registration of virtual and real objects. The overlaid sensory information can be constructive (i.e. additive to the natural environment), or destructive (i.e. masking of the natural environment). This experience is seamlessly interwoven with the physical world such that it is perceived as an immersive aspect of the real environment. See paragraph [0005]);
receiving a first user input, the first user input including an first verbal input to control navigation to a destination within the 3D model (In particular embodiments, a client system may implement voice commands within an augmented reality (AR), virtual reality (VR), mixed reality (MR), or extended reality (XR) environment via a voice SDK, which allows XR applications to easily integrate voice commands on XR devices (e.g., the client system). In particular embodiments, XR may be one or more of a combination of AR, VR, or MR. As an example and not by way of limitation, XR may include elements of AR or VR or MR. In particular embodiments, the client system may implement voice commands in combination with gestures within XR environments via a voice SDK. As XR content becomes more immersive, voice integration may make the content more engaging to interact with various applications. That is, people typically interact with the real world using their voice and hands. Currently, navigating through XR content may be cumbersome as the user may need to navigate through unfamiliar nested menus. As an example and not by way oflimitation, a user may need to click through various virtual menus by ray-casting from a controller or inputting text through a virtual keyboard to navigate to a desired destination. To improve upon the user experience, the XR system may implement voice commands, combined with one or more other modalities ( e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0009])( The solution presented by the embodiments disclosed herein to address this challenge may be the XR system may implement voice commands, combined with one or more other modalities ( e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/ search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0011])( As an example and not by way of limitation, the XR display device may render a navigation menu to access information and/or other parts of the application associated with the XR environment the user is located. In particular embodiments, the XR display device may render, for one or more displays of the XR display device, information corresponding to frequently asked questions responsive to executing the first task. Although this disclosure describes rendering one or more elements in a particular manner, this disclosure contemplates rendering one or more elements in any suitable manner. See paragraph [0172]);
translating the first verbal input into a first text query using a first machine learning model (In particular embodiments, the user input may comprise speech input. The speech input may be received at the ASR module 208 for extracting the text transcription from the speech input. The ASR module 208 may use statistical models to determine the most likely sequences of words that correspond to a given portion of speech received by the assistant system 140 as audio input. The models may include one or more of hidden Markov models, neural networks, deep learning models, or any combination thereof. The received audio input may be encoded into digital data at a particular sampling rate ( e.g., 16, 44.1, or 96 kHz) and with a particular number of bits representing each sample ( e.g., 8, 16, of 24 bits). See paragraph [0134]);
analyzing, by a second machine learning model, the first text query to determine a desired navigation (the text transcription from the ASR module 208 may be sent to the NLU module 210. The NLU module 210 may process the text transcription and extract the user intention (i.e., intents) and parse the slots or parsing result based on the linguistic ontology. In particular embodiments, the intents and slots from the NLU module 210 and/or the events and contexts from the context engine 220 may be sent to the entity resolution module 212. See paragraph [0137])( As an example and not by way of limitation, the user input may comprise "direct me to my next meeting." The assistant system 140 may use a calendar agent to retrieve the location of the next meeting. The assistant system 140 may then use a navigation agent to direct the user to the next meeting. See paragraph [0126])( FIG. 7 illustrates an example flow diagram of processing an audio input. In particular embodiments, a client system 702 ( e.g., client system 130) may be embodied as a XR display device, such as an AR headset or a VR headset. The NLU model may match the text to intent and extract entities for the slots. In particular embodiments, the NLU model 714 may send the results back to the application 706. In particular embodiments, the NLU model 714 may send the intents, slots, and entities to a dialog manager 716. The dialog manager 716 may resolve the intents and slots received from the NLU model. In particular embodiments, the assistant system 704 may perform a task or send instructions to execute a task to the client system 702. The dialog manager 716 may send instructions to the application 706 to perform a task based on the received intents, slots, and entities from the NLU model. In particular embodiments, the assistant system 704 may generate a result from processing the intents, slots, and entities from the NLU model 714 and render a response through the text-to-speech module 718. See paragraph [0176] ); and
providing one or more navigational commands to control navigation to a destination within a graphical user interface associated with the 3D model responsive to the verbal input ( As an example and not by way of limitation, the user input may comprise "direct me to my next meeting." The assistant system 140 may use a calendar agent to retrieve the location of the next meeting. The assistant system 140 may then use a navigation agent to direct the user to the next meeting. See paragraph [0126])( A XR system may use a language model to identify phrases from an audio input and compare that to a list of phrases associated with the application. The XR system may identify a corresponding gesture performed and compare that to a list associated with a combination of phrases and gestures. When a particular phrase and gesture combination is detected, then the XR system may present a corresponding output, such as a resulting wizard spell. The XR system may perform object recognition on real-world objects and allow the user to interact with the object (e.g., see a real-world menu, CV recognizes it, then point at an item on the menu and ask the assistant to tell you about it). See paragraph [0154])( the XR system may implement voice commands, combined with one or more other modalities (e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0152]).
Regarding claim 15, Li teaches The non-transitory computer-readable medium of claim 12, wherein the first verbal response is generated based on context from the user (Para 0175, FIGS. 6A-6B illustrate an example flow diagram of processing an audio input. the process 600 may start at step 602, where a client system receives an utterance from a user. At step 604, the client system may use ASR and a natural language model to process the utterance. The client system may process the utterance to generate one or more intents and one or more entities. the client system may present an audio output to the user to respond to the user utterance. See paragraph [0175]) (A XR system may use a language model to identify phrases from an audio input and compare that to a list of phrases associated with the application. The XR system may identify a corresponding gesture performed and compare that to a list associated with a combination of phrases and gestures. When a particular phrase and gesture combination is detected, then the XR system may present a corresponding output, such as a resulting wizard spell. The XR system may perform object recognition on real-world objects and allow the user to interact with the object (e.g., see a real-world menu, CV recognizes it, then point at an item on the menu and ask the assistant to tell you about it) See paragraph [0155]).
Regarding claim 20, Li teaches the non-transitory computer-readable medium of claim 12, wherein the method further comprises: receiving a second user input, the second first user input including a second verbal input to request information of an aspect of the physical environment(the one or more client systems may receive, by the one or more microphones of the one or more client systems, a second audio input comprising a second voice command of a second plurality of voice commands associated with a second application. In particular embodiments, the second plurality of voice commands may comprise a third set of commands executable by a second customized on-device natural language understanding (NLU) model installed on the one or more client systems and a fourth set of commands executable by the server-side assistant system) See paragraph [0026]);
translating the second verbal input into a second text query using the first machine learning model (The speech input may be received at the ASR module 208 for extracting the text transcription from the speech input. The ASR module 208 may use statistical models to determine the most likely sequences of words that correspond to a given portion of speech received by the assistant system 140 as audio input. The models may include one or more of hidden Markov models, neural networks, deep learning models, or any combination thereof. See paragraph [0134])(the ASR module 208 may comprise one or more of a grapheme-to-phoneme (G2P) model, a pronunciation learning model, a personalized acoustic model, a personalized language model (PLM), or an end-pointing model. In particular embodiments, the grapheme-to-phoneme (G2P) model may be used to determine a user's grapheme-to phoneme style (i.e., what it may sound like when a particular user speaks a particular word). In particular embodiments, the personalized acoustic model may be a model of the relationship between audio signals and the sounds of phonetic units in the language. See Paragraph [0135]);
analyzing, by the second machine learning model, the second text query to determine an inquiry result((The speech input may be received at the ASR module 208 for extracting the text transcription from the speech input. The ASR module 208 may use statistical models to determine the most likely sequences of words that correspond to a given portion of speech received by the assistant system 140 as audio input. The models may include one or more of hidden Markov models, neural networks, deep learning models, or any combination thereof. See paragraph [0134]) (the ASR module 208 may comprise one or more of a grapheme-to-phoneme (G2P) model, a pronunciation learning model, a personalized acoustic model, a personalized language model (PLM), or an end-pointing model. In particular embodiments, the grapheme-to-phoneme (G2P) model may be used to determine a user's grapheme-to-phoneme style (i.e., what it may sound like when a particular user speaks a particular word). In particular embodiments, the personalized acoustic model may be a model of the relationship between audio signals and the sounds of phonetic units in the language). See paragraph [0135]); and providing a second verbal response based on the inquiry result (( As an example and not by way of limitation, the user input may comprise "direct me to my next meeting." The assistant system 140 may use a calendar agent to retrieve the location of the next meeting. The assistant system 140 may then use a navigation agent to direct the user to the next meeting. See paragraph [0126])(A XR system may use a language model to identify phrases from an audio input and compare that to a list of phrases associated with the application. The XR system may identify a corresponding gesture performed and compare that to a list associated with a combination of phrases and gestures. When a particular phrase and gesture combination is detected, then the XR system may present a corresponding output, such as a resulting wizard spell. The XR system may perform object recognition on real-world objects and allow the user to interact with the object (e.g., see a real-world menu, CV recognizes it, then point at an item on the menu and ask the assistant to tell you about it). See paragraph [0154]) (the XR system may implement voice commands, combined with one or more other modalities (e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/search, voice FAQ, and voice driven gameplay & experiences). See paragraph [0152]).
Regarding claim 21, Li teaches The non-transitory computer-readable medium of claim 20, wherein the analyzing the second text query to determine the inquiry result is based on data from external data sources (agents 228 a/b may provide several functionalities for the assistant system 140 including, for example, native template generation, task specific business logic, and querying external APls; Para 0126, The third-party agents may be associated with third-party providers that provide content objects and/or services hosted by the third-party system 170). See paragraph [0093]).
Regarding claim 22, Li teaches The non-transitory computer-readable medium of claim 20, wherein the first verbal input is received from a chat session (An assistant system can provide information or services on behalf of a user based on a combination of user input, location awareness, and the ability to access information from a variety of online sources (such as weather conditions, traffic congestion, news, stock prices, user schedules, retail prices, etc.). The user input may include text (e.g., online chat), especially in an instant messaging application or other applications, voice, images, motion, or a combination of them. See paragraph [0003]) ( FIG. 2 illustrates an example architecture 200 of the assistant system 140. In particular embodiments, the assistant system 140 may assist a user to obtain information or services. The assistant system 140 may enable the user to interact with the assistant system 140 via user inputs of various modalities ( e.g., audio, voice, text, vision, image, video, gesture, motion, activity, location, orientation) in stateful and multi-tum conversations to receive assistance from the assistant system 140. See paragraph [0071]).
Regarding claim 23, Li teaches A method (In one embodiment, a method includes receiving, by a XR
display device, a gesture-based input from a first user of the XR display device, processing, using a gesture-detection model, the gesture-based input to identify a first gesture, receiving, by the XR display device, an audio input from the first user, where the audio input includes a first voice command, processing, using a natural-language model, the audio input to identify one or more intents or one or more slots associated with the first voice command, determining whether the identified first gesture matches the first voice command, and executing, responsive to the identified first
gesture matching the first voice command and by the XR display device, a first task corresponding to the first voice command based on the identified first gesture and the identified one or more intents or one or more slots. See abstract)(FIG. 1 illustrates an example network environment 100 associated with an assistant system. Network environment 100 includes a client system 130, an assistant system 140, a social-networking system 160, and a third-party system 170 connected to each other by a network 110. See paragraph [0051])( FIG. 2 illustrates an example architecture 200 of the assistant system 140. In particular embodiments, the assistant system 140 may assist a user to obtain information or services. The assistant system 140 may enable the user to interact with the assistant system 140 via user inputs of various modalities ( e.g., audio, voice, text, vision, image, video, gesture, motion, activity, location, orientation) in stateful and multi-tum conversations to receive assistance from the assistant system 140. See paragraph [0071])comprising:
receiving a 3D digital model representing a physical environment (Augmented reality can be defined as a system that incorporates three basic features: a combination of real and virtual worlds, real-time interaction, and accurate 3D registration of virtual and real objects. The overlaid sensory information can be constructive (i.e. additive to the natural environment), or destructive (i.e. masking of the natural environment). This experience is seamlessly interwoven with the physical world such that it is perceived as an immersive aspect of the real environment. See paragraph [0005]);
receiving a first user input, the first user input including an first verbal input to control navigation to a destination within the 3D model (In particular embodiments, a client system may implement voice commands within an augmented reality (AR), virtual reality (VR), mixed reality (MR), or extended reality (XR) environment via a voice SDK, which allows XR applications to easily integrate voice commands on XR devices (e.g., the client system). In particular embodiments, XR may be one or more of a combination of AR, VR, or MR. As an example and not by way of limitation, XR may include elements of AR or VR or MR. In particular embodiments, the client system may implement voice commands in combination with gestures within XR environments via a voice SDK. As XR content becomes more immersive, voice integration may make the content more engaging to interact with various applications. That is, people typically interact with the real world using their voice and hands. Currently, navigating through XR content may be cumbersome as the user may need to navigate through unfamiliar nested menus. As an example and not by way oflimitation, a user may need to click through various virtual menus by ray-casting from a controller or inputting text through a virtual keyboard to navigate to a desired destination. To improve upon the user experience, the XR system may implement voice commands, combined with one or more other modalities ( e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0009])( The solution presented by the embodiments disclosed herein to address this challenge may be the XR system may implement voice commands, combined with one or more other modalities ( e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/ search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0011])( As an example and not by way of limitation, the XR display device may render a navigation menu to access information and/or other parts of the application associated with the XR environment the user is located. In particular embodiments, the XR display device may render, for one or more displays of the XR display device, information corresponding to frequently asked questions responsive to executing the first task. Although this disclosure describes rendering one or more elements in a particular manner, this disclosure contemplates rendering one or more elements in any suitable manner. See paragraph [0172]);
translating the first verbal input into a first text query using a first machine learning model (In particular embodiments, the user input may comprise speech input. The speech input may be received at the ASR module 208 for extracting the text transcription from the speech input. The ASR module 208 may use statistical models to determine the most likely sequences of words that correspond to a given portion of speech received by the assistant system 140 as audio input. The models may include one or more of hidden Markov models, neural networks, deep learning models, or any combination thereof. The received audio input may be encoded into digital data at a particular sampling rate ( e.g., 16, 44.1, or 96 kHz) and with a particular number of bits representing each sample ( e.g., 8, 16, of 24 bits). See paragraph [0134]);
analyzing, by a second machine learning model, the first text query to determine a desired navigation (the text transcription from the ASR module 208 may be sent to the NLU module 210. The NLU module 210 may process the text transcription and extract the user intention (i.e., intents) and parse the slots or parsing result based on the linguistic ontology. In particular embodiments, the intents and slots from the NLU module 210 and/or the events and contexts from the context engine 220 may be sent to the entity resolution module 212. See paragraph [0137])( As an example and not by way of limitation, the user input may comprise "direct me to my next meeting." The assistant system 140 may use a calendar agent to retrieve the location of the next meeting. The assistant system 140 may then use a navigation agent to direct the user to the next meeting. See paragraph [0126])( FIG. 7 illustrates an example flow diagram of processing an audio input. In particular embodiments, a client system 702 ( e.g., client system 130) may be embodied as a XR display device, such as an AR headset or a VR headset. The NLU model may match the text to intent and extract entities for the slots. In particular embodiments, the NLU model 714 may send the results back to the application 706. In particular embodiments, the NLU model 714 may send the intents, slots, and entities to a dialog manager 716. The dialog manager 716 may resolve the intents and slots received from the NLU model. In particular embodiments, the assistant system 704 may perform a task or send instructions to execute a task to the client system 702. The dialog manager 716 may send instructions to the application 706 to perform a task based on the received intents, slots, and entities from the NLU model. In particular embodiments, the assistant system 704 may generate a result from processing the intents, slots, and entities from the NLU model 714 and render a response through the text-to-speech module 718. See paragraph [0176] ); and
providing one or more navigational commands to control navigation to a destination within a graphical user interface associated with the 3D model responsive to the verbal ( As an example and not by way of limitation, the user input may comprise "direct me to my next meeting." The assistant system 140 may use a calendar agent to retrieve the location of the next meeting. The assistant system 140 may then use a navigation agent to direct the user to the next meeting. See paragraph [0126])( A XR system may use a language model to identify phrases from an audio input and compare that to a list of phrases associated with the application. The XR system may identify a corresponding gesture performed and compare that to a list associated with a combination of phrases and gestures. When a particular phrase and gesture combination is detected, then the XR system may present a corresponding output, such as a resulting wizard spell. The XR system may perform object recognition on real-world objects and allow the user to interact with the object (e.g., see a real-world menu, CV recognizes it, then point at an item on the menu and ask the assistant to tell you about it). See paragraph [0154])( the XR system may implement voice commands, combined with one or more other modalities (e.g., gesture, pose, eye gaze, etc.) to allow a user to perform voice navigation/search, voice FAQ, and voice-driven gameplay & experiences. See paragraph [0152]).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 2, 3, 5, 6, 7, 8, 13, 14, 16, 18, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Li et al. (US 2023/0128422)(Hereinafter referred to as Li) in view of Moore et al. (US 2020/0333157)(Hereinafter referred to as Moore).
Regarding claim 2, Li teaches The system of claim 1, but is silent to wherein the memory containing instructions to further control the one or more processors to: generate a first verbal response describing the destination based on metadata associated with the destination and the 3D model; and provide the first verbal response.
However, Li further teaches wherein the memory containing instructions to further control the one or more processors to: generate a first verbal response based on data associated with the 3D model and provide the first verbal response (FIGS. 6A-6B illustrate an example flow diagram of processing an audio input. the process 600 may start at step 602, where a client system receives an utterance from a user. At step 604, the client system may use ASR and a natural language model to process the utterance. The client system may process the utterance to generate one or more intents and one or more entities . the client system may present an audio output to the user to respond to the user utterance. See paragraph [0175])(The first task may also comprise presenting, by the XR display device, the translated text to the first user. As an example and not by way of limitation, the XR display device may present the translated text through rendering the text to be displayed on one or more displays of the XR display device or through outputting an audio output through speakers of the XR display device. See paragraph [0174]).
Moore teaches to generate a first verbal response describing the destination based on metadata associated with the destination and the 3D model; and provide the first verbal response (the navigation application includes an interactive map application ( or interactive navigation application) that provides information to the user in an audible form (e.g., as natural language speech), and receives user inputs and commands in a verbal form (e.g., as natural language speech). See paragraph [0382])(The speech input may be a verbal request for directions, a verbal request to perform a local search (e.g., a search for restaurants, gas stations, lodging, etc.) [metadata], or simply a verbal command to activate the interactive map. In response to the speech input, the device makes at least a subset of functionalities ( e.g., providing directions and performing searches) of the interactive map available to the user through an audio-output-and-speech-input user interface; See paragraph [0385])(in response to the speech input, visual information (e.g., a list of search results) is provided to the user on the lock-screen along with an audio output (e.g., a reading of the information to the user) [describing the destination]; Para 0405, the user is allowed to request point to point directions from the interactive map via a natural language speech query, such as "How do I get from Time Square to the Empire State building?" The interactive map responds to the user's inquiry by providing point to point directions to the user, for example, either visually and/or audibly. As the user travels from one location to the next location, the interactive map optionally (e.g., upon user's verbal request) provides information to the user in an audible form, such as time to destination, distance to destination, and current location) See paragraph [0386]).
Li and Moore both teach of navigations systems and Moore teaches that by freeing visual and physical attention the system allows the user to engage in other activities, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Li with the navigation techniques of Moore such that the system could fee up the user’s attention to allow them to perform other activities simultaneously.
Regarding claim 3, Li in view of Moore teaches the system of claim 2, wherein the first verbal response includes a position of the destination relative to other locations within the physical environment ((Moore; FIG. 51 illustrates user device 4900 of FIG. 49 after the user makes an inquiry based on the current dialog. The figure is shown in three stages 5101-5103. In stage 5101, the user asks (as shown by arrow 5105) "when will I get there". In stage 5102, the screen optionally displays the audible interaction between the user and the voice-activated service. In stage 5103, since the current dialog was about the next turn in the route, the voice-activated navigation responds (as
shown by arrow 5110) "in two minutes". The voice-activated navigation makes the response based on the current position, the distance to the next waypoint [position of destination], the current speed, the traffic condition between the current position and the next waypoint, etc.) See paragraph [0419]).
Regarding claim 5, Li in view of Moore teaches the system of claim 2, wherein the first verbal input is received via a microphone (Li; the XR display device may receive an audio input comprising a first voice command by one or more microphones of the XR display device. See paragraph [0195]) and the first verbal response is generated to be provided by an audio speaker (Li; As an example and not by way of limitation, the XR display device may present the translated text through rendering the text to be displayed on one or more displays of the XR display device or through outputting an audio output through speakers of the XR display device. See paragraph [0175]).
Regarding claim 6, Li in view of Moore teaches the system of claim 2, wherein the first verbal input is received via a microphone (Li; the XR display device may receive an audio input comprising a first voice command by one or more microphones of the XR display device. See paragraph [0195]), but is silent to and the first verbal response is to be provided as text in the current combination.
Moore further teaches providing text based on a verbal response (conjunctively with, the voice guidance, the navigation application can provide text and/or graphical instructions in at least two modes while operating in the background. See paragraph [0023]) (FIG. 60 illustrates 4 stages 6001-6004 of a user interface of some embodiments where navigation is incorporated into voiceactivated service output. As shown, a map 6025 and navigational directions 6090 are shown on the screen. The map identifies the current location 6030 of the user device and a route 6035 that is currently set for navigation. In this example, the navigation application provides a verbal guidance when the device reaches within 50 feet of the next turn. As shown in stage 6001, the user device is still 60 feet from the next turn (as shown by 6090)). See paragraph [0497])
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the system of Li in view of Moore with the text presentation techniques of Moore such that users with audio impairment could still utilize the system properly.
Regarding claim 7, Li in view of Moore teaches The system of claim 2, but is silent to wherein the memory containing instructions to further control the one or more processors to: generate a textual response based on some or all of the first verbal response; and provide to the graphical user interface the textual response.
Moore further teaches providing text based on a verbal response (conjunctively with, the voice guidance, the navigation application can provide text and/or graphical instructions in at least two modes while operating in the background. See paragraph [0023]) (FIG. 60 illustrates 4 stages 6001-6004 of a user interface of some embodiments where navigation is incorporated into voiceactivated service output. As shown, a map 6025 and navigational directions 6090 are shown on the screen. The map identifies the current location 6030 of the user device and a route 6035 that is currently set for navigation. In this example, the navigation application provides a verbal guidance when the device reaches within 50 feet of the next turn. As shown in stage 6001, the user device is still 60 feet from the next turn (as shown by 6090)). See paragraph [0497])
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the system of Li in view of Moore with the text presentation techniques of Moore such that users with audio impairment could still utilize the system properly.
Regarding claim 8, Li in view of Moore teaches the system of claim 1, but is silent to wherein the memory containing instructions to further control the one or more processors to: generate a textual response based on some or all of the first verbal input response; and provide to the graphical user interface the textual response.
Moore further teaches providing text based on a verbal response (conjunctively with, the voice guidance, the navigation application can provide text and/or graphical instructions in at least two modes while operating in the background. See paragraph [0023]) (FIG. 60 illustrates 4 stages 6001-6004 of a user interface of some embodiments where navigation is incorporated into voiceactivated service output. As shown, a map 6025 and navigational directions 6090 are shown on the screen. The map identifies the current location 6030 of the user device and a route 6035 that is currently set for navigation. In this example, the navigation application provides a verbal guidance when the device reaches within 50 feet of the next turn. As shown in stage 6001, the user device is still 60 feet from the next turn (as shown by 6090)). See paragraph [0497])
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the system of Li in view of Moore with the text presentation techniques of Moore such that users with audio impairment could still utilize the system properly.
Regarding claim 13, Li teaches the non-transitory computer-readable medium of claim 12, but is silent to wherein the method further comprises: generating a first verbal response describing the destination based on metadata associated with the destination and the 3D model; and providing the first verbal response.
However, Li further teaches wherein the memory containing instructions to further control the one or more processors to: generate a first verbal response based on data associated with the 3D model and provide the first verbal response (FIGS. 6A-6B illustrate an example flow diagram of processing an audio input. the process 600 may start at step 602, where a client system receives an utterance from a user. At step 604, the client system may use ASR and a natural language model to process the utterance. The client system may process the utterance to generate one or more intents and one or more entities . the client system may present an audio output to the user to respond to the user utterance. See paragraph [0175])(The first task may also comprise presenting, by the XR display device, the translated text to the first user. As an example and not by way of limitation, the XR display device may present the translated text through rendering the text to be displayed on one or more displays of the XR display device or through outputting an audio output through speakers of the XR display device. See paragraph [0174]).
Moore teaches to generate a first verbal response describing the destination based on metadata associated with the destination and the 3D model; and provide the first verbal response (the navigation application includes an interactive map application ( or interactive navigation application) that provides information to the user in an audible form (e.g., as natural language speech), and receives user inputs and commands in a verbal form (e.g., as natural language speech). See paragraph [0382])(The speech input may be a verbal request for directions, a verbal request to perform a local search (e.g., a search for restaurants, gas stations, lodging, etc.) [metadata], or simply a verbal command to activate the interactive map. In response to the speech input, the device makes at least a subset of functionalities ( e.g., providing directions and performing searches) of the interactive map available to the user through an audio-output-and-speech-input user interface; See paragraph [0385])(in response to the speech input, visual information (e.g., a list of search results) is provided to the user on the lock-screen along with an audio output (e.g., a reading of the information to the user) [describing the destination]; Para 0405, the user is allowed to request point to point directions from the interactive map via a natural language speech query, such as "How do I get from Time Square to the Empire State building?" The interactive map responds to the user's inquiry by providing point to point directions to the user, for example, either visually and/or audibly. As the user travels from one location to the next location, the interactive map optionally (e.g., upon user's verbal request) provides information to the user in an audible form, such as time to destination, distance to destination, and current location) See paragraph [0386]).
Li and Moore both teach of navigations systems and Moore teaches that by freeing visual and physical attention the system allows the user to engage in other activities, therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine the system of Li with the navigation techniques of Moore such that the system could fee up the user’s attention to allow them to perform other activities simultaneously.
Regarding claim 14, Li in view of Moore teaches The non-transitory computer-readable medium of claim 13, wherein the first verbal response includes a position of the destination relative to other locations within the physical environment ((Moore; FIG. 51 illustrates user device 4900 of FIG. 49 after the user makes an inquiry based on the current dialog. The figure is shown in three stages 5101-5103. In stage 5101, the user asks (as shown by arrow 5105) "when will I get there". In stage 5102, the screen optionally displays the audible interaction between the user and the voice-activated service. In stage 5103, since the current dialog was about the next turn in the route, the voice-activated navigation responds (as shown by arrow 5110) "in two minutes". The voice-activated navigation makes the response based on the current position, the distance to the next waypoint [position of destination], the current speed, the traffic condition between the current position and the next waypoint, etc.) See paragraph [0419]).
Regarding claim 16, Li in view of Moore teaches the non-transitory computer-readable medium of claim 13, wherein the first verbal input is received via a microphone (Li; the XR display device may receive an audio input comprising a first voice command by one or more microphones of the XR display device. See paragraph [0195]) and the first verbal response is generated to be provided by an audio speaker (Li; As an example and not by way of limitation, the XR display device may present the translated text through rendering the text to be displayed on one or more displays of the XR display device or through outputting an audio output through speakers of the XR display device. See paragraph [0175]).
Regarding claim 17, Li in view of Moore teaches the non-transitory computer-readable medium of claim 13, wherein the first verbal input is received via a microphone (Li; the XR display device may receive an audio input comprising a first voice command by one or more microphones of the XR display device. See paragraph [0195]), but is silent to and the first verbal response is to be provided as text in the current combination.
Moore further teaches providing text based on a verbal response (conjunctively with, the voice guidance, the navigation application can provide text and/or graphical instructions in at least two modes while operating in the background. See paragraph [0023]) (FIG. 60 illustrates 4 stages 6001-6004 of a user interface of some embodiments where navigation is incorporated into voiceactivated service output. As shown, a map 6025 and navigational directions 6090 are shown on the screen. The map identifies the current location 6030 of the user device and a route 6035 that is currently set for navigation. In this example, the navigation application provides a verbal guidance when the device reaches within 50 feet of the next turn. As shown in stage 6001, the user device is still 60 feet from the next turn (as shown by 6090)). See paragraph [0497])
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the system of Li in view of Moore with the text presentation techniques of Moore such that users with audio impairment could still utilize the system properly.
Regarding claim 18, Li teaches the non-transitory computer-readable medium of claim 12, but is silent to wherein the method further comprises: generating a textual response based on some or all of the first verbal response; and providing to the graphical user interface the textual response.
Moore further teaches providing text based on a verbal response (conjunctively with, the voice guidance, the navigation application can provide text and/or graphical instructions in at least two modes while operating in the background. See paragraph [0023]) (FIG. 60 illustrates 4 stages 6001-6004 of a user interface of some embodiments where navigation is incorporated into voiceactivated service output. As shown, a map 6025 and navigational directions 6090 are shown on the screen. The map identifies the current location 6030 of the user device and a route 6035 that is currently set for navigation. In this example, the navigation application provides a verbal guidance when the device reaches within 50 feet of the next turn. As shown in stage 6001, the user device is still 60 feet from the next turn (as shown by 6090)). See paragraph [0497])
Li and Moore both teach of navigations systems and Moore teaches that by freeing visual and physical attention the system allows the user to engage in other activities and presenting information in a textual manner, therefore, It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the system of Li in view of Moore with the text presentation techniques of Moore such that users with audio impairment could still utilize the system properly.
Regarding claim 19, Li teaches the non-transitory computer-readable medium of claim 12, but is silent to wherein the method further comprises: generating a textual response based on some or all of the first verbal input response; and providing to the graphical user interface the textual response.
Moore further teaches providing text based on a verbal response (conjunctively with, the voice guidance, the navigation application can provide text and/or graphical instructions in at least two modes while operating in the background. See paragraph [0023]) (FIG. 60 illustrates 4 stages 6001-6004 of a user interface of some embodiments where navigation is incorporated into voiceactivated service output. As shown, a map 6025 and navigational directions 6090 are shown on the screen. The map identifies the current location 6030 of the user device and a route 6035 that is currently set for navigation. In this example, the navigation application provides a verbal guidance when the device reaches within 50 feet of the next turn. As shown in stage 6001, the user device is still 60 feet from the next turn (as shown by 6090)). See paragraph [0497])
Li and Moore both teach of navigations systems and Moore teaches that by freeing visual and physical attention the system allows the user to engage in other activities and presenting information in a textual manner, therefore, It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to modify the system of Li in view of Moore with the text presentation techniques of Moore such that users with audio impairment could still utilize the system properly.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Mowatt (US 2018/0068474)(Hereinafter referred to as Mowatt), generally teaches Augmented reality customiztaion.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICHOLAS R WILSON whose telephone number is (571)272-0936. The examiner can normally be reached M-F 7:30-5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (572)-272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NICHOLAS R WILSON/Primary Examiner, Art Unit 2611