DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 9 is objected to because of the following informalities: ‘to’ after onsite should be removed (“based on conversation data related to the on-site ). Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-2, 10-11, 14, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Bhowmik et al. (hereinafter Bhowmik) (US 11264029 B2) (see attached copy for paragraph numbers) in view of Schmirler et al. (hereinafter Schmirler) (US 20180131907 A1).
Regarding claim 1, Bhowmik discloses:
the method comprising:
receiving a voice input of the user (Bhowmik, P(2): "the system receives a first voice input communicating a first query content" (system receives user's spoken query as claimed voice input));
pre-processing a digital signal corresponding to the voice input (Bhowmik, P(2): "speech recognition circuitry programmed to recognize speech within the first voice input to determine the first query content" (speech recognition processing of the electronically captured voice necessarily processes the digital signal corresponding to that voice input to determine its content));
generating a response signal based on a result of processing the pre-processed digital signal (Bhowmik, P(2): "The system processes the first voice input to determine the first query content, and determines whether the first query content matches one of the locally-handled user inputs. If the first audio input matches one of the locally-handled user inputs, then the system takes a local responsive action" (Bhowmik generates a responsive result by applying its local AI assistant and stored input database to the processed voice query result)) and
generating an output in response to the voice input based on the response signal and providing the same to the user (Bhowmik, P(42): "If the first audio input matches one of the locally-handled user inputs, then the system takes a local responsive action. One example of a local responsive action is to provide locally-available information to the user by playing an audio response on the ear-wearable device." (the system generates an audio output based on the responsive action produced from the voice query and plays that output to the user)).
Bhowmik does not explicitly disclose:
A method for providing an artificial intelligence-based voice conversation environment capable of enabling interaction between an artificial intelligence secretary and a user who performs work on site
using an artificial intelligence model pre-trained based on an on-site data set related to an on-site;
However, Schmirler discloses:
A method for providing an artificial intelligence-based voice conversation environment capable of enabling interaction between an artificial intelligence secretary and a user who performs work on site (Schmirler, P[0050]: "one or more embodiments of the present disclosure provide a system that generates and delivers augmented reality (AR) or virtual reality (VR) presentations (referred to collectively herein as “VR/AR presentations”) to a user via a wearable computer or other client device.", "P[0052]: "For users that are physically located on the plant floor, the VR/AR presentation system can provide automation system data, notifications, and proactive guidance to the user" (wearable computer provides interactive electronic assistance and proactive guidance to a user performing work at an industrial site, reading on claimed AI assistant environment for on-site worker)),
using an artificial intelligence model pre-trained based on an on-site data set related to an on-site (Schmirler, P[0085]: "presentation system 302 collects device data 704 from industrial devices or systems 702 across the plant environment.", P[0086]: "Rendering component 308 populates this virtual reality presentation with selected subsets of collected plant data" (teaches constructing responsive computerized worker-assistance output using a stored data set collected from the particular industrial site));
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Bhowmik and Schmirler. Doing so would have brought the collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably improved the relevance and usefulness of the granted responses through combining known industrial information management techniques to a known wearable voice assistant architecture to achieve the predictable result of providing context-specific assistance to workers.
Regarding claim 2, the combination of Bhowmik and Schmirler discloses:
The method of claim 1, wherein the generating a response signal comprises:
generating first text data corresponding to the pre-processed digital signal using a voice input conversion artificial intelligence model (Bhowmik, P(2): "speech recognition circuitry programmed to recognize speech within the first voice input to determine the first query content" (The speech recognition circuitry analyzes the processed voice signal and produces recognized query content corresponding to the claimed first text data and speech recognition was conventionally implemented using trained speech recognition models));
extracting a request word related to the on-site from the first text data (Bhowmik, P(2): "The system processes the first voice input to determine the first query content, and determines whether the first query content matches one of the locally-handled user inputs" (determining the query content and matching that content against stored locally handled inputs reads on extracting the request represented by one or more words withing the recognized text.)); and
generating second text data for an answer word corresponding to the request word by analyzing the request word based on the pre-trained artificial intelligence model (Bhowmik, P(42): "If the first audio input matches one of the locally-handled user inputs, then the system takes a local responsive action. One example of a local responsive action is to provide locally-available information to the user by playing an audio response on the ear-wearable device. " (generates responsive answer content corresponding to the recognized request using stored local assistant data)) based on conversation data related to the on-site (Schmirler, P[0085]: "presentation system 302 collects device data 704 from industrial devices or systems 702 across the plant environment.", P[0086]: "Rendering component 308 populates this virtual reality presentation with selected subsets of collected plant data" (teaches site specific information from which an answer concerning the industrial site is generated)); and
the generating an output and providing the same to the user comprises:
generating the output in a voice format by analyzing the second text data based on a text input conversion artificial intelligence model and providing the same to the user (Bhowmik, P(2): "speech generation circuitry programmed to generate speech output to the first speaker of the ear-wearable device" (the response content is processed by speech-generation circuitry to produce a voice format output that is played to the user through the wearable speaker)).
Regarding claim 10, Bhowmik discloses:
A smart electronic device for providing an artificial intelligence-based voice conversation environment (Bhowmik, P(9): "a local assistant system for responding to voice input includes an ear-wearable device"), the device comprising:
a sensor system for receiving a voice input of a user (Bhowmik, P(9): "The ear-wearable device can include a first speaker, a first microphone, a first processor, a first memory storage, and a first wireless communication device." (microphone captures user's voice input));
an output device for providing an output for the voice input to the user (Bhowmik, P(2): "speech generation circuitry programmed to generate speech output to the first speaker of the ear-wearable device" (the speaker constitutes the output device and provides speech responsive to the received input));
a memory for storing at least one instruction (Bhowmik, P(9): "The ear-wearable device can include a first speaker, a first microphone, a first processor, a first memory storage"); and
a processor assembly for executing the at least one instruction (Bhowmik, P(9): "The ear-wearable device can include a first speaker, a first microphone, a first processor, a first memory storage" (processor and programmed circuitry execute the stored instructions that implement the claimed voice-assistant operations)), wherein the processor assembly, by executing the at least one instruction, is configured to:
pre-process a digital signal corresponding to the voice input of the user received from the sensor system (Bhowmik, P(2): "speech recognition circuitry programmed to recognize speech within the first voice input to determine the first query content" (the processor executes speech recognition processing on the microphone captured digital voice signal to determine its content));
generate a response signal based on a result of processing the pre-processed digital signal (Bhowmik, P(2): "The system processes the first voice input to determine the first query content, and determines whether the first query content matches one of the locally-handled user inputs. If the first audio input matches one of the locally-handled user inputs, then the system takes a local responsive action" (Bhowmik's programmed local assistant generates the responsive result from the processed voice input)) and
generate the output in response to the voice input based on the response signal and provide the same to the user through the output device (Bhowmik, P(42): "If the first audio input matches one of the locally-handled user inputs, then the system takes a local responsive action. One example of a local responsive action is to provide locally-available information to the user by playing an audio response on the ear-wearable device." (the system generates an audio output based on the responsive action produced from the voice query and plays that output to the user through device's speaker)).
Bhowmik does not explicitly disclose:
a display device for providing an augmented reality view to the user
using an artificial intelligence model pre-trained based on an on-site data set related to an on-site
However, Schmirler discloses:
a display device for providing an augmented reality view to the user (Schmirler, P[0050]: "a system that generates and delivers augmented reality (AR) or virtual reality (VR) presentations (referred to collectively herein as “VR/AR presentations”) to a user via a wearable computer", P[0051]: "can enhance or augment this live view with superimposed operational or status data" (the wearable computer's display provides an augmented view containing virtual information superimposed over the user's environmental view));
using an artificial intelligence model pre-trained based on an on-site data set related to an on-site (Schmirler, P[0085]: "presentation system 302 collects device data 704 from industrial devices or systems 702 across the plant environment.", P[0086]: "Rendering component 308 populates this virtual reality presentation with selected subsets of collected plant data" (teaches constructing responsive computerized worker-assistance output using a stored data set collected from the particular industrial site));
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Bhowmik and Schmirler. Doing so would have brought the collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably improved the relevance and usefulness of the granted responses through combining known industrial information management techniques to a known wearable voice assistant architecture to achieve the predictable result of providing context-specific assistance to workers.
Regarding claim 11, claim 11 recites the device corresponding to the method described in claim 2 and is rejected under the same grounds as above.
Regarding claim 14, Bhowmik discloses:
A system for providing an artificial intelligence-based voice conversation environment, the system comprising (Bhowmik, Abstract: "Embodiments herein relate to a local assistant system responding to voice input using an ear-wearable device."):
a computing device comprising a processor for performing calculation to provide an environment capable of operating an artificial intelligence-based voice conversation environment application (Bhowmik, P(3): "the local assistant responds to voice input using the ear-wearable device and a gateway device, wherein the gateway device includes a gateway processor, a gateway memory storage" (the gateway is a processor-based computing device that executes the local voice-assistant functions and provides the operating environment for the application));
at least one smart electronic device for providing the artificial intelligence-based voice conversation environment to a user (Bhowmik, P(9): "a local assistant system for responding to voice input includes an ear-wearable device" (the processor-based ear wearable is the smart electronic device through which the user accesses the voice-assistant environment));
an input device for receiving a voice input of the user (Bhowmik, P(9): "The ear-wearable device can include a first speaker, a first microphone, a first processor, a first memory storage, and a first wireless communication device"); and
an output device for generating an output in response to the voice input and providing the same to the user (Bhowmik, P(2): "speech generation circuitry programmed to generate speech output to the first speaker of the ear-wearable device"), wherein the processor is configured to:
pre-process a digital signal corresponding to the voice input of the user (Bhowmik, P(2): "speech recognition circuitry programmed to recognize speech within the first voice input to determine the first query content" (the processor executes speech recognition processing on the microphone captured digital voice signal to determine its content));
generate a response signal based on a result of processing the pre-processed digital signal (Bhowmik, P(2): "The system processes the first voice input to determine the first query content, and determines whether the first query content matches one of the locally-handled user inputs. If the first audio input matches one of the locally-handled user inputs, then the system takes a local responsive action" (Bhowmik's programmed local assistant generates the responsive result from the processed voice input)) and
generate an output in response to the voice input based on the response signal and provide the same to the user (Bhowmik, P(42): "If the first audio input matches one of the locally-handled user inputs, then the system takes a local responsive action. One example of a local responsive action is to provide locally-available information to the user by playing an audio response on the ear-wearable device." (the system generates an audio output based on the responsive action produced from the voice query and plays that output to the user)).
Bhowmik does not explicitly disclose:
using an artificial intelligence model pre-trained based on an on-site data set related to an on-site
However, Schmirler discloses:
using an artificial intelligence model pre-trained based on an on-site data set related to an on-site (Schmirler, P[0085]: "presentation system 302 collects device data 704 from industrial devices or systems 702 across the plant environment.", P[0086]: "Rendering component 308 populates this virtual reality presentation with selected subsets of collected plant data" (teaches constructing responsive computerized worker-assistance output using a stored data set collected from the particular industrial site));
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Bhowmik and Schmirler. Doing so would have brought the collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably improved the relevance and usefulness of the granted responses through combining known industrial information management techniques to a known wearable voice assistant architecture to achieve the predictable result of providing context-specific assistance to workers.
Regarding claim 16, claim 16 recites the system corresponding to the methods described in claim 2 and is rejected under the same grounds as above.
Claims 3-4 are rejected under 35 U.S.C. 103 as being unpatentable over Bhowmik et al. (hereinafter Bhowmik) (US 11264029 B2) (see attached copy for paragraph numbers) in view of Schmirler et al. (hereinafter Schmirler) (US 20180131907 A1) in further view of Yoon et al. (hereinafter Yoon) (US 20210341991 A1).
Regarding claim 3, the combination of Bhowmik and Schmirler discloses the method of claim 1.
The combination of Bhowmik and Schmirler does not explicitly disclose:
further comprising collecting data of on-site information related to a surrounding environment of the user;
However, Yoon discloses:
further comprising collecting data of on-site information related to a surrounding environment of the user (Yoon, P[0023]: "Accordingly, non-limiting examples of data captured by the AR device 102 may include location data (e.g., GPS geolocation data), biometric data (e.g., heartrate data), motion/acceleration data, image/video data, environmental data (e.g., heat data, light data, moisture data, etc.), posture data (e.g., information indicating wherein a user's body a sensor/AR device 102 is located), other data, and any combination thereof." (wearable AR device collects multiple forms of information concerning the environment surrounding the on-site user)).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Yoon with Bhowmik and Schmirler. Doing so would have brought the wearable augmented reality devices capable of collecting real time environmental information of Yoon (Yoon, Abstract, P[0023]) collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably used known sensor-based context gathering for the use of enhancing an existing wearable industrial assistant.
Regarding claim 4, The combination of combination of Bhowmik, Schmirler, and Yoon discloses the method of claim 3.
The combination further discloses:
further comprising displaying the on-site information into field of view of the user (Schmirler, P[0070]: "Rendering component 308 can generate the presentation such that items of the plant data 610 are overlaid on or near graphical representations of the industrial assets to which the items of data relate." (plant or on-site information is displayed as an AR overlay within the view rendered by the wearable appliance. Its wearable presentations encompass the wearer's field of view.)).
Claims 5, 6, 12, 13, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Bhowmik et al. (hereinafter Bhowmik) (US 11264029 B2) (see attached copy for paragraph numbers) in view of Schmirler et al. (hereinafter Schmirler) (US 20180131907 A1) in further view of Yoon et al. (hereinafter Yoon) (US 20210341991 A1) and Montgomerie et al. (hereinafter Montgomerie) (US 20160292925 A1).
Regarding claim 5, The combination of Bhowmik, Schmirler, and Yoon discloses the method of claim 3.
The combination of Bhowmik, Schmirler, and Yoon does not explicitly disclose:
wherein the generating a response signal comprises generating the response signal based on the data of the on-site information and the result of processing the pre-processed digital signal
However, Montgomerie discloses:
wherein the generating a response signal comprises generating the response signal based on the data of the on-site information and the result of processing the pre-processed digital signal (Montgomerie, P[0168]: "voice commands may be used to navigate the entire process, as well as to control content." (content provided by AR system is generated or selected in response to processed voice command information), P[0066]: "a method for interaction using augmented reality may include capturing a video frame (1800), generating AR coordinates from captured video frame (1805), updating a scene view according to AR coordinates (1810), creating new content by adding 3D models or annotating the rendered view (1815)" (The resulting AR content is also based on captured on-stie image information, thus teaching generation based on both the processed voice command and surrounding site data)).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Montgomerie with Yoon, Bhowmik and Schmirler. Doing so would have brought the voice-controlled augmented reality systems utilizing captured images of user’s surroundings and computer vision techniques of Montgomerie (Montgomerie, Abstract, P[0168]) wearable augmented reality devices capable of collecting real time environmental information of Yoon (Yoon, Abstract, P[0023]) collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably enhanced the usefulness of the wearable assistant while employing only know techniques for integrating voice interaction, computer vision, and wearable augmented reality systems.
Regarding claim 6, The combination of Bhowmik, Schmirler, Yoon, and Montgomerie discloses the method of claim 5.
The combination further discloses:
wherein the data of the on-site information comprises data of an on-site image acquired by capturing the surrounding environment of the user (Montgomerie, P[0075]: "a “Local user/Field service” view 400 may include a camera 410, an AR toolkit 420, a renderer 430, a native plugin 440, and a video encoder 450." (the field user's wearable or mobile camera captures image/video data depicting user's surrounding work environment)), wherein the generating a response signal comprises generating the response signal based on the result of processing the pre-processed digital signal and a result of analyzing the acquired on-site image using the pre-trained artificial intelligence model based on image data related to the on-site (Montgomerie, P[0168]: "voice commands may be used to navigate the entire process, as well as to control content.", P[0083]: "computer vision techniques that may be able analyze the image to detect real-time state changes" (Uses both processed voice commands and computer-vision analysis of a work-site image to determine or control the responsive AR content provided to the worker. The computer-vision model is necessarily trained using representative image data for the objects or states being detected.)).
Regarding claim 12, The combination of Bhowmik and Schmirler discloses the device of claim 10.
The combination of Bhowmik and Schmirler does not explicitly disclose:
wherein the processor assembly is configured to:
generate the response data based on data of on-site information related to the on-site collected through the sensor system and the result of processing the pre-processed digital signal
However, Yoon discloses:
wherein the processor assembly is configured to:
generate the response data based on data of on-site information related to the on-site collected through the sensor system (Yoon, P[0023]: "non-limiting examples of data captured by the AR device 102 may include location data (e.g., GPS geolocation data), biometric data (e.g., heartrate data), motion/acceleration data, image/video data, environmental data (e.g., heat data, light data, moisture data, etc.), posture data (e.g., information indicating wherein a user's body a sensor/AR device 102 is located), other data, and any combination thereof." (teaches collecting on site information related to the surrounding environment through the AR device sensor system))
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Yoon with Bhowmik and Schmirler. Doing so would have brought the wearable augmented reality devices capable of collecting real time environmental information of Yoon (Yoon, Abstract, P[0023]) collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably used known sensor-based context gathering for the use of enhancing an existing wearable industrial assistant.
Yoon in combination with Bhowmik and Schmirler does not explicitly disclose:
and the result of processing the pre-processed digital signal However, Montgomerie discloses:
and the result of processing the pre-processed digital signal (Montgomerie, P[0168]: "voice commands may be used to navigate the entire process, as well as to control content" (teaches generating the responsive content based on the processed voice input (ex. the processed digital signal))).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Montgomerie with Yoon, Bhowmik and Schmirler. Doing so would have brought the voice-controlled augmented reality systems utilizing captured images of user’s surroundings and computer vision techniques of Montgomerie (Montgomerie, Abstract, P[0168]) wearable augmented reality devices capable of collecting real time environmental information of Yoon (Yoon, Abstract, P[0023]) collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably enhanced the usefulness of the wearable assistant while employing only know techniques for integrating voice interaction, computer vision, and wearable augmented reality systems.
Regarding claim 13, claim 13 recites the system corresponding to the methods described in claim 6 and is rejected under the same grounds as above.
Regarding claim 17, claim 17 recites the system corresponding to the device described in claim 12 and is rejected under the same grounds as above.
Claims 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Bhowmik et al. (hereinafter Bhowmik) (US 11264029 B2) (see attached copy for paragraph numbers) in view of Schmirler et al. (hereinafter Schmirler) (US 20180131907 A1) in further view of Yoon et al. (hereinafter Yoon) (US 20210341991 A1), Montgomerie et al. (hereinafter Montgomerie) (US 20160292925 A1) and Hou et al. (hereinafter Hou) (US 20230351724 A1).
Regarding claim 7, the combination of Bhowmik, Schmirler, Yoon, and Montgomerie discloses the method of claim 6.
The combination further discloses:
generating the response signal based on the result of processing the pre-processed digital signal and a result of analyzing an image of the target object using the pre-trained artificial intelligence model based on the image data related to the on-site (Montgomerie, P[0168]: "voice commands may be used to navigate the entire process, as well as to control content.", P[0083]: "computer vision techniques that may be able analyze the image to detect real-time state changes" (Montgomerie generates or selects the responsive content based jointly on the processed voice command and machine analysis of the work site image and objects or states depicted)).
The combination does not explicitly disclose:
wherein the generating a response signal comprises:
detecting a target object from the on-site image using an object detection algorithm;
However, Hou discloses:
wherein the generating a response signal comprises:
detecting a target object from the on-site image using an object detection algorithm (Hou, Abstract: "Object detection can be performed by a machine-learned model configured to determine various object properties." (Hou directly detects an object depicted in an acquired image using ML object detection algorithm);
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Hou with Montgomery, Yoon, Bhowmik and Schmirler. Doing so would have brought the ML learned object detection and estimation of object pose (via image date) of Hou (Hou, Abstract) with the voice-controlled augmented reality systems utilizing captured images of user’s surroundings and computer vision techniques of Montgomerie (Montgomerie, Abstract, P[0168]) wearable augmented reality devices capable of collecting real time environmental information of Yoon (Yoon, Abstract, P[0023]) collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably improved the precision of augmented reality guidance while preserving the operation of the underlying wearable voice assistant system.
Regarding claim 8, the combination of Bhowmik, Schmirler, Yoon, and Montgomerie discloses the method of claim 6.
The combination further discloses:
and
displaying virtual content corresponding to related information of the target object, the virtual content being matched to the detected target object (Schmirler, P[0086]: "Rendering component 308 can generate the presentation such that items of the plant data 610 are overlaid on or near graphical representations of the industrial assets to which the items of data relate." (The virtual information associated with an identified industrial asset is spatially matched and displayed on or near that target object.)).
The combination does not explicitly disclose:
further comprising:
detecting a target object from the on-site image using an object detection algorithm;
estimating a pose of the target object based on the data of the on-site image and pose information of the user;
extracting information related to the target object by analyzing an image of the target object using the pre-trained artificial intelligence model based on the image data related to the on- site;
However, Hou discloses:
further comprising:
detecting a target object from the on-site image using an object detection algorithm (Hou, Abstract: "Object detection can be performed by a machine-learned model configured to determine various object properties." (the ML model directly detects a target object from input image data));
estimating a pose of the target object based on the data of the on-site image and pose information of the user (Hou, P[0081]: "computing systems according to the present disclosure can perform 3D object detection using 2D imagery that in certain implementations can be combined with additional data such as camera intrinsics and/or AR data." (Estimates the object's 3D pose from the acquired image together with AR/camera pose information associated with user's camera device), P[0018]: "These properties can include an x, y, z-coordinate position, a yaw, pitch, roll rotational orientation (together 6DoF pose)" (The resulting position and rotation orientation constitute the claimed pose of the target object));
extracting information related to the target object by analyzing an image of the target object using the pre-trained artificial intelligence model based on the image data related to the on- site (Hou, P[0025]: "This image data can be provided a machine-learned object detection model which can be trained according to example implementations of the present disclosure. " (The trained ML model analyzes the target object image and extracts object properties, including identity-related segmentation, location, orientation, and bounding box information));
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Hou with Montgomery, Yoon, Bhowmik and Schmirler. Doing so would have brought the ML learned object detection and estimation of object pose (via image date) of Hou (Hou, Abstract) with the voice-controlled augmented reality systems utilizing captured images of user’s surroundings and computer vision techniques of Montgomerie (Montgomerie, Abstract, P[0168]) wearable augmented reality devices capable of collecting real time environmental information of Yoon (Yoon, Abstract, P[0023]) collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably improved the precision of augmented reality guidance while preserving the operation of the underlying wearable voice assistant system.
Claims 9 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Bhowmik et al. (hereinafter Bhowmik) (US 11264029 B2) (see attached copy for paragraph numbers) in view of Schmirler et al. (hereinafter Schmirler) (US 20180131907 A1) in further view of Montgomerie et al. (hereinafter Montgomerie) (US 20160292925 A1).
Regarding claim 9, the combination of Bhowmik and Schmirler disclose the method of claim 1.
The combination further discloses:
wherein the generating a response signal comprises:
generating first text data corresponding to the pre-processed digital signal using a voice input conversion artificial intelligence model (Bhowmik, P(2): "speech recognition circuitry programmed to recognize speech within the first voice input to determine the first query content" (Speech recognition model converts the processed voice input into recognized query content corresponding to first text data));
The combination does not explicitly disclose:
extracting an augmented reality manual content request word related to the on-site from the first text data; and
generating an augmented reality manual content signal corresponding to the request word by analyzing the request word based on the pre-trained artificial intelligence model based on conversation data related to the on-site to; and
the generating the output and providing the same to the user comprises:
generating augmented reality manual content corresponding to the request word based on the augmented reality manual content signal and providing the same into field of view of the user.
However. Montgomerie discloses:
extracting an augmented reality manual content request word related to the on-site from the first text data (Montgomerie, P[0083]: "Moving from step to step may be achieved in a variety of ways—for example, via a simple interface such as buttons; voice control with simple commands such as “next” or “back”, or in the case of a branched workflow via speaking choices listed as commands" (the AR workflow recognizes spoken command words identifying requested manual steps or content within the AR procedure)); and
generating an augmented reality manual content signal corresponding to the request word by analyzing the request word based on the pre-trained artificial intelligence model based on conversation data related to the on-site to (Montgomerie, P[0168]: "voice commands may be used to navigate the entire process, as well as to control content." (The AR system analyzes the recognized spoken command and generates or selects the corresponding AR workflow/manual content), P[0069]: "The Scope SDK may provide software tools to support creation of step-by-step visual instructions in Augmented Reality (AR) and Virtual Reality (VR)" (The content controlled by the voice request is site/task related AR manual content previously generated and stored for the relevant procedure)); and
the generating the output and providing the same to the user comprises:
generating augmented reality manual content corresponding to the request word based on the augmented reality manual content signal and providing the same into field of view of the user (Montgomerie, P[0064]: "The instructor 300 may load content into the field of view of each student 310, and manipulate it in real time in order to provide an immersive learning experience." (The requested AR instructed/manual content is generated or retrieved and displayed directly within the wearable user's field of view)).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Bhowmik in view of Schmirler and Montgomerie. Doing so would have brought the voice-controlled augmented reality systems utilizing captured images of a user’s surroundings and computer vision techniques (Montgomerie, Abstract, P[0168]) with the collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably improved enhanced the usefulness of the wearable assistant while employing only known techniques for integrating voice interaction, computer vision, and wearable augmented reality systems.
Regarding claim 15, the combination of Bhowmik and Schmirler disclose the method of claim 14.
The combination does not explicitly disclose:
a first smart electronic device that a first user uses; and
a second smart electronic device that a second user different from the first user uses, wherein the processor is configured to generate an output in response to a first voice input based on the first voice input of the first user and provide the same to the second user through the second smart electronic device
However, Montgomerie discloses:
wherein the at least one smart electronic device comprises:
a first smart electronic device that a first user uses (Montgomerie, See mapping below); and
a second smart electronic device that a second user different from the first user uses (Montgomerie, P[0063]: "Display devices 330 and 340 used by the instructor 300 and the student 310, respectively, may be, for example, a tablet, mobile phone, laptop computer, head-mounted display, etc. " (provides separate smart electronic devices respectively used by two different users)), wherein the processor is configured to generate an output in response to a first voice input based on the first voice input of the first user and provide the same to the second user through the second smart electronic device (Montgomerie, P[0063]: "Instructor 300 and student 310 may share a live connection 3220 in a network such as the Internet and may also share audio, video, data, etc." (The processor generates or selects responsive content under the first user's control and provides that output to the second user through the second smart device)).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have combined the teachings of Bhowmik in view of Schmirler and Montgomerie. Doing so would have brought the voice-controlled augmented reality systems utilizing captured images of a user’s surroundings and computer vision techniques (Montgomerie, Abstract, P[0168]) with the collecting, organizing, and presenting industrial site specific information through wearable augmented reality systems of Schmirler (Schmirler, Abstract, P[0050]-P[0052], P[0085]) with the wearable voice assistant system that receives spoken input, performs speech recognition, generates responsive information, and outputs synthesized speech of Bhowmik (Bhowmik, Abstract, P(2)). This modification would have predictably improved enhanced the usefulness of the wearable assistant while employing only known techniques for integrating voice interaction, computer vision, and wearable augmented reality systems.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHASHIDHAR S MANOHARAN whose telephone number is (571)272-6772. The examiner can normally be reached M-F 8:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHASHIDHAR SHANKAR MANOHARAN/Examiner, Art Unit 2655
/ANDREW C FLANDERS/Supervisory Patent Examiner, Art Unit 2655