Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This action is in response to amendments and remarks filed on 01/15/2026. Claims 2-3, 10, 12, 17, 20, 26-27, and 30 are amended. Claim 19 is canceled. Claims 1-18, and 20-30 are considered in this office action. Claims 1-18, and 20-30 are pending examination. Claims 1-18, and 20-30 are rejected necessitated by the amendments. This action is Final.
Response to Arguments
Applicant presents the following arguments regarding the previous office action:
Lock does not disclose determining a location of the objects in the area via a mobile device, and Lock does not disclose tactile and/or audible instructions via a mobile device to the user to guide the user and the camera to the determined location of the object.
Lydecker fails to disclose a pre-selected number of location samples are generated, or that a location is determined via averaging predicted locations determined for the pre-selected number of location samples such the determined location is an average predicted location, Lydecker does not suggest or use ray casting in any way especially not for the use determining the location of an object, Lydecker's process is not applied to camera data of any kind, and that the prior art is silent on generating a pre-selected number of location samples via ray casting of the image data.
Lydecker is silent on providing hand guidance to a user, and Hirano is silent on providing guidance to a user to move closure to a selected object and providing guidance to a user so the user can move vertically and horizontally to be closure to an object.
Regarding applicant argument A, that the cited art fails to teach, for claims 1 and 9, determining a location of the objects in the area via a mobile device, tactile and or audible instructions via a mobile device to the user to guide the user and the camera to the determine location of the object, the use of ray casting in for the use of determining the location of an object applied to camera data, and a pre-selected number of location samples are generated, or that a location is determined via averaging predicted locations determined for the pre-selected number of location samples such the determined location is an average predicted location. Lock teaches the claimed hand held mobile guidance context as claimed where the guidance is generated based on the camera's current view and the cameras pose which allows the app to choose the next action. The guidance actions are relative to the current camera pan-tilt orientation such that the "Up" action generates a waypoint above the camera's current orientation. Lock goes on to state that for visually impaired users the waypoint position would be provided via audio or vibrotactile instructions (4.2.3 Smartphone Application, tracking the camera’s pose allows the app to infer the current state and choose the optimal action to take next) … (4.2.3 Smartphone Application, for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions). Thus, Lock teaches a hand held mobile device that provides direction guidance to move the camera vertically or horizontally based on the current camera position.
Regarding the applicants further allegation that Lock does not disclose the location based object guidance in the way that is claimed. Even if deemed convincing, Lydecker teaches mapping data that includes object identification and navigational instructions or indicators (Lydecker, Paragraph 0098, Lines 1-8, The lighting device processor 47 is configured to update the mapping data based on the user input and the location indicated by the user's smartphone 63 or mobile device. Subsequently, when a visually impaired user needs the water bottle, the visually impaired user speaks a query (e.g., “where is my water bottle”). In response to the visually impaired user's query, the lighting device may generate an indication of an object location based on the updated mapping data), and further teaches that when using the updated mapping data objects are identified and the system determines a location of the user device and the determined position of the user data (Lydecker, 18, determining a location of the user device in the premises using visible light communication between the lighting device and the user device; and based on the determined position of the user device, accessing mapping data corresponding to the determined location of the user device in the premises to be provided to the user device). Lydecker further teaches audible instructions via a mobile device to the user to guide the user and the camera to the determine location of the object (0061, examples may include audio examples that use sound to generate a scene in the user's brain. The audio examples may also include ultrasonic systems that facilitate the generation of directed audio).
A person of ordinary skill in the art would have found it obvious to incorporate Lydecker's object-location mapping and audio navigational output into Lock's smartphone based active search system in order to provide location based assistance to a visually impaired user with predictable results. Therefore, applicants arguments with respect to the independent claims 1 and 9 are not persuasive.
Regarding applicants argument B, the applicants arguments have been fully considered and are mute in light of the new grounds for rejection below.
Regarding applicant’s argument C the applicants arguments have been fully considered and are not persuasive. The claim does not require that the object to be only limited to independent real world articles discovered by the camera. Hirano teaches detecting and specifying the position of the operator’s hand from a 3D camera to obtain the position of the operators hand, and that the hand is guided for a target objective that is expressly described. Hirano further teaches that the notification of the operation target may be a sound or a display object on the HUD where the operators hand is guided toward based on the detected hand position.
Lydecker then goes on to supply the real world location and sensor navigation aspect that the applicant alleges are absent. Lydecker teaches stored and collected data for objects in the premises including object coordinate data, 3D object coordinate views and transformation of those object coordinates relative to a user’s view based on the user device camera. A person of ordinary skill in the art would have found it obvious to combine Hirano’s camera based hand position guidance with Lydecker’s real world object location and navigational teachings to guide a user’s hand toward an object based on camera derived positional information.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4-7, 9, and 13-16 are rejected under 35 U.S.C. 103 as being unpatentable over Lock et al (Active Object Search with a Mobile Device for People with Visual Impairments), in view of Lydecker et al (US20170035645A1).
Regarding claim 1, Lock discloses, a mobile device comprising: a processor connected to a non-transitory computer readable medium having an application stored thereon, the application defining a method that is performed when the processor runs the application (Lock, 4.2.3, we incorporated the trained system into an app (see Figure 6) for an Asus ZenPhone AR smartphone, running Android 7.0, with Google’s augmented reality toolkit (ARCore), which provides the necessary 3D pose of the device), identifying an object that is the selected item to be found from the camera data; in response to identifying the object, determining a location of the object in the area around the mobile device; (Lock, 4, the actual location of the latter is unknown, meaning that the system will guide the user towards the most likely location where the object might be found, based on its internal knowledge of spatial relations between objects (e.g. a computer monitor is more likely to be above than below a desk). This path is generated one waypoint at a time and is updated with every new object observation captured by the camera, or after a re-orientation of the latter beyond a certain angle); and providing tactile instructions and/or audible instructions via the mobile device to the user while the mobile device is held by at least one hand of the user to instruct the user where to move vertically and horizontally based on a position of the camera and the determined location of the object in the area around the mobile device (Lock, 4.2.3, in a real application for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions. However, since we are mainly interested in evaluating the control algorithm of our system and not the interface) … (Lock, 3, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone, as pictured in Figure 1) …
PNG
media_image1.png
154
191
media_image1.png
Greyscale
Fig. 1
(Lock, 4.1.2 and Figure 3, the policy produced by an MDP defines the action the agent will take when it finds itself in any given state. In this case, the action is the direction of the next waypoint relatively to the current device’s pose. The possible actions are given by A = {UP,DOWN,LEFT,RIGHT}), identifying, based on the camera data, that the object is out of a camera view of the camera of the mobile device (2, PREVIOUS WORK, in this paper, we implement such an active vision system with a human in the loop that guides the user towards an out-of-view target object. Our system exploits prior knowledge of the objects spatial distribution within an indoor environment, learned from a dataset of real-world images, and the history of past object observations made during the search), and in response to identifying that the object is out of the camera view of the camera of the mobile device, providing tactile instructions and/or audible instructions via the mobile device to the user (4.2.3 Smartphone Application, tracking the camera’s pose allows the app to infer the current state and choose the optimal action to take next) … (4.2.3 Smartphone Application, for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions), while the mobile device is held by at least one hand of the user to instruct the user where to move the camera view (1, Introduction, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone), See figure 1. vertically and horizontally to place the object in the camera view of the camera of the mobile device, (Figure 5, examples of the spatial relationships between the desk, keyboard and mouse objects. Each square corresponds to the probability of executing an action (top square for UP, left square for LEFT, etc.) … (Figure 6, a screenshot of the smartphone interface showing an example of guidance instruction (down-left in this case) towards a waypoint and the QR-object scanner area).
PNG
media_image2.png
247
327
media_image2.png
Greyscale
Fig. 2
Additionally, Lydecker who is in the same field of endeavor of devices for guiding the visually impaired discloses, (Lydecker, Paragraph 0098, Lines 1-8, The lighting device processor 47 is configured to update the mapping data based on the user input and the location indicated by the user's smartphone 63 or mobile device. Subsequently, when a visually impaired user needs the water bottle, the visually impaired user speaks a query (e.g., “where is my water bottle”). In response to the visually impaired user's query, the lighting device may generate an indication of an object location based on the updated mapping data), and in response to identifying that the object is out of the camera view of the camera of the mobile device, providing tactile instructions and/or audible instructions via the mobile device to the user (0061, examples may include audio examples that use sound to generate a scene in the user's brain. The audio examples may also include ultrasonic systems that facilitate the generation of directed audio).
A person of ordinary skill in the art would have found it obvious to incorporate Lydecker's object-location mapping and audio navigational output into Lock's smartphone based active search system in order to provide location based assistance to a visually impaired user with predictable results. The results being the 3D object localization guiding the user toward the target via audio/tactile cues while the mobile device is in a user’s hand. This would solve the same problem addressed by both references using known AR toolkit capabilities and well understood signal smoothing techniques.
Further justification for combining or modifying this disclosure not only comes from the state of the art but from Lydecker (0133, the hardware elements, operating systems and programming languages of such computer and/or mobile user terminal devices also are conventional in nature, and it is presumed that those skilled in the art are adequately familiar therewith).
Regarding claim 4, Lock and Lydecker disclose, the mobile device of claim 1, as discussed supra. Additionally, Lock discloses the mobile device is a cell phone, a mobile communication terminal, a smart phone, or a smart watch (Lock, 3, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone, as pictured in Figure 1).
Regarding claim 5, Lock and Lydecker disclose the mobile device of claim 1, as discussed supra. Additionally, Lock discloses, the method also comprises: determining the position of the camera relative to the determined location of the object in the area around the mobile device (Lock, 4.2.2, the system uses the waypoint’s location to provide the user with guidance instructions (i.e. u in Figure 2). The policy actions, and waypoints by extension, are relative to the current camera’s pan-tilt orientation), the position of the camera being a proxy for the at least one a hand of the user (Lock, 3, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone, as pictured in Figure 1). updating the determined location of the camera relative to the determined location of the object in the area around the mobile device to account for movement of the camera that occurs in response to the providing of the tactile instructions and/or audible instructions (Lock, 4, this path is generated one waypoint at a time and is updated with every new object observation captured by the camera, or after a re-orientation of the latter beyond a certain angle); and providing updated tactile instructions and/or audible instructions via the mobile device to the user to instruct the user where to move vertically and horizontally based on the determined updated position of the camera and the determined location of the object in the area around the mobile device (Lock, 3, the reference signal, r, is the object the user wishes to capture with the smartphone’s camera. The goal of the control block, K, is to generate human interpretable instructions, u, to guide the user towards the target object. The process to be controlled involves a human, H, who interprets the instruction and executes a physical action, u∗ to actually manipulate the smartphone’s camera, P. A new observation, y, with the camera is then fed back to the loop and the error signal, e, is updated accordingly).
Regarding claim 6, Lock and Lydecker disclose the mobile device of claim 1, as discussed supra. Additionally, Lock discloses generating a graphical user interface (GUI) on a display of the mobile device to display location information based on the determined location of the object in the area around the mobile device and the position of the camera (Lock, 4.2.3, we are mainly interested in evaluating the control algorithm of our system and not the interface (K and not u), our current prototype generates guidance instructions with four on-screen arrows (see Figure 6). Obviously, this visual interface is only used for debugging and experimental evaluation of the controller, and it will be replaced by an opportune audio interface, e.g. (Bellotto, 2013), at a second stage).
Regarding claim 7, Lock and Lydecker disclose the mobile device of claim 6, as discussed supra. Additionally, Lock discloses, the method comprises: updating the GUI in response to selection of a guide icon to initiate the mobile device performing the providing of the tactile instructions and/or audible instructions via the mobile device to the user to instruct the user where to move so the user moves toward the object based on the position of the camera and the determined location of the object in the area around the mobile device (Lock, 5.1, the participant started by pressing a button on the app, which randomly selected a target object and then guided the user towards it. Since the participants were allowed to use the smartphone’s display, the target was randomly selected by the app without informing them) … (Lock, 4.2.3, in a real application for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions) … (Obviously, this visual interface is only used for debugging and experimental evaluation of the controller, and it will be replaced by an opportune audio interface, e.g. (Bellotto, 2013), at a second stage).
Regarding claim 9, Lock discloses, identifying an object that is the selected item to be found from the camera data; in response to identifying the object, determining a location of the object in the area around the mobile device; (Lock, 4, the actual location of the latter is unknown, meaning that the system will guide the user towards the most likely location where the object might be found, based on its internal knowledge of spatial relations between objects (e.g. a computer monitor is more likely to be above than below a desk). This path is generated one waypoint at a time and is updated with every new object observation captured by the camera, or after a re-orientation of the latter beyond a certain angle); and providing tactile instructions and/or audible instructions via the mobile device to the user while the mobile device is held by at least one hand of the user to instruct the user where to move vertically and horizontally based on a position of the camera and the determined location of the object in the area around the mobile device (Lock, 4.2.3, in a real application for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions. However, since we are mainly interested in evaluating the control algorithm of our system and not the interface) … (Lock, 3, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone, as pictured in Figure 1) … (Lock, 4.1.2 and Figure 3, the policy produced by an MDP defines the action the agent will take when it finds itself in any given state. In this case, the action is the direction of the next waypoint relatively to the current device’s pose. The possible actions are given by A = {UP,DOWN,LEFT,RIGHT}), identifying, based on the camera data, that the object is out of a camera view of the camera of the mobile device (2, PREVIOUS WORK, in this paper, we implement such an active vision system with a human in the loop that guides the user towards an out-of-view target object. Our system exploits prior knowledge of the objects spatial distribution within an indoor environment, learned from a dataset of real-world images, and the history of past object observations made during the search), and in response to identifying that the object is out of the camera view of the camera of the mobile device, providing tactile instructions and/or audible instructions via the mobile device to the user (4.2.3 Smartphone Application, tracking the camera’s pose allows the app to infer the current state and choose the optimal action to take next) … (4.2.3 Smartphone Application, for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions), while the mobile device is held by at least one hand of the user to instruct the user where to move the camera view (1, Introduction, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone), See figure 1. vertically and horizontally to place the object in the camera view of the camera of the mobile device, (Figure 5, examples of the spatial relationships between the desk, keyboard and mouse objects. Each square corresponds to the probability of executing an action (top square for UP, left square for LEFT, etc.) … (Figure 6, a screenshot of the smartphone interface showing an example of guidance instruction (down-left in this case) towards a waypoint and the QR-object scanner area).
Additionally, Lydecker discloses, responding to input selecting an item to be found by utilizing at least one camera of the mobile device to receive camera data of an area around the mobile device; identifying an object that is the selected item to be found from the camera data; in response to identifying the object, determining a location of the object in the area around the mobile device (Lydecker, Paragraph 0098, Lines 1-8, The lighting device processor 47 is configured to update the mapping data based on the user input and the location indicated by the user's smartphone 63 or mobile device. Subsequently, when a visually impaired user needs the water bottle, the visually impaired user speaks a query (e.g., “where is my water bottle”). In response to the visually impaired user's query, the lighting device may generate an indication of an object location based on the updated mapping data), and in response to identifying that the object is out of the camera view of the camera of the mobile device, providing tactile instructions and/or audible instructions via the mobile device to the user (0061, examples may include audio examples that use sound to generate a scene in the user's brain. The audio examples may also include ultrasonic systems that facilitate the generation of directed audio).
Regarding claim 13, Lock and Lydecker disclose, the non-transitory computer readable medium of claim 9 as discussed supra. Additionally Lock discloses the mobile device is a cell phone, a mobile communication terminal, a smart phone, or a smart watch (Lock, 3, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone, as pictured in Figure 1).
Regarding claim 14, Lock and Lydecker disclose, the non-transitory computer readable medium of claim 9 as discussed supra. Additionally Lock discloses generating a graphical user interface (GUI) on a display of the mobile device to display location information based on the determined location of the object in the area around the mobile device and the position of the camera (Lock, 4.2.3, we are mainly interested in evaluating the control algorithm of our system and not the interface (K and not u), our current prototype generates guidance instructions with four on-screen arrows (see Figure 6). Obviously, this visual interface is only used for debugging and experimental evaluation of the controller, and it will be replaced by an opportune audio interface, e.g. (Bellotto, 2013), at a second stage).
Regarding claim 15, Lock and Lydecker disclose, the non-transitory computer readable medium of claim 14 as discussed supra. Additionally Lock discloses updating the GUI in response to selection of a guide icon to initiate the mobile device performing the providing of the tactile instructions and/or audible instructions via the mobile device to the user to instruct the user where to move so the user moves toward the object based on the position of the camera and the determined location of the object in the area around the mobile device (Lock, 5.1, the participant started by pressing a button on the app, which randomly selected a target object and then guided the user towards it. Since the participants were allowed to use the smartphone’s display, the target was randomly selected by the app without informing them) … (Lock, 4.2.3, in a real application for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions) … (Obviously, this visual interface is only used for debugging and experimental evaluation of the controller, and it will be replaced by an opportune audio interface, e.g. (Bellotto, 2013), at a second stage).
Regarding claim 16 Lock and Lydecker disclose the non-transitory computer readable medium of claim 9 as discussed supra. Additionally, Lydecker discloses the providing of the audible instructions comprises: periodic emission of sound to indicate a proximity of the camera to the object (Lydecker, Paragraph 0126, Lines 29-31, the outputted navigational instructions may be visual stimuli, such as arrows or boundary lines, or may be audio stimuli, such as “turn right” or “turn left,” or may be tactile stimuli provided by the user wearable device); and audible directional instruction output to the user to change a direction of the camera to move the mobile device closer to the object based on the determined location of the object in the area around the mobile device and a determined position of the camera relative to the object (Lydecker, Paragraph 0124, Lines 16-18, in addition, audio navigational commands may be generated such as “turn left,” “turn right,” “move to your right,” “stop,” “the tablet is in front of you,” and the like).
Claims 2 and 10-11 are all rejected under 35 U.S.C. 103 as being unpatentable over Lock et al (Active Object Search with a Mobile Device for People with Visual Impairments), in view of Lydecker et al (US20170035645A1), further in view of Huynh (Semantic Labeling and Object Registration for Augmented Reality Language Learning)
Regarding claim 2, Lock and Lydecker disclose the mobile device of claim 1 as discussed supra. Additionally, Huynh who is in the same field of endeavor of object registration for augmented reality discloses determining of the location of the object in the area around the mobile device includes: generating a pre-selected number of location samples via ray casting of the camera data (2.1 Semantic Labeling, We stream video frames from the built-in HoloLens front facing camera to a server running on an MSI VR backpack … The resulting 2D bounding boxes and labels are sent back to the HoloLens, along with the original camera pose, where we project the center point of each 2D bounding box onto the 3D mesh via raycast from the original camera pose) and determining the location via averaging predicted locations determined for the pre- selected number of location samples such that the determined location is an average predicted location (2.2 Object Registration, if the Euclidean distance between these subsequent 3D positions are within a threshold D (e.g. 50 centimeters away for a keyboard object), we average these positions and affix the object).
It would be obvious to one skilled in the art to combine the combination of Lock and Lydecker with Huynh. Huynh teaches a known way to improve the precision and stability of the object location determination by projecting camera detection in 3D space via ray casting from the original camera pose, and using multiple steamed frames to establish the object location which are then averaged. Applying Huynh’s ray cast and averaging technique to the Lock-Lydecker combination would have predictably yielded more accurate location based audio guidance for a user searching for an object in an environment.
Further justification for combining or modifying the combination of Lock and Lydecker with Huynh not only comes from the state of the art but from Huynh (4. Conclusion, We described how to integrate eye tracking into our framework to allow for user selection or activation of annotations. We discuss how the combi nation of these technologies opens up new and interesting research directions for the growing field of AR language learning).
Regarding claim 10, Lock and Lydecker disclose the non-transitory computer readable medium of claim 9 as discussed supra. Additionally, Huynh discloses, determining of the location of the object in the area around the mobile device includes: generating a pre-selected number of location samples via ray casting of the camera data (2.1 Semantic Labeling, We stream video frames from the built-in HoloLens front facing camera to a server running on an MSI VR backpack … The resulting 2D bounding boxes and labels are sent back to the HoloLens, along with the original camera pose, where we project the center point of each 2D bounding box onto the 3D mesh via raycast from the original camera pose) and determining the location via averaging predicted locations determined for the pre- selected number of location samples such that the determined location is an average predicted location (2.2 Object Registration, if the Euclidean distance between these subsequent 3D positions are within a threshold D (e.g. 50 centimeters away for a keyboard object), we average these positions and affix the object).
Regarding claim 11, Lock, Lydecker, and Huynh disclose, the non-transitory computer readable medium of claim 10 as discussed supra. Additionally Lock discloses determining the position of the camera relative to the determined location of the object in the area around the mobile device (Lock, 4.2.2, the system uses the waypoint’s location to provide the user with guidance instructions (i.e. u in Figure 2). The policy actions, and waypoints by extension, are relative to the current camera’s pan-tilt orientation), the position of the camera being a proxy for the at least one a hand of the user (Lock, 3, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone, as pictured in Figure 1). updating the determined location of the camera relative to the determined location of the object in the area around the mobile device to account for movement of the camera that occurs in response to the providing of the tactile instructions and/or audible instructions (Lock, 4, this path is generated one waypoint at a time and is updated with every new object observation captured by the camera, or after a re-orientation of the latter beyond a certain angle); and providing updated tactile instructions and/or audible instructions via the mobile device to the user to instruct the user where to move vertically and horizontally based on the determined updated position of the camera and the determined location of the object in the area around the mobile device (Lock, 3, the reference signal, r, is the object the user wishes to capture with the smartphone’s camera. The goal of the control block, K, is to generate human interpretable instructions, u, to guide the user towards the target object. The process to be controlled involves a human, H, who interprets the instruction and executes a physical action, u∗ to actually manipulate the smartphone’s camera, P. A new observation, y, with the camera is then fed back to the loop and the error signal, e, is updated accordingly).
Claims 3 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Lock et al (Active Object Search with a Mobile Device for People with Visual Impairments), in view of Lydecker et al (US20170035645A1), further in view of Huynh (Semantic Labeling and Object Registration for Augmented Reality Language Learning), further in view of Xue et al. (US20190228559A1).
Regarding claim 3, Lock, Lydecker, and Huynh disclose the mobile device of claim 2, as discussed supra. Additionally, Lydecker discloses the user moves the mobile device toward the object in response to the provided tactile instructions and/or audible instructions (Lydecker, Paragraph 0031, Lines 9-11, in response to a user device query, generate an indication of the movable object location based on the updated mapping data; and deliver the indication based on the updated mapping data to the user device).
Furthermore, Xue who is in the same field of endeavor of ray casting projections from perspective views discloses, updating the determined location by obtaining a moving average based on a generation of additional location samples obtained via ray casting (Xue, 0056, when associating masks, a geometry consistency check is performed to take past frames into consideration so that the masks of the corresponding objects do not make a precipitous movement that is unlikely. In one example, a 5-frame window moving average with decaying weights for past frames may be used to perform this geometry check).
It would be obvious to one skilled in the art to combine the combination of Lock, Lydecker, and Huynh with Xue. Applying Xue’s ray based 3D localization and moving average to Lock and Lydecker’s system would have predictably improved the accuracy of the determined object location used for guidance.
Further justification for combining or modifying this disclosure not only comes from the state of the art but from Xue (0011, Those skilled in the art will further understand that the exemplary embodiments may be used in sporting events not involving a goal or non-sport related scenarios, such as those that include recognizing an object(s) and casting a ray projection from a desired perspective view to an area).
Regarding claim 12, Lock, Lydecker, and Huynh disclose the non-transitory computer readable medium of claim 11 as discussed supra. Additionally, Lydecker discloses the user moves the mobile device toward the object in response to the provided tactile instructions and/or audible instructions (Lydecker, Paragraph 0031, Lines 9-11, in response to a user device query, generate an indication of the movable object location based on the updated mapping data; and deliver the indication based on the updated mapping data to the user device).
Furthermore, Xue discloses, updating the determined location by obtaining a moving average based on a generation of additional location samples obtained via ray casting (Xue, 0056, when associating masks, a geometry consistency check is performed to take past frames into consideration so that the masks of the corresponding objects do not make a precipitous movement that is unlikely. In one example, a 5-frame window moving average with decaying weights for past frames may be used to perform this geometry check).
Claims 8 is rejected under 35 U.S.C. 103 as being unpatentable over Lock et al (Active Object Search with a Mobile Device for People with Visual Impairments), in view of Lydecker et al (US20170035645A1), further in view of Bellotto (A Portable Navigation System with an Adaptive Multimodal Interface for the Blind).
Regarding claim 8, Lock and Lydecker disclose the mobile device of claim 1 as discussed supra. Additionally, Bellotto who is in the same field of endeavor of adaptive multimodal interfaces for the blind discloses, providing of the tactile instructions and/or the audible instructions comprises: periodic emission of sound to indicate a proximity of the camera to the object such that the emission of sound is increased in volume as the camera is moved closer to the object and the emission of sound is decreased in volume as the camera is moved farther from the object (Bellotto, Planned Experiments, set of experiments have been planned to determine how the four feedback parameters of the multimodal UI affect user performance. These parameters are the pitch gain rate, the volume gain rate, vibration period and voice command frequency. The sound panning is controlled by a separate library so it is not included as a variable in these tests. The first set of experiments is to have the user traverse through an obstacle course on a 2D plane where the exit of the course will be the final target. Since the latter is at a fixed height, the pitch variable is not considered here, which simplifies the experiment. Only the vibration, voice commands and volume gain are used for feedback. A separate set of tests will be conducted to determine the effect of the pitch and spatial sound, where the user will be directed to point the Tango to look at virtual targets floating at different levels and distances to the user. Here the vibration parameter will be eliminated and only the voice commands, pitch and volume gain will be used as feedback parameters);
It would be obvious to one skilled in the art to combine the combination of Lock and Lydecker with Bellotto to utilize the disclosed tactile and audible instructions. This would again enable a visually impaired person to be assisted in a greater more accurate capacity for the exact positioning of their hand to complete an objective.
Further justification for combining or modifying the combination of Lock and Lydecker with Bellotto not only comes from the state of the art but from Lock (Lock, 4.2.3, we are mainly interested in evaluating the control algorithm of our system and not the interface (K and not u), our current prototype generates guidance instructions with four on-screen arrows (see Figure 6). Obviously, this visual interface is only used for debugging and experimental evaluation of the controller, and it will be replaced by an opportune audio interface, e.g. (Bellotto, 2013), at a second stage).
Claim 17 and 21-24 is rejected under 35 U.S.C. 103 as being unpatentable over Lock et al (Active Object Search with a Mobile Device for People with Visual Impairments), in view of Lydecker et al (US20170035645A1), further in view of Huynh (Semantic Labeling and Object Registration for Augmented Reality Language Learning), further in view of Billings et al. (US20190329364A1).
Regarding claim 17, Lock discloses, the mobile device responding to input selecting an item to be found by utilizing at least one camera of the mobile device to receive camera data of an area around the mobile device; the mobile device identifying an object that is the selected item to be found from the camera data (Lock, 5.1, for the experiment, the MDP policies were integrated into an Android application that uses the camera to provide observation data and track the pose and viewing direction) … (Lock, 5.1, the participant started by pressing a button on the app, which randomly selected a target object and then guided the user towards it. Since the participants were allowed to use the smartphone’s display, the target was randomly selected by the app without informing them); providing audible instructions and/or tactile instructions via the mobile device to the user to instruct the user where to move vertically and horizontally based on a position of the camera and the determined location of the object in the area around the mobile device (Lock, 4.2.3, in a real application for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions. However, since we are mainly interested in evaluating the control algorithm of our system and not the interface) …
(Lock, 3, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone, as pictured in Figure 1) …
(Lock, 4.1.2 and Figure 3, the policy produced by an MDP defines the action the agent will take when it finds itself in any given state. In this case, the action is the direction of the next waypoint relatively to the current device’s pose. The possible actions are given by A = {UP,DOWN,LEFT,RIGHT}) identifying, based on the camera data, that the object is out of a camera view of the camera of the mobile device (2, PREVIOUS WORK, in this paper, we implement such an active vision system with a human in the loop that guides the user towards an out-of-view target object. Our system exploits prior knowledge of the objects spatial distribution within an indoor environment, learned from a dataset of real-world images, and the history of past object observations made during the search), and in response to identifying that the object is out of the camera view of the camera of the mobile device, providing tactile instructions and/or audible instructions via the mobile device to the user (4.2.3 Smartphone Application, tracking the camera’s pose allows the app to infer the current state and choose the optimal action to take next) … (4.2.3 Smartphone Application, for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions), while the mobile device is held by at least one hand of the user to instruct the user where to move the camera view (1, Introduction, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone), See figure 1 vertically and horizontally to place the object in the camera view of the camera of the mobile device (Figure 5, examples of the spatial relationships between the desk, keyboard and mouse objects. Each square corresponds to the probability of executing an action (top square for UP, left square for LEFT, etc.) … (Figure 6, a screenshot of the smartphone interface showing an example of guidance instruction (down-left in this case) towards a waypoint and the QR-object scanner area). See figure 5.
PNG
media_image2.png
247
327
media_image2.png
Greyscale
Additionally, Lydecker discloses, in response to identifying the object, the mobile device determining a location of the object in the area around the mobile device (Lydecker, Paragraph 0029, Lines 18-21, examples of the different techniques the different sensors may use include, but are not limited to, a structured-light technique or a time of flight technique for determining the layout and/or structure of an area as well as determining the location of objects with the area).
Furthermore, Huynh discloses, the mobile device determining the location of the object in the area around the mobile device includes: generating a pre-selected number of location samples via ray casting of the camera data (2.1 Semantic Labeling, We stream video frames from the built-in HoloLens front facing camera to a server running on an MSI VR backpack … The resulting 2D bounding boxes and labels are sent back to the HoloLens, along with the original camera pose, where we project the center point of each 2D bounding box onto the 3D mesh via raycast from the original camera pose). and determining the location via averaging predicted locations determined for the pre-selected number of location samples such that the determined location is an average predicted location (2.2 Object Registration, if the Euclidean distance between these subsequent 3D positions are within a threshold D (e.g. 50 centimeters away for a keyboard object), we average these positions and affix the object).
Finally, Billings discloses, a method of providing hand guidance to direct a user toward an object via a mobile device (Billings, 0103, the processing circuitry 602 may be configured to cue the user move, for example, their head to reposition the camera 620 to be centered at the selected object or move their hand in the captured image in a direction to engage the selected object. For example, the assistance feedback may be verbal cues from an output device indicated that the user should, for example, move their hand to the left).
It would be obvious to one skilled in the art to combine the combination of Lock, Lydecker, and Huynh with Billings to utilize the disclosed hand guidance instructions. This would again enable a visually impaired person to be assisted in a greater more accurate capacity for the exact positioning of their hand to complete an objective.
Further justification for combining or modifying Lock and Billings not only comes from the state of the art but from Billings (0116, many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented).
Claims 18 is rejected under 35 U.S.C. 103 as being unpatentable over Lock et al (Active Object Search with a Mobile Device for People with Visual Impairments), in view of Lydecker et al (US20170035645A1), further in view of Huynh (Semantic Labeling and Object Registration for Augmented Reality Language Learning), further in view of Bellotto (A Portable Navigation System with an Adaptive Multimodal Interface for the Blind), further in view of Billings et al. (US20190329364A1).
Regarding claim 18, Lock, Lydecker, Huynh, and Billings disclose the method of claim 17 as discussed supra. Additionally, Bellotto discloses, providing of the tactile instructions and/or the audible instructions comprises: periodic emission of sound to indicate a proximity of the camera to the object such that the emission of sound is increased in volume as the camera is moved closer to the object and the emission of sound is decreased in volume as the camera is moved farther from the object (Bellotto, Planned Experiments, set of experiments have been planned to determine how the four feedback parameters of the multimodal UI affect user performance. These parameters are the pitch gain rate, the volume gain rate, vibration period and voice command frequency. The sound panning is controlled by a separate library so it is not included as a variable in these tests. The first set of experiments is to have the user traverse through an obstacle course on a 2D plane where the exit of the course will be the final target. Since the latter is at a fixed height, the pitch variable is not considered here, which simplifies the experiment. Only the vibration, voice commands and volume gain are used for feedback. A separate set of tests will be conducted to determine the effect of the pitch and spatial sound, where the user will be directed to point the Tango to look at virtual targets floating at different levels and distances to the user. Here the vibration parameter will be eliminated and only the voice commands, pitch and volume gain will be used as feedback parameters); and audible directional instruction output to the user to change a direction of the camera vertically and horizontally to move the mobile device closer to the object based on the determined location of the object in the area around the mobile device and a position of the camera relative to the determined location of the object (Bellotto, 3, the navigation system in its current form guides the user toward a particular target using voice commands and spatialised sound, while vibration is used for obstacle avoidance. The spatial sound has 3 parameters that convey the target’s position relative to where the user is pointing the Tango’s camera and its current position: the pitch (elevation angle to the target), gain (absolute distance to the target) and 3D panning (whether the target is to the left or right of the user)).
It would be obvious to one skilled in the art to combine Lock, Lydecker, Huynh, and Billings with Bellotto to utilize the disclosed tactile and audible instructions. This would again enable a visually impaired person to be assisted in a greater more accurate capacity for the exact positioning of their hand to complete an objective.
Further justification for combining or modifying Lock and Bellotto not only comes from the state of the art but from Lock (Lock, 4.2.3, we are mainly interested in evaluating the control algorithm of our system and not the interface (K and not u), our current prototype generates guidance instructions with four on-screen arrows (see Figure 6). Obviously, this visual interface is only used for debugging and experimental evaluation of the controller, and it will be replaced by an opportune audio interface, e.g. (Bellotto, 2013), at a second stage).
Claims 20 is rejected under 35 U.S.C. 103 as being unpatentable over Lock et al (Active Object Search with a Mobile Device for People with Visual Impairments), in view of Lydecker et al (US20170035645A1), further in view of Huynh (Semantic Labeling and Object Registration for Augmented Reality Language Learning), further in view of Bellotto (A Portable Navigation System with an Adaptive Multimodal Interface for the Blind), further in view of Billings et al. (US20190329364A1), further in view of Xue et al. (US20190228559A1).
Regarding claim 20, Lock, Lydecker, Huynh, and Billings and Bellotto disclose the method of claim 18 as discussed supra. Additionally, Xue discloses updating the determined location by obtaining a moving average based on a generation of additional location samples obtained via ray casting of the camera data (Xue, 0056, when associating masks, a geometry consistency check is performed to take past frames into consideration so that the masks of the corresponding objects do not make a precipitous movement that is unlikely. In one example, a 5-frame window moving average with decaying weights for past frames may be used to perform this geometry check).
Additionally, Lydecker discloses, while the user moves the mobile device toward the object in response to the provided audible instructions and/or tactile instructions (Lydecker, Paragraph 0031, Lines 9-11, in response to a user device query, generate an indication of the movable object location based on the updated mapping data; and deliver the indication based on the updated mapping data to the user device).
It would be obvious to one skilled in the art to combine the combination of Lock, Lydecker, and Huynh, Billings, and Bellotto with Xue . Applying Xue’s ray based 3D localization and moving average to Lock and Lydecker’s system would have predictably improved the accuracy of the determined object location used for guidance.
Further justification for combining or modifying this disclosure not only comes from the state of the art but from Xue (0011, Those skilled in the art will further understand that the exemplary embodiments may be used in sporting events not involving a goal or non-sport related scenarios, such as those that include recognizing an object(s) and casting a ray
projection from a desired perspective view to an area).
Regarding claim 21, Lock, Lydecker, Huynh, and Billings disclose the method of claim 17 as discussed supra. Additionally, Lock discloses generating a graphical user interface (GUI) on a display of the mobile device that displays a list of selectable items to facilitate the receipt of the input selecting the item to be found (Lock, 5.1, the participant started by pressing a button on the app, which randomly selected a target object and then guided the user towards it. Since the participants were allowed to use the smartphone’s display, the target was randomly selected by the app without informing them) … (Lock, 4.2.3, in a real application for people with visual impairments, the position of the waypoint would be given to the user by a set of audio or vibrotactile instructions) … (Obviously, this visual interface is only used for debugging and experimental evaluation of the controller, and it will be replaced by an opportune audio interface, e.g. (Bellotto, 2013), at a second stage).
Regarding claim 22, Lock, Lydecker, Huynh, and Billings disclose the method of claim 21 as discussed supra. Additionally, Lock discloses, updating the display of the mobile device in response to receipt of the input selecting the item to be found on the display of the mobile device to display a selectable guide icon in the GUI that is selectable to initiate the mobile device performing the providing of the audible instructions and/or the tactile instructions (Lock, 4.2.3, we are mainly interested in evaluating the control algorithm of our system and not the interface (K and not u), our current prototype generates guidance instructions with four on-screen arrows (see Figure 6). Obviously, this visual interface is only used for debugging and experimental evaluation of the controller, and it will be replaced by an opportune audio interface, e.g. (Bellotto, 2013), at a second stage).
Regarding claim 23, Lock, Lydecker, Huynh, and Billings disclose the method of claim 17 as discussed supra. Additionally, Lock discloses determining the position of the camera relative to the determined location of the object in the area around the mobile device (Lock, 4.2.2, the system uses the waypoint’s location to provide the user with guidance instructions (i.e. u in Figure 2). The policy actions, and waypoints by extension, are relative to the current camera’s pan-tilt orientation), the position of the camera being a proxy for the at least one a hand of the user (Lock, 3, the proposed system implements ideas from the field of active vision (Bajcsy et al., 2017), but replaces the typical electro-mechanical actuators of a moving camera with the body (i.e. arm, hand) of the user holding the smartphone, as pictured in Figure 1). updating the determined location of the camera relative to the determined location of the object in the area around the mobile device to account for movement of the camera that occurs in response to the providing of the tactile instructions and/or audible instructions (Lock, 4, this path is generated one waypoint at a time and is updated with every new object observation captured by the camera, or after a re-orientation of the latter beyond a certain angle); and providing updated tactile instructions and/or audible instructions via the mobile device to the user to instruct the user where to move vertically and horizontally based on the determined updated position of the camera and the determined location of the object in the area around the mobile device (Lock, 3, the reference signal, r, is the object the user wishes to capture with the smartphone’s camera. The goal of the control block, K, is to generate human interpretable instructions, u, to guide the user towards the target object. The process to be controlled involves a human, H, who interprets the instruction and executes a physical action, u∗ to actually manipulate the smartphone’s camera, P. A new observation, y, with the camera is then fed back to the loop and the error signal, e, is updated accordingly).
Regarding claim 24, Lock, Lydecker, Huynh, and Billings disclose the method of claim 23 as discussed supra. Additionally, Lock discloses determining of the position of the camera relative to the determined location of the object is based on sensor data obtained via at least one sensor of the mobile device (Lock, 3, the reference signal, r, is the object the user wishes to capture with the smartphone’s camera. The goal of the control block, K, is to generate human interpretable instructions, u, to guide the user towards the target object. The process to be controlled involves a human, H, who interprets the instruction and executes a physical action, u∗ to actually manipulate the smartphone’s camera, P. A new observation, y, with the camera is then fed back to the loop and the error signal, e, is updated accordingly).
Claims 25 and 29 are all rejected under 35 U.S.C. 103 as being unpatentable over Hirano (WO2015125213A1) in view of Lydecker et al (US20170035645A1).
Regarding claim 25, Hirano discloses, a mobile device for providing hand guidance to direct a user toward an object, the mobile device comprising: a processor connected to a camera and a non-transitory computer readable medium having an application stored thereon (Hirano, Paragraph 12, Lines 13-16, the moving body gesture guidance device 1 is connected to a 3D camera 6, a speaker 7, a HUD (head-up display) 8, and a center display 9 via the I / F unit 3. The calculation unit 2 corresponds to a CPU mounted on the moving body gesture guidance device 1 and performs calculation processing in various controls) … (Hirano, Paragraph 11, Line 7, for example, a device or application that is a target of a gesture operation); the mobile device configured to concurrently track a position of a hand of the user and a location of an object based on camera data of an area around the object (Hirano, Paragraph 8, Lines 1-5, the gesture guidance device for a moving body according to the present invention specifies a position of the operator's hand based on detection information of a sensor that detects the operator's hand, and performs a gesture operation with the position of the operator's hand. Each time the emphasis degree calculation unit calculates the emphasis degree according to the detection of the operator's hand by the sensor, the emphasis degree calculation unit that calculates the emphasis degree according to the difference from the detected position, and the gesture degree with the calculated emphasis degree. A notification control unit that notifies the operation device of the operation to the notification device and guides the operator's hand to a predetermined position) … (Hirano, Abstract, Lines1-3, in the present invention, calculated is a degree of emphasis R that is in accordance with the difference between the position of the hand of an operator (A) specified on the basis of a detection signal from a 3D camera). However, Hirano does not explicitly disclose sensor data of the mobile device to generate audible instructions and/or tactile instructions to output to a user to guide the user to the object via vertical motion and horizontal motion independent of whether the object is within a line of sight of the camera after the location of the object is determined and also independent of whether a hand of the user is in sight of the camera.
Nevertheless, Lydecker discloses, sensor data of the mobile device to generate audible instructions and/or tactile instructions to output to a user to guide the user to the object via vertical motion and horizontal motion independent of whether the object is within a line of sight of the camera after the location of the object is determined and also independent of whether a hand of the user is in sight of the camera (Lydecker, Paragraph 0052 and 0053, Lines 3-7 and 1-4, the mapping sensor 44 is configured to collect mapping data using sensors described above with respect to other examples. The collected mapping data relates to the area of the premises in which the lighting device is located. In this example, the lighting system 35 includes a data storage 46, accessible by the lighting device 37 processor 47, for maintaining area mapping data collected by the mapping sensor of an area in which the lighting device is located in the premises) …. (Lydecker, Paragraph 0126, Lines 29-34, the outputted navigational instructions may be visual stimuli, such as arrows or boundary lines, or may be audio stimuli, such as “turn right” or “turn left,” or may be tactile stimuli provided by the user wearable device. In the case of the user wearable device being a bracelet, an example of tactile stimuli may be pressure applied to the skin of a user at different locations).
It would be obvious to one skilled in the art to combine Hirano with Lydecker to utilize the disclosed hand positioning method. Hirano’s system monitors a user’s hand movements and provides data for the system to calculate guidance and feedback. This would again enable a visually impaired person to be assisted in a greater more accurate capacity for the exact positioning of their hand to complete an objective. Further justification for combining or modifying this disclosure not only comes from the state of the art but from Hirano (Hirano, Description, Paragraph 73, Lines 1-3, in the present invention, within the scope of the invention, any combination of each embodiment, any component of each embodiment can be modified, or any component can be omitted).
Regarding claim 29, Hirano and Lydecker disclose the mobile device of claim 25, as discussed supra. Additionally, Lydecker discloses, the mobile device is a cell phone, a mobile communication terminal, a smart phone, or a smart watch (Lydecker, Paragraph 0031, Lines 9-11, the user terminal equipment such as that shown at 59 may be implemented with any suitable processing device that can communicate and offer a suitable user interface. The terminal 59, for example, is shown as a desktop computer with a wired link into the public WAN 55. However, other terminal types, such as laptop computers, notebook computers, netbook computers, tablet computers, and smartphones may serve as the user terminal computers) … (Lydecker, Paragraph 0097, Lines 12-16, the lighting device 37 processor 47 may respond to inputs received via the additional I/O 44S, such as microphone or other form of input, to a request from a user device, such as wearable device 11, smartphone 63 or other user wearable devices 61, such as a smartwatch, in the area).
Claim 26 is rejected under 35 U.S.C. 103 as being unpatentable over Hirano (WO2015125213A1) in view of Lydecker et al (US20170035645A1), further in view of Huynh (Semantic Labeling and Object Registration for Augmented Reality Language Learning).
Regarding claim 26, Hirano and Lydecker disclose the mobile device of claim 25, as discussed supra. Additionally, Huynh discloses the location of the object in the area around the mobile device is determined via a location determination process that includes: generating a pre-selected number of location samples via ray casting of the camera data (2.1 Semantic Labeling, We stream video frames from the built-in HoloLens front facing camera to a server running on an MSI VR backpack … The resulting 2D bounding boxes and labels are sent back to the HoloLens, along with the original camera pose, where we project the center point of each 2D bounding box onto the 3D mesh via raycast from the original camera pose), and determining the location of the object via averaging predicted locations determined for the pre-selected number of location samples such that the determined location of the object is an average predicted location (2.2 Object Registration, if the Euclidean distance between these subsequent 3D positions are within a threshold D (e.g. 50 centimeters away for a keyboard object), we average these positions and affix the object).
It would be obvious to one skilled in the art to combine the combination of Hirano and Lydecker with Huynh. Huynh teaches a known way to improve the precision and stability of the object location determination by projecting camera detection in 3D space via ray casting from the original camera pose, and using multiple steamed frames to establish the object location which are then averaged. Applying Huynh’s ray cast and averaging technique to the Hirano and Lydecker combination would have predictably yielded more accurate location based audio guidance for a user searching for an object in an environment.
Further justification for combining or modifying the combination of Hirano and Lydecker with Huynh not only comes from the state of the art but from Huynh (4. Conclusion, We described how to integrate eye tracking into our framework to allow for user selection or activation of annotations. We discuss how the combi nation of these technologies opens up new and interesting research directions for the growing field of AR language learning).
Claim 27 is rejected under 35 U.S.C. 103 as being unpatentable over Hirano (WO2015125213A1) in view of Lydecker et al (US20170035645A1), further in view of Huynh (Semantic Labeling and Object Registration for Augmented Reality Language Learning), further in view of Xue.
Regarding claim 27, Hirano, Lydecker, and Huynh disclose the mobile device of claim 26, as discussed supra. Additionally, Lydecker discloses while the user moves the mobile device toward the object in response to the provided audible instructions and/or tactile instructions (Lydecker, Paragraph 0031, Lines 9-11, in response to a user device query, generate an indication of the movable object location based on the updated mapping data; and deliver the indication based on the updated mapping data to the user device).
Additionally, Xue discloses, the location of the object is updated by obtaining a moving average based on a generation of additional location samples obtained via ray casting of the camera data (Xue, 0056, when associating masks, a geometry consistency check is performed to take past frames into consideration so that the masks of the corresponding objects do not make a precipitous movement that is unlikely. In one example, a 5-frame window moving average with decaying weights for past frames may be used to perform this geometry check).
It would be obvious to one skilled in the art to combine the combination of Hirano, Lydecker, and Huynh with Xue. Applying Xue’s ray based 3D localization and moving average to Hirano, Lydecker, and Huynh’s system would have predictably improved the accuracy of the determined object location used for guidance.
Further justification for combining or modifying this disclosure not only comes from the state of the art but from Xue (0011, Those skilled in the art will further understand that the exemplary embodiments may be used in sporting events not involving a goal or non-sport related scenarios, such as those that include recognizing an object(s) and casting a ray projection from a desired perspective view to an area).
Claim 28 is rejected under 35 U.S.C. 103 as being unpatentable over Lydecker et al (US20170035645A1), in view of Hirano (WO2015125213A1), further in view of Lee (Hand-Priming in Object Localization for Assistive Egocentric Vision), further in view of Frankel et al. (US20190026939A1).
Regarding claim 28, Lydecker and Hirano disclose, the mobile device of claim 25, as discussed supra. Additionally, Lydecker discloses the mobile device is configured to update the tactile instructions and/or audible instructions to instruct the user where to move based on the determined updated position of the camera and the determined location of the object in the area around the mobile device (Lydecker, Paragraph 0127 and 0128, Lines 15-27, based on the results of the comparison, the processor or server, at 480, generates updated mapping data and updated indications of the identified objects to the user device. The generated updated mapping data and updated indications may only be the data in the updated collected data that changed in comparison to the first area mapping data. The updating of mapping data may be continuous. Alternatively, the comparison step at 470 may be omitted and all of the processed updated data is provided to the user wearable device to replace the first area mapping data in its entirety. Using the updated mapping data, the navigational instructions are updated to indicate an updated unobstructed pathway through the area based on the provided updated mapping data and updated indications of the identified objects). However, the combination of Lydecker and Hirano does not explicitly disclose, mobile device is configured to determine the position of the camera relative to the determined location of the object in the area around the mobile device, the position of the camera being a proxy for a hand of the user, and the mobile device is configured to update the determined location of the camera relative to the determined location of the object in the area around the mobile device to account for movement of the camera that occurs in response to the tactile instructions and/or the audible instructions.
Additionally, Lee discloses, the position of the camera being a proxy for a hand of the user, (Lee, Hand-Primed Object Localization, Lines 43-46, our model expects egocentric images that are obtained by users taking photography or by extracting salient frames from videos recorded using a wearable camera on the head, the chest, or eyeglasses); and the mobile device is configured to update the determined location of the camera relative to the determined location of the object in the area around the mobile device to account for movement of the camera that occurs in response to the tactile instructions and/or the audible instructions (Lee, Introduction, Lines 47-51, the intuition is that the hand segmentation can help capture semantic relations between the hand and the object of interest, such as the relationship between the hand position and the object position and the relationship between the hand pose and the object size). However, even the combination of Lydecker and Hirano with Lee does not explicitly disclose, the mobile device is configured to determine the position of the camera relative to the determined location of the object in the area around the mobile device.
Nevertheless, Frankel who is in the same field of endeavor of methods for blind and visually impaired person environment navigational assistance discloses, the mobile device is configured to determine the position of the camera relative to the determined location of the object in the area around the mobile device (Frankel, Paragraph 0012, the mobile device can be configured to perform one of recognizing one or more objects in the image using image detection and processing algorithms to visually identify the objects. The location of the objects and environmental features such as open spaces and floors can be based on determining the orientation of the mobile device, camera focus distance to the objects, and on correlation of expected common object dimensions relative to the image).
It would be obvious to one skilled in the art to combine the combination of Lydecker, Hirano, and Lee with Frankel to determine the position of the camera relative to the location of objects around the device. This would enable a visually impaired person to be assisted in a greater more accurate capacity.
Further justification for combining or modifying the combination of Lydecker, Hirano, and Lee with Frankel not only comes from the state of the art but from Frankel (0030, the algorithms for mapping image points to corresponding 3D real space coordinates are then applied, based on the dimensional information obtained either from multiple images, or from well-known everyday object size estimates. These algorithms are well known in the art).
Claim 30 is rejected under 35 U.S.C. 103 as being unpatentable over Lydecker et al (US20170035645A1), in view of Hirano (WO2015125213A1), further in view of Lee (Hand-Priming in Object Localization for Assistive Egocentric Vision).
Regarding claim 30, Hirano and Lydecker disclose the mobile device of claim 25, as discussed supra. Additionally, Lydecker discloses, the mobile device is configured to determine the location of the object via one of: (i) utilization of ray casting, (Lydecker, Paragraph 0034, Lines 7-14, a transformation may be applied to the data that facilitates the transformation to 3D world coordinates. The 3D world coordinates may then be provided to a user wearable device … (Lydecker, Paragraph 0034, Lines 7-14, the reflected energy data collected by the respective mapping sensors 44 may be shared with other lighting devices, or with a LEADER lighting device (e.g., selected as a leader according to known techniques of a group of lighting devices), in the room so a complete mapping of a room may be generated based on the data collected by the individual lighting devices. The complete mapping uses different processing techniques, such affine transformations and other perspective transformations, the processor 47, in each lighting device 37 or the LEADER lighting device, may generate a complete mapping of the area. The complete mapping may be a data representation of the view of the room (i.e., area) based on a combination of mapping sensor data collected by each of the lighting devices located in the room), and (iii) utilizing feature matching to obtain a three dimensional location of the object with respect to the camera (Lydecker, Paragraph 0005, with the advent of modern electronics has come advancement, including advances in the types of light sources as well as advancements in networking and control capabilities of the lighting devices. For example, lighting devices include wireless communication systems that facilitate networking and may include sensors that detect environmental condition data relative to the location of the lighting device). However Lydecker nor Hirano disclose using an artificial intelligence to directly get a three dimensional location of the object from the camera data so that the location of the object is determined with respect to the camera.
Nevertheless, Lee who is in the same field of endeavor of object localization for assistive vision discloses, using an artificial intelligence to directly get a three dimensional location of the object from the camera data so that the location of the object is determined with respect to the camera, (Lee, Introduction, 3.2 Implementation Details, Lines 1-15, our model expects egocentric images that are obtained by users taking photography or by extracting salient frames from videos recorded using a wearable camera on the head, the chest, or eyeglasses. 3424 3.2. Implementation Details As the object localization output is dependent on the hand segmentation output, we separately trained these two networks. First, we trained the hand segmentation network. Then, while freezing the weights of the hand segmentation network, we trained the object localization model. Adam [24] was used to train our hand segmentation and object localization networks. Following are hyperparameters that we set for training: (hand segmentation) 10,000 training steps, 0.00001 learning rate, 16 batch size, and 10−9 Adam’s epsilon; (object localization) 20,000 training steps, 0.00001 learning rate, 8 batch size, and 10−9 Adam’s epsilon. In both of the models, we initialized the weights of the first five convolutional layers (conv1–conv5) with those of the VGG-16 network model [48] pre-trained on ImageNet),
It would be obvious to one skilled in the art to combine the combination of Lydecker and Hirano with Lee to utilize AI to synchronize the location of objects with respect to the camera. Lee’s system trains object localization networks to monitor the movements of objects and provides data for the system to calculate guidance and feedback. This would enable a visually impaired person to be assisted in a greater more accurate capacity.
Further justification for combining or modifying the combination of Lydecker and Hirano with Lee not only comes from the state of the art but from Lee (6. Conclusion, we believe that our method can be further employed in other applications that need to understand hand–object interactions, such as object/action recognition and assistive systems for people with visual impairments).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHANE E DOUGLAS whose telephone number is (703)756-1417. The examiner can normally be reached Monday - Friday 7:30AM - 5:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Christian Chace can be reached on (571) 272-4190.. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.E.D./Examiner, Art Unit 3665
/CHRISTIAN CHACE/Supervisory Patent Examiner, Art Unit 3665