Prosecution Insights
Last updated: October 01, 2026
Application No. 19/063,820

REMOTE ANNOTATION AND NAVIGATION USING AN AR WEARABLE DEVICE

Non-Final OA §103
Filed
Feb 26, 2025
Priority
Oct 13, 2022 — continuation of 12/266,060
Examiner
LIU, GORDON G
Art Unit
Tech Center
Assignee
Snap Inc.
OA Round
1 (Non-Final)
83%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
581 granted / 701 resolved
+22.9% vs TC avg
Moderate +15% lift
Without
With
+14.8%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
36 currently pending
Career history
720
Total Applications
across all art units

Statute-Specific Performance

§101
7.1%
-32.9% vs TC avg
§103
77.3%
+37.3% vs TC avg
§102
3.4%
-36.6% vs TC avg
§112
2.7%
-37.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 701 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending under this Office action. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-8, 10, 12-15, 17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Gauglitz, etc. (US 20160358383 A1) in view of Parkinson, etc. (US 20150220506 A1), further in view of Gavriliuc, etc. (US 20160371885 A1). Regarding claim 1, Gauglitz teaches that a computing device (See Gauglitz: Fig. 20, and [0190], "FIG. 20 is a block diagram of a computing device 2000, according to an embodiment. In one embodiment, multiple such computer systems are used in a distributed network to implement multiple components in a transaction-based environment. An object-oriented, service-oriented, or other architecture may be used to implement such functions and communicate between the multiple systems and components. In some embodiments, the computing device of FIG. 20 is an example of a client device that may invoke methods described herein over a network. In other embodiments, the computing device is an example of a computing device that may be included in or connected to a motion interactive video projection system, as described elsewhere herein. In some embodiments, the computing device of FIG. 20 is an example of one or more of the personal computer, smartphone, tablet, or various servers") comprising: at least one processor (See Gauglitz: Fig. 20, and [0191], "One example computing device in the form of a computer 2010, may include a processing unit 2002, memory 2004, removable storage 2012, and non-removable storage 2014. Although the example computing device is illustrated and described as computer 2010, the computing device may be in different forms in different embodiments. For example, the computing device may instead be a smartphone, a tablet, or other computing device including the same or similar elements as illustrated and described with regard to FIG. 20. Further, although the various data storage elements are illustrated as part of the computer 2010, the storage may include cloud-based storage accessible via a network, such as the Internet"); and at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations (See Gauglitz: Fig. 20, and [0191], "One example computing device in the form of a computer 2010, may include a processing unit 2002, memory 2004, removable storage 2012, and non-removable storage 2014. Although the example computing device is illustrated and described as computer 2010, the computing device may be in different forms in different embodiments. For example, the computing device may instead be a smartphone, a tablet, or other computing device including the same or similar elements as illustrated and described with regard to FIG. 20. Further, although the various data storage elements are illustrated as part of the computer 2010, the storage may include cloud-based storage accessible via a network, such as the Internet"; and [0193], “Computer-readable instructions stored on a computer-readable medium are executable by the processing unit 2002 of the computer 2010. A hard drive (magnetic disk or solid state), CD-ROM, and RAM are some examples of articles including a non-transitory computer-readable medium. For example, various computer programs 2025 or apps, such as one or more applications and modules implementing one or more of the methods illustrated and described herein or an app or application that executes on a mobile device or is accessible via a web browser, may be stored on a non-transitory computer-readable medium”) comprising: receiving an image captured by an extended reality (XR) wearable device, a plurality of 3-dimensional (3D) coordinates corresponding to a plurality of positions within the image (See Gauglitz: Figs. 2 and 10, and [0043], “Under the hood, the system runs a SLAM system and sends the tracked camera pose along with the encoded live video stream to the remote system. The local user's system receives information about annotations from the remote system and uses this information together with the live video to render the augmented view”; [0050], “A 3D surface model is constructed on the fly from the live video stream and from associated camera poses. Keyframes were selected based on a set of heuristics (good tracking quality, low device movement, minimum time interval & translational distance between keyframes), then detect and describe features in the new frame using SIFT. Four closest existing keyframes were chosen and matched against their features (one frame at a time) via an approximate nearest neighbor algorithm and collect matches that satisfy the epipolar constraint (which is known due to the received camera poses) within some tolerance as tentative 3D points. If a feature has previously been matched to features from other frames, we check for mutual epipolar consistency of all observations and merge them into a single 3D point if possible; otherwise, the two 3D points remain as competing hypotheses”; and [0118], " An operational mobile system may be used to implement this idea. The remote user's system received the live video and computer vision-based tracking information from a local user's smartphone or tablet and modeled the environment in 3D based on this data. It enabled the remote user to navigate the scene and to create annotations in it that were then sent back and visualized to the local user in AR. This earlier work focused on the overall system and enabling the navigation and communication within this framework. However, the remote user's interface was rather simple: most notably, it used a standard mouse for most interactions, supplemented by keyboard shortcuts for certain functions, and supported only single-point-based markers". Note that 3D pose of the cameras is mapped to the 3D coordinates corresponding to a plurality of positions within the image, and the keyframes is mapped to the identification of the images), and an identification of the image; displaying the image on a display of the computing device (See Gauglitz: Figs. 1,3, 9, and 11, and [0045], "The system operates at 30 frames per second. System latencies were measured using a camera with 1000 fps, which observed a change in the physical world (a falling object passing a certain height) as well as its image on the respective screen. The latency between physical effect and the local user's tablet display—including image formation on the sensor, retrieval, processing by the SLAM system, rendering of the image, and display on the screen—was measured as 205±22.6 ms"); accessing an indication of an augmentation to the image, the augmentation generated by a user of the computing device (See Gauglitz: Fig. 9, and [0116], "FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings"; and [0136], “In contrast, touchscreens are not only ubiquitous today, but they afford direct interaction without the need for an intermediate representation (i.e., a mouse cursor), and provide haptic feedback during the touch. In this context, however, they have the downside of providing 2D input only. Discussing the implications of this limitation and describing and evaluating appropriate solutions in the context of live remote collaboration is one of the main contributions of this application”. Note that the freehand drawing as the 2D input to the system is mapped to the indication of augmentation generated by the user); determining 3D coordinates of the augmentation based on the plurality of 3D coordinates (See Gauglitz: Figs. 9-12, and [0116], “FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings”; [0145], "Given an individual 2D input location p=(x, y), a depth d can thus be obtained by un-projecting p onto the model of the scene. For a sequence of 2D inputs (a 2D drawing) p.sub.1, . . . , p.sub.n, this results in several alternatives to interpret the depth of the drawing as a whole, of which we consider the following"; and [0178], “FIG. 18 shows a screenshot during an orbit operation 1800, according to an embodiment. As a third gesture, we added orbiting around a given world point p.sup.3D. The corresponding gesture is to keep one finger (relatively) static at a point p and move the second finger in an arc around it (FIG. 15(e)). The orbit center p.sup.3D is determined by un-projecting p onto the scene and remains fixed during the movement; additionally, we maintain the gravity vector (as reported by the local user's device). The rotation around the gravity vector at p.sup.3D is then specified by the movement of the second finger”. Note that 2D screen annotations are un-projected into 3Dworld coordinates using received 3D model, camera pose and depth information so the annotations become world-stabilized, sine the camera is tracked with respect to the scene, these annotations automatically obtain a position in world coordinates, and depth d can thus be obtained by un-projecting p onto the model of the scene.); and sending (See Gauglitz: Figs. 2 and 10, and [0026], "This AR solution enables collaboration between two collaborators in different physical locations. As described herein, a first collaborator is described as being located in physical location A, and may be referred to as a “local user” or “local collaborator.” Similarly, a second collaborator is described as being located in physical location B, and may be referred to as a “remote user” or “remote collaborator.” This AR solution enables the remote collaborator to control his or her viewpoint of physical location A, such as by adjusting camera controls or other viewpoint adjustments. This AR solution enables the remote collaborator to communicate information with visual or spatial reference to physical objects in physical location A, where the information is sent from the remote collaborator to the collaborator in this physical location A. The visual and/or spatial reference information may include identifying objects, locations, directions, or spatial instructions by creating annotations, where the annotations are transmitted to the collaborator in physical location A and visualized using appropriate display technology. For example, the AR solution may include a visual display for the collaborator in physical location A, where the annotations are overlaid onto and anchored to the respective real world object or location"; [0076], ‘Currently, the 3D model is available only on the remote user's side. Thus, annotations can be correctly occluded by the physical scene by the remote user's renderer, but not by the local user's renderer. While other depth cues (most notably, parallax) still indicate the annotation's location, it would be desirable to enable occlusion. The remote system can send either the model geometry or alternatively local visibility information per annotation back to the local device in order to enable occlusion on the local user's side”; and [0127], “The remote user's system may be run on one of a variety of devices, such as a commodity PC or laptop, a cell phone, tablet, virtual reality device, or other device. The embodiment shown in FIG. 10 was generated using a commodity PC. It receives the live video stream and the associated camera poses and models the environment in real time from this data. The remote user can set world-anchored annotations, which are sent back and displayed, to the local user. Further, thanks to the constructed 3D model, he/she can move away from the live view and choose to look at the scene from a different viewpoint via a set of virtual navigation controls. Transitions from one viewpoint to another (including the live view) are rendered seamlessly via image based rendering techniques using the 3D model, cached keyframes, and the live frame appropriately”) the identification, the indication of the augmentation, and the 3D coordinates of the augmentation to the XR wearable device. However, Gauglitz fails to explicitly disclose that an identification of the image; and sending the identification, the indication of the augmentation, and the 3D coordinates of the augmentation to the XR wearable device. However, Parkinson teaches that an identification of the image (See Parkinson: Fig. 1, and [0015], “In another aspect, the invention may be a computer-assisted method of remote document annotation, including selecting, at a host computing platform, a document to be annotated, and providing, at the host computing platform, one or more annotations to the document. The method may further include submitting, at the host computing platform, a location identifier of a desired recipient of the document, and transmitting, by the host computing platform, information representative of the document and information representative of the one or more annotations, to a head mounted display device associated with the location identifier. The method may further include receiving, by the head mounted display device, the information representative of the document and the information representative of the one or more annotations. The method may further include applying, by the head mounted display device, the information representative of the one or more annotations to the document, so as to recreate the one or more annotations provided at the host computing platform. The method may further include displaying, by the head mounted display device, the document together with the annotations”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gauglitz to have an identification of the image as taught by Parkinson in order to provide greater convenience and mobility through hands dependent devices (See Parkinson: Fig. 1, and [0004], “A wireless computing headset contains one or more wireless computing and communication interfaces, enabling data and streaming video capability, and provides greater convenience and mobility through hands dependent devices. For more information concerning such devices, see co-pending patent applications entitled "Mobile Wireless Display Software Platform for Controlling Other Systems and Devices," U.S. application Ser. No. 12/348,648 filed Jan. 5, 2009, "Handheld Wireless Display Devices Having High Resolution Display Suitable For Use as a Mobile Internet Device," PCT International Application No. PCT/US09/38601 filed Mar. 27, 2009, and "Improved Headset Computer," U.S. Application No. 61/638,419 filed Apr. 25, 2012, each of which are incorporated herein by reference in their entirety”). Gauglitz teaches a method and system that may provide for an augmented shared visual space for live mobile remote collaboration on physical tasks with live image, 3D/pose data, remote annotation, un-project to world coordinates, and return the world-stabilized annotation to a local device for displaying; while Parkinson teaches a system and method that may transmit coordinate-bearing annotation among multiple users in the HMD architectures with identifier associated with the location of the HMD device. Therefore, it is obvious to one of ordinary skill in the art to modify Gauglitz by Parkinson to use standard computing devices on the remote site, and transmit spatial coordinate annotation data to a local display device. The motivation to modify Gauglitz by Parkinson is “Use of known technique to improve similar devices (methods, or products) in the same way”. However, Gauglitz, modified by Parkinson, fails to explicitly disclose that sending the identification, the indication of the augmentation, and the 3D coordinates of the augmentation to the XR wearable device. However, Gavriliuc teaches that sending the identification, the indication of the augmentation, and the 3D coordinates of the augmentation to the XR wearable device (See Gavriliuc: Fig. 1, and [0019], “The display device may send captured image data, and also depth data representing environment 100, to a remote display device 106 for presentation to user 108. User 108 may input markup to the image data, for example by entering a handmade sketch via touch input, by inserting an existing drawing or image, etc. using the remote display device 106. The remote display device 106 associates the markup with a three-dimensional location within the environment based upon the location in the image data at which the markup is made and then may send the markup back to the display device 102 for presentation to user 104 via augmented reality display device 102. Via the markup, user 108 may provide instructions, questions, comments, and/or other information to user 106 in the context of the image data. Further, since the markup is associated with a three-dimensional location in a coordinate frame of the augmented reality environment, user 104 may view the markup as spatially augment reality imagery, such that the markup remains in a selected orientation and location relative to environment 100. This may allow user 104 to view the markup from different perspectives by moving within the environment. It will be understood that, in other implementations, any other suitable type of computing device than a wearable augmented reality display device may be used. For example, a video-based augmentation mode that combines a camera viewfinder view with the markup may be used with any suitable display device comprising a camera to present markup-augmented imagery according to the present disclosure”; and [0023], “Each input of markup may be associated with a three-dimensional location in the environment that is mapped to the two-dimensional (i.e. in the plane of the image) location in the image at which the markup was made. For example, in FIG. 2B, the user 108 input the drawing 202 at a location corresponding to the top of a table, and the drawing 204 at a location corresponding to a wall. Thus, the drawings 202, 204 may be associated with the depth locations in the real world scene that are mapped to those locations in the image displayed on display device 106. After receiving input of the markup, computing device 106 may send the markup and associated three-dimensional location of each item of markup to display device 102 for presentation to user 104, to a server for storage and later retrieval, and/or to any other suitable device”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gauglitz to have sending the identification, the indication of the augmentation, and the 3D coordinates of the augmentation to the XR wearable device as taught by Gavriliuc in order to enable operating in environments of arbitrary geometric complexity, and providing a feature-rich virtual navigation that allows remote users to explore the environment independently of local users current camera positions (See Gavriliuc: Fig. 1, and [0019], “The display device may send captured image data, and also depth data representing environment 100, to a remote display device 106 for presentation to user 108. User 108 may input markup to the image data, for example by entering a handmade sketch via touch input, by inserting an existing drawing or image, etc. using the remote display device 106. The remote display device 106 associates the markup with a three-dimensional location within the environment based upon the location in the image data at which the markup is made and then may send the markup back to the display device 102 for presentation to user 104 via augmented reality display device 102”). Gauglitz teaches a method and system that may provide for an augmented shared visual space for live mobile remote collaboration on physical tasks with live image, 3D/pose data, remote annotation, un-project to world coordinates, and return the world-stabilized annotation to a local device for displaying; while Gavriliuc teaches a system and method that may display the image data, receiving an input of a markup to the image data, and associating the markup with a three-dimensional location in the real world scene based on the depth data, and send the markup and the three-dimensional location associated with the markup to another device. Therefore, it is obvious for one of ordinary skills in the arts to modify Gauglitz by Gavriliuc to send the markup and 3D location back to the local users. The motivation to modify Gauglitz by Gavriliuc is “Use of known technique to improve similar devices (methods, or products) in the same way”. Regarding claim 2, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. Further, Gavriliuc teaches that the computing device of claim 1, wherein the identification of the image is an indication of a Coordinated Universal Time (UTC) or an indication of an image identification number (See Gavriliuc: Fig. 1, and [0020], “In some implementations, depth image data may be acquired via a depth camera integrated with the augmented reality display device 102. In such implementations, the fields of view of the depth camera and RGB camera may have a calibrated or otherwise known spatial relationship. This may facilitate integrating or otherwise associating each frame of RBG data with corresponding depth image data”; and Fig. 6, and [0030], “Method 600 further includes, at 608, sending the image data and depth data to device B. Device B receives the image data and depth data at 610, and displays the received image data at 612. At 614, method 600 comprises receiving one or more input(s) of markup to the image data. Any suitable input of markup may be received. For example, as described above, the input of markup may comprise an input of a drawing made by touch or other suitable input. The input of markup further may comprise an input of an image, video, animated item, executable item, text, and/or any other suitable content. At 616, method 600 includes associating each item of markup with a three-dimensional location in the real world scene. Each markup and the associated three-dimensional location are then sent back to device A, as shown at 618. As described above, the data shared between devices may also include an identifier of the real world scene. Such an identifier may facilitate storage and later retrieval of the markup, so that other devices can obtain the markup based upon the identifier when the devices are at the identified location. It will be appreciated that the identifier may be omitted in some implementations”. Note that real-time image and depth streams are conventionally timestamped or numbered so that the correct frame’s depth data is used for the 3D association. Thus, UTC or an image ID number (timestamps) is an obvious species of such an identifier). Regarding claim 3, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. Further, Gauglitz teaches that the computing device of claim 1, wherein the operations further comprise: receiving a video stream from the XR wearable device, the video stream comprising a plurality of images comprising the image (See Gauglitz: Figs. 1-3, and [0050], “A 3D surface model is constructed on the fly from the live video stream and from associated camera poses. Keyframes were selected based on a set of heuristics (good tracking quality, low device movement, minimum time interval & translational distance between keyframes), then detect and describe features in the new frame using SIFT. Four closest existing keyframes were chosen and matched against their features (one frame at a time) via an approximate nearest neighbor algorithm and collect matches that satisfy the epipolar constraint (which is known due to the received camera poses) within some tolerance as tentative 3D points. If a feature has previously been matched to features from other frames, we check for mutual epipolar consistency of all observations and merge them into a single 3D point if possible; otherwise, the two 3D points remain as competing hypotheses”); displaying on the display the video stream and a user interface comprising indications of directions (See Gauglitz Fig. 3, and [0058], “The present subject matter provides the ability to freeze and return to live views. The application starts with the remote view coupled to the local user's live view. With a single right-click, the remote user “freezes” his/her camera at the current pose, for example in order to set annotations precisely, or as a starting point for further navigation. Whenever the remote user's view is not coupled to the local user's view, the latter is displayed to the remote user as an inset (cf. FIG. 3). A click onto this inset or pressing 0 immediately transitions back to the live view”; and [0062], “The present subject matter provides the ability to save and revisit viewpoints. Further, the user can actively save a particular viewpoint to revisit it later. Pressing Alt plus any number key saves the current viewpoint; pressing the respective number key alone revisits this view later. Small numbers along the top of the screen indicate which numbers are currently in use (see FIG. 3”. Note that the viewpoints is mapped to the directions); receiving a selection of a direction of the directions (See Gauglitz Fig. 3, and [0058], “The present subject matter provides the ability to freeze and return to live views. The application starts with the remote view coupled to the local user's live view. With a single right-click, the remote user “freezes” his/her camera at the current pose, for example in order to set annotations precisely, or as a starting point for further navigation. Whenever the remote user's view is not coupled to the local user's view, the latter is displayed to the remote user as an inset (cf. FIG. 3). A click onto this inset or pressing 0 immediately transitions back to the live view”; and [0062], “The present subject matter provides the ability to save and revisit viewpoints. Further, the user can actively save a particular viewpoint to revisit it later. Pressing Alt plus any number key saves the current viewpoint; pressing the respective number key alone revisits this view later. Small numbers along the top of the screen indicate which numbers are currently in use (see FIG. 3”; and [0059], “The present subject matter provides panning and zooming capabilities. By moving the mouse while its right button is pressed, the user can pan the view in a panorama-like fashion (rotate around the current camera position). To prevent the user from getting lost in unmapped areas, we constrain the panning to the angular extent of the modeled environment. To ensure that the system does not appear unresponsive to the user's input while enforcing this constraint, we allow a certain amount of “overshoot” beyond the allowed extent. In this range, further mouse movement away from the modeled environment causes an exponentially declining increase in rotation and visual feedback in the form of an increasingly intense blue gradient along the respective screen border (FIG. 3). Once the mouse button is released, the panning quickly snaps back to the allowed range. Thus, the movement appears to be constrained by a (nonlinear) spring rather a hard wall”. Note that the viewpoints is mapped to the directions); and sending to the XR wearable device an indication of the direction (See Gauglitz: Fig. 1, and [0076], “Currently, the 3D model is available only on the remote user's side. Thus, annotations can be correctly occluded by the physical scene by the remote user's renderer, but not by the local user's renderer. While other depth cues (most notably, parallax) still indicate the annotation's location, it would be desirable to enable occlusion. The remote system can send either the model geometry or alternatively local visibility information per annotation back to the local device in order to enable occlusion on the local user's side”. Note that the navigation information and world-stabilized annotation include the positional and directional information). Regarding claim 4, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 3 as outlined above. Further, Parkinson teaches that the computing device of claim 3, wherein the operations further comprise: receiving an audio stream from the XR wearable device (See Parkinson: Fig. 3, and [0032], “FIGS. 1A and 1B show an example embodiment of a wireless computing headset device 100 (also referred to herein as a head mounted display (HMD) or headset computer (HSC)) that incorporates a high-resolution (VGA or better) micro-display element 1010, and other features described below. HMD 100 can include audio input and/or output devices, including one or more microphones, input and output speakers”); playing the audio stream on the computing device (See Parkinson: Figs. 1A-B, and [0033], “The schematic diagram of FIG. 3 illustrates some of the modules of the HMD 100. FIG. 3 includes a schematic diagram of the operative modules of the HMD 100. For the case of speech recognition processing, controller 9100 accesses speech-to-text module 9036, which can be located locally to each HMD 100 or located remotely at a host 200 (FIG. 1A). Speech-to-text software module 9036 contains instructions to display to a user an image of processed text (e.g. dictation transcription) and menus (or navigation and other prompts). The graphics converter module 9040 converts the image instructions received from the module 9036 via bus 9103 and converts the instructions into graphics to display on the monocular display 9010. At the same time text-to-speech module 9035b converts instructions received from speech-to-text software module 9036 to create sounds representing the contents for the image to be displayed. The instructions are converted into digital sounds representing the corresponding image contents that the text-to-speech module 9035b feeds to the digital-to-analog converter 9021b, which in turn feeds speaker 9006 to present the audio to the user. Speech processing software module 9036 can be stored locally at memory 9120 or remotely at a host 200 (FIG. 1A). The user can speak/utter dictation and/or command selection and the user's speech 9090 is received at microphone 9020. The received speech is then converted from an analog signal into a digital signal at analog-to-digital converter 9021a. Once the speech is converted from an analog to a digital signal speech recognition module 9035a processes the speech into recognized speech. The recognized speech is compared against known speech and processed into text according to instructions of speech-to-text module 9036”); receiving audio input by the computing device (See Parkinson: Figs. 1A-B, and [0033], “Example embodiments of the HMD 100 can receive user input through sensing voice commands, head movements, 110, 111, 112 and hand gestures 113, or any combination thereof. Microphone(s) operatively coupled or preferably integrated into the HMD 100 can be used to capture speech commands which are then digitized and processed using automatic speech recognition techniques. Gyroscopes, accelerometers, and other micro-electromechanical system sensors can be integrated into the HMD 100 and used to track the user's head movement 110, 111, 112 to provide user input commands. Cameras or other motion tracking sensors can be used to monitor a user's hand gestures 113 for user input commands. Such a user interface overcomes the hands-dependent formats of other mobile devices”); and sending the audio input to the XR wearable device (See Parkinson: Figs. 1A-B, and [0036], “A head worn frame 1000 and strap 1002 are generally configured so that a user can wear the HMD 100 on the user's head. A housing 1004 is generally a low profile unit which houses the electronics, such as the microprocessor, memory or other storage device, along with other associated circuitry. Speakers 1006 provide audio output to the user so that the user can hear information. Micro-display subassembly 1010 is used to render visual information to the user. It is coupled to the arm 1008. The arm 1008 generally provides physical support such that the micro-display subassembly is able to be positioned within the user's field of view 300 (FIG. 1A), preferably in front of the eye of the user or within its peripheral vision preferably slightly below or above the eye. Arm 1008 also provides the electrical or optical connections between the micro-display subassembly 1010 and the control circuitry housed within housing unit 1004”). Regarding claim 5, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 4 as outlined above. Further, Gavriliuc teaches that the computing device of claim 4, wherein the user interface further comprises an indication of an edit command, and wherein the operations further comprise: receiving an indication of a selection, by the user, of the edit command while the image is being displayed on the display of the XR wearable device (See Gavriliuc: Figs. 4A-B, and [0026], “Some implementations may allow a recipient of markup to interact with and/or manipulate the markup, and to modify display of the markup based upon the interaction and/or manipulation. For example, FIG. 4A shows the user 104 moving the holographic object 302 to a different location within the environment 100, e.g. to suggest a possible different place to put a vase. In this instance, updated three-dimensional positional information for the markup may be sent, to computing device 106 along with image data capturing environment 100 from a current perspective, so that the markup may be displayed as drawing 202 in the updated location on computing device 106. The user 104 also may enter additional markup via computing device 102 to send to computing device 106. FIG. 4B shows the user 104 inserting additional markup 504 in the form of an arrow indicating another possible location to which to move the picture. Hand 502 in FIGS. 4A and 4B may represent either a cursor displayed on the see-through display of computing device 102, or an actual hand of the user 104 making hand gestures. A cursor may be controlled, for example, by user inputs using motion sensors (e.g. via head gestures), image sensors (e.g. inward facing image sensors to detect eye gestures or outward facing image sensors to detect hand gestures), touch sensors (e.g. a touch sensor incorporated into a portion of the computing device 104), external input devices (e.g. a mouse, touch sensor, or other position signal controller in communication with the computing device 102), microphones (e.g. via speech inputs), or in any other suitable manner”). Regarding claim 6, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. Further, Gauglitz, Parkinson and Gavriliuc teach that the computing device of claim 1, wherein the image is a first image, the plurality of 3D coordinates is a first plurality of 3D coordinates, the augmentation is a first augmentation, the 3D coordinates are first 3D coordinates, and wherein the operations (See Gauglitz: Fig. 1-3, and [0037], “Support for spatial references to the remote scene in video-mediated collaboration or telepresence has been an active research topic. Notable early works include “VideoDraw” and the “DoubleDigitalDesk.” Modalities that have been investigated include remote pointers, drawing onto a live video stream, and transferring videos of hand gestures. These annotations are then displayed to the collaborator on a separate screen in a third-person perspective, on a head-worn display, or via projectors”; and [0266], “In Example 73, the subject matter of any one or more of Examples 54-72 optionally include wherein the rendering component is further configured to select a target perspective from a set of acceptable target perspectives in response to receiving the image capture device control input”. Note that the live video stream includes plurality of frames, and each frame may be selected for collaborating and annotation. Repeating the same steps are not inventive features) further comprise: accessing a second image captured by the XR wearable device and an identification of the second image, the image being receive from the XR wearable device (See Gauglitz: Fig. 9, and [0116], "FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings"; and [0136], “In contrast, touchscreens are not only ubiquitous today, but they afford direct interaction without the need for an intermediate representation (i.e., a mouse cursor), and provide haptic feedback during the touch. In this context, however, they have the downside of providing 2D input only. Discussing the implications of this limitation and describing and evaluating appropriate solutions in the context of live remote collaboration is one of the main contributions of this application”. Note that the freehand drawing as the 2D input to the system is mapped to the indication of augmentation generated by the user. Repeating the same steps are not inventive features); sending, to the XR wearable device, the identification of the second image (See Gavriliuc: Fig. 1, and [0019], “The display device may send captured image data, and also depth data representing environment 100, to a remote display device 106 for presentation to user 108. User 108 may input markup to the image data, for example by entering a handmade sketch via touch input, by inserting an existing drawing or image, etc. using the remote display device 106. The remote display device 106 associates the markup with a three-dimensional location within the environment based upon the location in the image data at which the markup is made and then may send the markup back to the display device 102 for presentation to user 104 via augmented reality display device 102. Via the markup, user 108 may provide instructions, questions, comments, and/or other information to user 106 in the context of the image data. Further, since the markup is associated with a three-dimensional location in a coordinate frame of the augmented reality environment, user 104 may view the markup as spatially augment reality imagery, such that the markup remains in a selected orientation and location relative to environment 100. This may allow user 104 to view the markup from different perspectives by moving within the environment. It will be understood that, in other implementations, any other suitable type of computing device than a wearable augmented reality display device may be used. For example, a video-based augmentation mode that combines a camera viewfinder view with the markup may be used with any suitable display device comprising a camera to present markup-augmented imagery according to the present disclosure”; and [0023], “Each input of markup may be associated with a three-dimensional location in the environment that is mapped to the two-dimensional (i.e. in the plane of the image) location in the image at which the markup was made. For example, in FIG. 2B, the user 108 input the drawing 202 at a location corresponding to the top of a table, and the drawing 204 at a location corresponding to a wall. Thus, the drawings 202, 204 may be associated with the depth locations in the real world scene that are mapped to those locations in the image displayed on display device 106. After receiving input of the markup, computing device 106 may send the markup and associated three-dimensional location of each item of markup to display device 102 for presentation to user 104, to a server for storage and later retrieval, and/or to any other suitable device”. Repeating the same steps are not inventive features); receiving, from the XR wearable device, a second plurality of 3D coordinates corresponding to a plurality of positions within the second image (See Gauglitz: Figs. 2 and 10, and [0043], “Under the hood, the system runs a SLAM system and sends the tracked camera pose along with the encoded live video stream to the remote system. The local user's system receives information about annotations from the remote system and uses this information together with the live video to render the augmented view”; [0050], “A 3D surface model is constructed on the fly from the live video stream and from associated camera poses. Keyframes were selected based on a set of heuristics (good tracking quality, low device movement, minimum time interval & translational distance between keyframes), then detect and describe features in the new frame using SIFT. Four closest existing keyframes were chosen and matched against their features (one frame at a time) via an approximate nearest neighbor algorithm and collect matches that satisfy the epipolar constraint (which is known due to the received camera poses) within some tolerance as tentative 3D points. If a feature has previously been matched to features from other frames, we check for mutual epipolar consistency of all observations and merge them into a single 3D point if possible; otherwise, the two 3D points remain as competing hypotheses”; and [0118], " An operational mobile system may be used to implement this idea. The remote user's system received the live video and computer vision-based tracking information from a local user's smartphone or tablet and modeled the environment in 3D based on this data. It enabled the remote user to navigate the scene and to create annotations in it that were then sent back and visualized to the local user in AR. This earlier work focused on the overall system and enabling the navigation and communication within this framework. However, the remote user's interface was rather simple: most notably, it used a standard mouse for most interactions, supplemented by keyboard shortcuts for certain functions, and supported only single-point-based markers". Note that 3D pose of the cameras is mapped to the 3D coordinates corresponding to a plurality of positions within the image, and the keyframes is mapped to the identification of the images. Repeating the same steps are not inventive features); displaying the second image on the display of the computing device (See Gauglitz: Figs. 1,3, 9, and 11, and [0045], "The system operates at 30 frames per second. System latencies were measured using a camera with 1000 fps, which observed a change in the physical world (a falling object passing a certain height) as well as its image on the respective screen. The latency between physical effect and the local user's tablet display—including image formation on the sensor, retrieval, processing by the SLAM system, rendering of the image, and display on the screen—was measured as 205±22.6 ms". Note that repeating the same steps are not inventive features); accessing an indication of a second augmentation for the second image, the second augmentation generated by the computing device in response to user input from the user (See Gauglitz: Fig. 9, and [0116], "FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings"; and [0136], “In contrast, touchscreens are not only ubiquitous today, but they afford direct interaction without the need for an intermediate representation (i.e., a mouse cursor), and provide haptic feedback during the touch. In this context, however, they have the downside of providing 2D input only. Discussing the implications of this limitation and describing and evaluating appropriate solutions in the context of live remote collaboration is one of the main contributions of this application”. Note that the freehand drawing as the 2D input to the system is mapped to the indication of augmentation generated by the user. Repeating the same steps are not inventive features); determining second 3D coordinates of the second augmentation based on the second plurality of 3D coordinates (See Gauglitz: Figs. 9-12, and [0116], “FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings”; [0145], "Given an individual 2D input location p=(x, y), a depth d can thus be obtained by un-projecting p onto the model of the scene. For a sequence of 2D inputs (a 2D drawing) p.sub.1, . . . , p.sub.n, this results in several alternatives to interpret the depth of the drawing as a whole, of which we consider the following"; and [0178], “FIG. 18 shows a screenshot during an orbit operation 1800, according to an embodiment. As a third gesture, we added orbiting around a given world point p.sup.3D. The corresponding gesture is to keep one finger (relatively) static at a point p and move the second finger in an arc around it (FIG. 15(e)). The orbit center p.sup.3D is determined by un-projecting p onto the scene and remains fixed during the movement; additionally, we maintain the gravity vector (as reported by the local user's device). The rotation around the gravity vector at p.sup.3D is then specified by the movement of the second finger”. Note that 2D screen annotations are un-projected into 3Dworld coordinates using received 3D model, camera pose and depth information so the annotations become world-stabilized, sine the camera is tracked with respect to the scene, these annotations automatically obtain a position in world coordinates, and depth d can thus be obtained by un-projecting p onto the model of the scene. Repeating the same steps are not inventive features); and sending the indication of the second augmentation and the second 3D coordinates of the second augmentation to the XR wearable device (See Gavriliuc: Fig. 1, and [0019], “The display device may send captured image data, and also depth data representing environment 100, to a remote display device 106 for presentation to user 108. User 108 may input markup to the image data, for example by entering a handmade sketch via touch input, by inserting an existing drawing or image, etc. using the remote display device 106. The remote display device 106 associates the markup with a three-dimensional location within the environment based upon the location in the image data at which the markup is made and then may send the markup back to the display device 102 for presentation to user 104 via augmented reality display device 102. Via the markup, user 108 may provide instructions, questions, comments, and/or other information to user 106 in the context of the image data. Further, since the markup is associated with a three-dimensional location in a coordinate frame of the augmented reality environment, user 104 may view the markup as spatially augment reality imagery, such that the markup remains in a selected orientation and location relative to environment 100. This may allow user 104 to view the markup from different perspectives by moving within the environment. It will be understood that, in other implementations, any other suitable type of computing device than a wearable augmented reality display device may be used. For example, a video-based augmentation mode that combines a camera viewfinder view with the markup may be used with any suitable display device comprising a camera to present markup-augmented imagery according to the present disclosure”; and [0023], “Each input of markup may be associated with a three-dimensional location in the environment that is mapped to the two-dimensional (i.e. in the plane of the image) location in the image at which the markup was made. For example, in FIG. 2B, the user 108 input the drawing 202 at a location corresponding to the top of a table, and the drawing 204 at a location corresponding to a wall. Thus, the drawings 202, 204 may be associated with the depth locations in the real world scene that are mapped to those locations in the image displayed on display device 106. After receiving input of the markup, computing device 106 may send the markup and associated three-dimensional location of each item of markup to display device 102 for presentation to user 104, to a server for storage and later retrieval, and/or to any other suitable device”. Note that repeating the same steps are not inventive features). Regarding claim 7, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 6 as outlined above. Further, Parkinson teaches that the computing device of claim 6, wherein the sending further comprises sending an indication of a command that the XR wearable device is to return the second plurality of 3D coordinates (See Parkinson: Fig. 1, and [0033], “Example embodiments of the HMD 100 can receive user input through sensing voice commands, head movements, 110, 111, 112 and hand gestures 113, or any combination thereof. Microphone(s) operatively coupled or preferably integrated into the HMD 100 can be used to capture speech commands which are then digitized and processed using automatic speech recognition techniques. Gyroscopes, accelerometers, and other micro-electromechanical system sensors can be integrated into the HMD 100 and used to track the user's head movement 110, 111, 112 to provide user input commands. Cameras or other motion tracking sensors can be used to monitor a user's hand gestures 113 for user input commands. Such a user interface overcomes the hands-dependent formats of other mobile devices”; and [0050], “The first application at the host 200 initiates communication (also referred to herein as a `call`) between the host 200 and the HMD 100 when the first application receives location information associated with the HMD 100. For example, a remote user at the host 200 may enter an IP address associated with the HMD 100 at the host 200. The location information associated with the HMD 100 may also be accompanied by an initiation command (e.g., the remote user may press a `send` button after entering the IP address)”. Note that extension the bidirectional communication between the remote host and the local HMD to include a command that requests spatial and coordinate data for the specific image is natural and obvious). Regarding claim 8, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. Further, Gauglitz teaches that the computing device of claim 1, wherein the augmentation to the image comprises a geometric shape or a line drawn based on user input (See Gauglitz: Fig. 1, and [0057], “Similarly, if the packet update contains information about coordinates, line width and color, then a line will be drawn on the displayed image based on the coordinates, the color and width information”; and Fig. 9, and [0116], “FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings”). Regarding claim 10, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. Further, Gauglitz teaches that the computing device of claim 1, wherein the plurality of 3D coordinates are 3D world coordinates in a frame of reference of the XR wearable device (See Gauglitz: Figa. 9-12, and [0064], “In addition to being able to control the viewpoint, the remote user can set and remove virtual annotations. Annotations are saved in 3D world coordinates, are shared with the local user's mobile device via the network, and immediately appear in all views of the world correctly anchored to their 3D world position (cf. FIGS. 1 and 3)”; [0121], “Existing videoconferencing or telepresence systems lack the ability to interact with the remote physical environment. Researchers have explored various methods to support spatial references to the remote scene, including pointers, hand gestures, and drawings. The use of drawings for collaboration has been investigated, including recognition of common shapes (for regularization, compression, and interpretation as commands). However, in all of these works, the remote user's view onto the scene is constrained to the current view of the local user's camera, and the support for spatially referencing the scene is contingent upon a stationary camera. More specifically, when drawings are used, they are created on a 2D surface as well as displayed in a 2D space (e.g., a live video), and it remains up to the user to mentally “un-project” them and interpret them in 3D. This is fundamentally different from creating annotations that are anchored and displayed in 3D space, that is, in AR”; [0122], “Some systems support world-stabilized annotations. Of those systems, some support only single-point-based markers, and some can cope with 2D/panorama scenes only. Some systems use active depth sensors on both the local and the remote user's side and thus support transmission of hand gestures in 3D (along with a different approach to navigating the remote scene); we will contrast their approach with ours in more detail”; and [0142], “Using a single finger, the remote user can draw annotations into the scene. A few examples are shown in FIG. 12(a). Since the camera is tracked with respect to the scene, these annotations automatically obtain a position in world coordinates. However, due to the use of a touchscreen, the depth of the drawing along the current viewpoint's optical axis is unspecified”). Regarding claim 12, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. Further, Gauglitz teaches that the computing device of claim 1, wherein the plurality of 3D coordinates is a point cloud of 3D coordinates (See Gauglitz: Fig. 1, and [0052], “Several steps are required to obtain a surface model from the 3D point cloud. First, a Delaunay tetrahedralization of the point cloud is created. Each tetrahedron is then labeled as “free” or “occupied,” and the interface between free and occupied tetrahedra is extracted as the scene's surface. The labeling of the tetrahedra works as follows: A graph structure is created in which each tetrahedron is represented by a node, and nodes of neighboring tetrahedra are linked by edges. Each node is further linked to a “free” (sink) node and an “occupied” (source) node. The weights of all edges depend on the observations that formed each vertex; for example, an observation ray that cuts through a cell indicates that this cell is free, while a ray ending in front of a cell indicates that the cell is occupied. Finally, the labels for all tetrahedra are determined by solving a dynamic graph cut problem. This modeling is refined further by taking the orientation of observation rays to cell interfaces into account, which reduces the number of “weak” links and thus the risk that the graph cut finds a minimum that does not correspond to a true surface”). Regarding claim 13, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. Further, Gauglitz, Parkinson and Gavriliuc teach that an apparatus of an extended reality (XR) wearable device (See Gauglitz: Fig. 20, and [0190], "FIG. 20 is a block diagram of a computing device 2000, according to an embodiment. In one embodiment, multiple such computer systems are used in a distributed network to implement multiple components in a transaction-based environment. An object-oriented, service-oriented, or other architecture may be used to implement such functions and communicate between the multiple systems and components. In some embodiments, the computing device of FIG. 20 is an example of a client device that may invoke methods described herein over a network. In other embodiments, the computing device is an example of a computing device that may be included in or connected to a motion interactive video projection system, as described elsewhere herein. In some embodiments, the computing device of FIG. 20 is an example of one or more of the personal computer, smartphone, tablet, or various servers") comprising: at least one processor (See Gauglitz: Fig. 20, and [0191], "One example computing device in the form of a computer 2010, may include a processing unit 2002, memory 2004, removable storage 2012, and non-removable storage 2014. Although the example computing device is illustrated and described as computer 2010, the computing device may be in different forms in different embodiments. For example, the computing device may instead be a smartphone, a tablet, or other computing device including the same or similar elements as illustrated and described with regard to FIG. 20. Further, although the various data storage elements are illustrated as part of the computer 2010, the storage may include cloud-based storage accessible via a network, such as the Internet"); and a memory storing instructions that, when executed by the at least one processor, configure the at least one processor to perform operations (See Gauglitz: Fig. 20, and [0191], "One example computing device in the form of a computer 2010, may include a processing unit 2002, memory 2004, removable storage 2012, and non-removable storage 2014. Although the example computing device is illustrated and described as computer 2010, the computing device may be in different forms in different embodiments. For example, the computing device may instead be a smartphone, a tablet, or other computing device including the same or similar elements as illustrated and described with regard to FIG. 20. Further, although the various data storage elements are illustrated as part of the computer 2010, the storage may include cloud-based storage accessible via a network, such as the Internet"; and [0193], “Computer-readable instructions stored on a computer-readable medium are executable by the processing unit 2002 of the computer 2010. A hard drive (magnetic disk or solid state), CD-ROM, and RAM are some examples of articles including a non-transitory computer-readable medium. For example, various computer programs 2025 or apps, such as one or more applications and modules implementing one or more of the methods illustrated and described herein or an app or application that executes on a mobile device or is accessible via a web browser, may be stored on a non-transitory computer-readable medium”) comprising: capturing, by an image capturing device of the XR wearable device, an image corresponding to a first user view of a real-world scene, the image associated with an identification (See Gauglitz: Figs. 2 and 10, and [0043], “Under the hood, the system runs a SLAM system and sends the tracked camera pose along with the encoded live video stream to the remote system. The local user's system receives information about annotations from the remote system and uses this information together with the live video to render the augmented view”; [0050], “A 3D surface model is constructed on the fly from the live video stream and from associated camera poses. Keyframes were selected based on a set of heuristics (good tracking quality, low device movement, minimum time interval & translational distance between keyframes), then detect and describe features in the new frame using SIFT. Four closest existing keyframes were chosen and matched against their features (one frame at a time) via an approximate nearest neighbor algorithm and collect matches that satisfy the epipolar constraint (which is known due to the received camera poses) within some tolerance as tentative 3D points. If a feature has previously been matched to features from other frames, we check for mutual epipolar consistency of all observations and merge them into a single 3D point if possible; otherwise, the two 3D points remain as competing hypotheses”; and [0118], " An operational mobile system may be used to implement this idea. The remote user's system received the live video and computer vision-based tracking information from a local user's smartphone or tablet and modeled the environment in 3D based on this data. It enabled the remote user to navigate the scene and to create annotations in it that were then sent back and visualized to the local user in AR. This earlier work focused on the overall system and enabling the navigation and communication within this framework. However, the remote user's interface was rather simple: most notably, it used a standard mouse for most interactions, supplemented by keyboard shortcuts for certain functions, and supported only single-point-based markers". Note that 3D pose of the cameras is mapped to the 3D coordinates corresponding to a plurality of positions within the image, and the keyframes is mapped to the identification of the images), the identification identifying the image (See Parkinson: Fig. 1, and [0015], “In another aspect, the invention may be a computer-assisted method of remote document annotation, including selecting, at a host computing platform, a document to be annotated, and providing, at the host computing platform, one or more annotations to the document. The method may further include submitting, at the host computing platform, a location identifier of a desired recipient of the document, and transmitting, by the host computing platform, information representative of the document and information representative of the one or more annotations, to a head mounted display device associated with the location identifier. The method may further include receiving, by the head mounted display device, the information representative of the document and the information representative of the one or more annotations. The method may further include applying, by the head mounted display device, the information representative of the one or more annotations to the document, so as to recreate the one or more annotations provided at the host computing platform. The method may further include displaying, by the head mounted display device, the document together with the annotations”); determining a plurality of 3-dimensional (3D) coordinates corresponding to a plurality of positions within the image(See Gauglitz: Figs. 9-12, and [0116], “FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings”; [0145], "Given an individual 2D input location p=(x, y), a depth d can thus be obtained by un-projecting p onto the model of the scene. For a sequence of 2D inputs (a 2D drawing) p.sub.1, . . . , p.sub.n, this results in several alternatives to interpret the depth of the drawing as a whole, of which we consider the following"; and [0178], “FIG. 18 shows a screenshot during an orbit operation 1800, according to an embodiment. As a third gesture, we added orbiting around a given world point p.sup.3D. The corresponding gesture is to keep one finger (relatively) static at a point p and move the second finger in an arc around it (FIG. 15(e)). The orbit center p.sup.3D is determined by un-projecting p onto the scene and remains fixed during the movement; additionally, we maintain the gravity vector (as reported by the local user's device). The rotation around the gravity vector at p.sup.3D is then specified by the movement of the second finger”. Note that 2D screen annotations are un-projected into 3Dworld coordinates using received 3D model, camera pose and depth information so the annotations become world-stabilized, sine the camera is tracked with respect to the scene, these annotations automatically obtain a position in world coordinates, and depth d can thus be obtained by un-projecting p onto the model of the scene); causing the image, the plurality of 3D coordinates, and the identification to be sent to a computing device (See Gauglitz: Figs. 2 and 10, and [0026], "This AR solution enables collaboration between two collaborators in different physical locations. As described herein, a first collaborator is described as being located in physical location A, and may be referred to as a “local user” or “local collaborator.” Similarly, a second collaborator is described as being located in physical location B, and may be referred to as a “remote user” or “remote collaborator.” This AR solution enables the remote collaborator to control his or her viewpoint of physical location A, such as by adjusting camera controls or other viewpoint adjustments. This AR solution enables the remote collaborator to communicate information with visual or spatial reference to physical objects in physical location A, where the information is sent from the remote collaborator to the collaborator in this physical location A. The visual and/or spatial reference information may include identifying objects, locations, directions, or spatial instructions by creating annotations, where the annotations are transmitted to the collaborator in physical location A and visualized using appropriate display technology. For example, the AR solution may include a visual display for the collaborator in physical location A, where the annotations are overlaid onto and anchored to the respective real world object or location"; [0076], ‘Currently, the 3D model is available only on the remote user's side. Thus, annotations can be correctly occluded by the physical scene by the remote user's renderer, but not by the local user's renderer. While other depth cues (most notably, parallax) still indicate the annotation's location, it would be desirable to enable occlusion. The remote system can send either the model geometry or alternatively local visibility information per annotation back to the local device in order to enable occlusion on the local user's side”; and [0127], “The remote user's system may be run on one of a variety of devices, such as a commodity PC or laptop, a cell phone, tablet, virtual reality device, or other device. The embodiment shown in FIG. 10 was generated using a commodity PC. It receives the live video stream and the associated camera poses and models the environment in real time from this data. The remote user can set world-anchored annotations, which are sent back and displayed, to the local user. Further, thanks to the constructed 3D model, he/she can move away from the live view and choose to look at the scene from a different viewpoint via a set of virtual navigation controls. Transitions from one viewpoint to another (including the live view) are rendered seamlessly via image based rendering techniques using the 3D model, cached keyframes, and the live frame appropriately”); and receiving, from the computing device, an indication of an augmentation and 3D coordinates associated with the augmentation (See Gavriliuc: Fig. 1, and [0019], “The display device may send captured image data, and also depth data representing environment 100, to a remote display device 106 for presentation to user 108. User 108 may input markup to the image data, for example by entering a handmade sketch via touch input, by inserting an existing drawing or image, etc. using the remote display device 106. The remote display device 106 associates the markup with a three-dimensional location within the environment based upon the location in the image data at which the markup is made and then may send the markup back to the display device 102 for presentation to user 104 via augmented reality display device 102. Via the markup, user 108 may provide instructions, questions, comments, and/or other information to user 106 in the context of the image data. Further, since the markup is associated with a three-dimensional location in a coordinate frame of the augmented reality environment, user 104 may view the markup as spatially augment reality imagery, such that the markup remains in a selected orientation and location relative to environment 100. This may allow user 104 to view the markup from different perspectives by moving within the environment. It will be understood that, in other implementations, any other suitable type of computing device than a wearable augmented reality display device may be used. For example, a video-based augmentation mode that combines a camera viewfinder view with the markup may be used with any suitable display device comprising a camera to present markup-augmented imagery according to the present disclosure”; and [0023], “Each input of markup may be associated with a three-dimensional location in the environment that is mapped to the two-dimensional (i.e. in the plane of the image) location in the image at which the markup was made. For example, in FIG. 2B, the user 108 input the drawing 202 at a location corresponding to the top of a table, and the drawing 204 at a location corresponding to a wall. Thus, the drawings 202, 204 may be associated with the depth locations in the real world scene that are mapped to those locations in the image displayed on display device 106. After receiving input of the markup, computing device 106 may send the markup and associated three-dimensional location of each item of markup to display device 102 for presentation to user 104, to a server for storage and later retrieval, and/or to any other suitable device”). Regarding claim 14, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 13 as outlined above. Further, Gavriliuc teaches that the XR wearable device of claim 13, wherein the identification of the image is an indication of a Coordinated Universal Time (UTC) or an indication of an image identification number (See Gavriliuc: Fig. 1, and [0020], “In some implementations, depth image data may be acquired via a depth camera integrated with the augmented reality display device 102. In such implementations, the fields of view of the depth camera and RGB camera may have a calibrated or otherwise known spatial relationship. This may facilitate integrating or otherwise associating each frame of RBG data with corresponding depth image data”; and Fig. 6, and [0030], “Method 600 further includes, at 608, sending the image data and depth data to device B. Device B receives the image data and depth data at 610, and displays the received image data at 612. At 614, method 600 comprises receiving one or more input(s) of markup to the image data. Any suitable input of markup may be received. For example, as described above, the input of markup may comprise an input of a drawing made by touch or other suitable input. The input of markup further may comprise an input of an image, video, animated item, executable item, text, and/or any other suitable content. At 616, method 600 includes associating each item of markup with a three-dimensional location in the real world scene. Each markup and the associated three-dimensional location are then sent back to device A, as shown at 618. As described above, the data shared between devices may also include an identifier of the real world scene. Such an identifier may facilitate storage and later retrieval of the markup, so that other devices can obtain the markup based upon the identifier when the devices are at the identified location. It will be appreciated that the identifier may be omitted in some implementations”. Note that real-time image and depth streams are conventionally timestamped or numbered so that the correct frame’s depth data is used for the 3D association. Thus, UTC or an image ID number (timestamps) is an obvious species of such an identifier). Regarding claim 15, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 13 as outlined above. Further, Gauglitz teaches that the XR wearable device of claim 13, wherein the operations further comprise: displaying, on a display of the XR wearable device, the augmentation, wherein a position of the augmentation is based on the 3D coordinates associated with the augmentation and a second user view of the real-world scene (See Gauglitz: Fig. 3, and [0055], “FIG. 3 shows a screenshot of the remote user's interface 300, according to an embodiment. Providing camera control (virtual navigation), or providing a remote user's ability to navigate the remote world via a virtual camera, independent of the local user's current location, is among the important contributions of this work. Some solutions use physical device movement for navigation. While this is arguably intuitive, using physical navigation has two disadvantages; one being generally true for physical navigation and the other one being specific to the application of live collaboration: First, the remote user needs to be able to physically move and track his/her movements in a space corresponding in size to the remote environment of interest, and “supernatural” movements or viewpoints (e.g., quickly covering large distances or adopting a bird's-eye view) are impossible. Second, it does not allow coupling of the remote user's view to the local user's view (and thus have the local user control the viewpoint) without breaking the frame of reference in which the remote user navigates. This application describes some embodiments using virtual navigation, and it is important that the navigation gives the remote user the option of coupling his/her view to that of the local user. The systems and methods described herein may also be applied to physical navigation”). Regarding claim 17, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 13 as outlined above. Further, Gauglitz teaches that the XR wearable device of claim 13, wherein the plurality of 3D coordinates are a point cloud of 3D coordinates (See Gauglitz: Fig. 1, and [0052], “Several steps are required to obtain a surface model from the 3D point cloud. First, a Delaunay tetrahedralization of the point cloud is created. Each tetrahedron is then labeled as “free” or “occupied,” and the interface between free and occupied tetrahedra is extracted as the scene's surface. The labeling of the tetrahedra works as follows: A graph structure is created in which each tetrahedron is represented by a node, and nodes of neighboring tetrahedra are linked by edges. Each node is further linked to a “free” (sink) node and an “occupied” (source) node. The weights of all edges depend on the observations that formed each vertex; for example, an observation ray that cuts through a cell indicates that this cell is free, while a ray ending in front of a cell indicates that the cell is occupied. Finally, the labels for all tetrahedra are determined by solving a dynamic graph cut problem. This modeling is refined further by taking the orientation of observation rays to cell interfaces into account, which reduces the number of “weak” links and thus the risk that the graph cut finds a minimum that does not correspond to a true surface”). Regarding claim 19, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. Further, Gauglitz, Parkinson and Gavriliuc teach that the method comprising: receiving an image captured by an extended reality (XR) wearable device, a plurality of 3-dimensional (3D) coordinates corresponding to a plurality of positions within the image (See Gauglitz: Figs. 2 and 10, and [0043], “Under the hood, the system runs a SLAM system and sends the tracked camera pose along with the encoded live video stream to the remote system. The local user's system receives information about annotations from the remote system and uses this information together with the live video to render the augmented view”; [0050], “A 3D surface model is constructed on the fly from the live video stream and from associated camera poses. Keyframes were selected based on a set of heuristics (good tracking quality, low device movement, minimum time interval & translational distance between keyframes), then detect and describe features in the new frame using SIFT. Four closest existing keyframes were chosen and matched against their features (one frame at a time) via an approximate nearest neighbor algorithm and collect matches that satisfy the epipolar constraint (which is known due to the received camera poses) within some tolerance as tentative 3D points. If a feature has previously been matched to features from other frames, we check for mutual epipolar consistency of all observations and merge them into a single 3D point if possible; otherwise, the two 3D points remain as competing hypotheses”; and [0118], " An operational mobile system may be used to implement this idea. The remote user's system received the live video and computer vision-based tracking information from a local user's smartphone or tablet and modeled the environment in 3D based on this data. It enabled the remote user to navigate the scene and to create annotations in it that were then sent back and visualized to the local user in AR. This earlier work focused on the overall system and enabling the navigation and communication within this framework. However, the remote user's interface was rather simple: most notably, it used a standard mouse for most interactions, supplemented by keyboard shortcuts for certain functions, and supported only single-point-based markers". Note that 3D pose of the cameras is mapped to the 3D coordinates corresponding to a plurality of positions within the image, and the keyframes is mapped to the identification of the images), and an identification of the image (See Parkinson: Fig. 1, and [0015], “In another aspect, the invention may be a computer-assisted method of remote document annotation, including selecting, at a host computing platform, a document to be annotated, and providing, at the host computing platform, one or more annotations to the document. The method may further include submitting, at the host computing platform, a location identifier of a desired recipient of the document, and transmitting, by the host computing platform, information representative of the document and information representative of the one or more annotations, to a head mounted display device associated with the location identifier. The method may further include receiving, by the head mounted display device, the information representative of the document and the information representative of the one or more annotations. The method may further include applying, by the head mounted display device, the information representative of the one or more annotations to the document, so as to recreate the one or more annotations provided at the host computing platform. The method may further include displaying, by the head mounted display device, the document together with the annotations”); displaying the image on a display of a computing device (See Gauglitz: Figs. 1,3, 9, and 11, and [0045], "The system operates at 30 frames per second. System latencies were measured using a camera with 1000 fps, which observed a change in the physical world (a falling object passing a certain height) as well as its image on the respective screen. The latency between physical effect and the local user's tablet display—including image formation on the sensor, retrieval, processing by the SLAM system, rendering of the image, and display on the screen—was measured as 205±22.6 ms"); accessing an indication of an augmentation to the image, the augmentation generated by a user of the computing device (See Gauglitz: Fig. 9, and [0116], "FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings"; and [0136], “In contrast, touchscreens are not only ubiquitous today, but they afford direct interaction without the need for an intermediate representation (i.e., a mouse cursor), and provide haptic feedback during the touch. In this context, however, they have the downside of providing 2D input only. Discussing the implications of this limitation and describing and evaluating appropriate solutions in the context of live remote collaboration is one of the main contributions of this application”. Note that the freehand drawing as the 2D input to the system is mapped to the indication of augmentation generated by the user); determining 3D coordinates of the augmentation based on the plurality of 3D coordinates (See Gauglitz: Figs. 9-12, and [0116], “FIG. 9 shows augmented reality annotations and virtual scene navigation 900, according to an embodiment. Augmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. This portion of the application presents a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. These elements are evaluated via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings”; [0145], "Given an individual 2D input location p=(x, y), a depth d can thus be obtained by un-projecting p onto the model of the scene. For a sequence of 2D inputs (a 2D drawing) p.sub.1, . . . , p.sub.n, this results in several alternatives to interpret the depth of the drawing as a whole, of which we consider the following"; and [0178], “FIG. 18 shows a screenshot during an orbit operation 1800, according to an embodiment. As a third gesture, we added orbiting around a given world point p.sup.3D. The corresponding gesture is to keep one finger (relatively) static at a point p and move the second finger in an arc around it (FIG. 15(e)). The orbit center p.sup.3D is determined by un-projecting p onto the scene and remains fixed during the movement; additionally, we maintain the gravity vector (as reported by the local user's device). The rotation around the gravity vector at p.sup.3D is then specified by the movement of the second finger”. Note that 2D screen annotations are un-projected into 3Dworld coordinates using received 3D model, camera pose and depth information so the annotations become world-stabilized, sine the camera is tracked with respect to the scene, these annotations automatically obtain a position in world coordinates, and depth d can thus be obtained by un-projecting p onto the model of the scene.); and sending (See Gauglitz: Figs. 2 and 10, and [0026], "This AR solution enables collaboration between two collaborators in different physical locations. As described herein, a first collaborator is described as being located in physical location A, and may be referred to as a “local user” or “local collaborator.” Similarly, a second collaborator is described as being located in physical location B, and may be referred to as a “remote user” or “remote collaborator.” This AR solution enables the remote collaborator to control his or her viewpoint of physical location A, such as by adjusting camera controls or other viewpoint adjustments. This AR solution enables the remote collaborator to communicate information with visual or spatial reference to physical objects in physical location A, where the information is sent from the remote collaborator to the collaborator in this physical location A. The visual and/or spatial reference information may include identifying objects, locations, directions, or spatial instructions by creating annotations, where the annotations are transmitted to the collaborator in physical location A and visualized using appropriate display technology. For example, the AR solution may include a visual display for the collaborator in physical location A, where the annotations are overlaid onto and anchored to the respective real world object or location"; [0076], ‘Currently, the 3D model is available only on the remote user's side. Thus, annotations can be correctly occluded by the physical scene by the remote user's renderer, but not by the local user's renderer. While other depth cues (most notably, parallax) still indicate the annotation's location, it would be desirable to enable occlusion. The remote system can send either the model geometry or alternatively local visibility information per annotation back to the local device in order to enable occlusion on the local user's side”; and [0127], “The remote user's system may be run on one of a variety of devices, such as a commodity PC or laptop, a cell phone, tablet, virtual reality device, or other device. The embodiment shown in FIG. 10 was generated using a commodity PC. It receives the live video stream and the associated camera poses and models the environment in real time from this data. The remote user can set world-anchored annotations, which are sent back and displayed, to the local user. Further, thanks to the constructed 3D model, he/she can move away from the live view and choose to look at the scene from a different viewpoint via a set of virtual navigation controls. Transitions from one viewpoint to another (including the live view) are rendered seamlessly via image based rendering techniques using the 3D model, cached keyframes, and the live frame appropriately”) the identification, the indication of the augmentation, and the 3D coordinates of the augmentation to the XR wearable device (See Gavriliuc: Fig. 1, and [0019], “The display device may send captured image data, and also depth data representing environment 100, to a remote display device 106 for presentation to user 108. User 108 may input markup to the image data, for example by entering a handmade sketch via touch input, by inserting an existing drawing or image, etc. using the remote display device 106. The remote display device 106 associates the markup with a three-dimensional location within the environment based upon the location in the image data at which the markup is made and then may send the markup back to the display device 102 for presentation to user 104 via augmented reality display device 102. Via the markup, user 108 may provide instructions, questions, comments, and/or other information to user 106 in the context of the image data. Further, since the markup is associated with a three-dimensional location in a coordinate frame of the augmented reality environment, user 104 may view the markup as spatially augment reality imagery, such that the markup remains in a selected orientation and location relative to environment 100. This may allow user 104 to view the markup from different perspectives by moving within the environment. It will be understood that, in other implementations, any other suitable type of computing device than a wearable augmented reality display device may be used. For example, a video-based augmentation mode that combines a camera viewfinder view with the markup may be used with any suitable display device comprising a camera to present markup-augmented imagery according to the present disclosure”; and [0023], “Each input of markup may be associated with a three-dimensional location in the environment that is mapped to the two-dimensional (i.e. in the plane of the image) location in the image at which the markup was made. For example, in FIG. 2B, the user 108 input the drawing 202 at a location corresponding to the top of a table, and the drawing 204 at a location corresponding to a wall. Thus, the drawings 202, 204 may be associated with the depth locations in the real world scene that are mapped to those locations in the image displayed on display device 106. After receiving input of the markup, computing device 106 may send the markup and associated three-dimensional location of each item of markup to display device 102 for presentation to user 104, to a server for storage and later retrieval, and/or to any other suitable device”). Regarding claim 20, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 19 as outlined above. Further, Gauglitz teaches that the method of claim 19, wherein the method further comprises: receiving a video stream from the XR wearable device, the video stream comprising a plurality of images comprising the image (See Gauglitz: Figs. 1-3, and [0050], “A 3D surface model is constructed on the fly from the live video stream and from associated camera poses. Keyframes were selected based on a set of heuristics (good tracking quality, low device movement, minimum time interval & translational distance between keyframes), then detect and describe features in the new frame using SIFT. Four closest existing keyframes were chosen and matched against their features (one frame at a time) via an approximate nearest neighbor algorithm and collect matches that satisfy the epipolar constraint (which is known due to the received camera poses) within some tolerance as tentative 3D points. If a feature has previously been matched to features from other frames, we check for mutual epipolar consistency of all observations and merge them into a single 3D point if possible; otherwise, the two 3D points remain as competing hypotheses”); displaying on the display the video stream and a user interface comprising indications of directions (See Gauglitz Fig. 3, and [0058], “The present subject matter provides the ability to freeze and return to live views. The application starts with the remote view coupled to the local user's live view. With a single right-click, the remote user “freezes” his/her camera at the current pose, for example in order to set annotations precisely, or as a starting point for further navigation. Whenever the remote user's view is not coupled to the local user's view, the latter is displayed to the remote user as an inset (cf. FIG. 3). A click onto this inset or pressing 0 immediately transitions back to the live view”; and [0062], “The present subject matter provides the ability to save and revisit viewpoints. Further, the user can actively save a particular viewpoint to revisit it later. Pressing Alt plus any number key saves the current viewpoint; pressing the respective number key alone revisits this view later. Small numbers along the top of the screen indicate which numbers are currently in use (see FIG. 3”. Note that the viewpoints is mapped to the directions); receiving a selection of a direction of the directions(See Gauglitz Fig. 3, and [0058], “The present subject matter provides the ability to freeze and return to live views. The application starts with the remote view coupled to the local user's live view. With a single right-click, the remote user “freezes” his/her camera at the current pose, for example in order to set annotations precisely, or as a starting point for further navigation. Whenever the remote user's view is not coupled to the local user's view, the latter is displayed to the remote user as an inset (cf. FIG. 3). A click onto this inset or pressing 0 immediately transitions back to the live view”; and [0062], “The present subject matter provides the ability to save and revisit viewpoints. Further, the user can actively save a particular viewpoint to revisit it later. Pressing Alt plus any number key saves the current viewpoint; pressing the respective number key alone revisits this view later. Small numbers along the top of the screen indicate which numbers are currently in use (see FIG. 3”; and [0059], “The present subject matter provides panning and zooming capabilities. By moving the mouse while its right button is pressed, the user can pan the view in a panorama-like fashion (rotate around the current camera position). To prevent the user from getting lost in unmapped areas, we constrain the panning to the angular extent of the modeled environment. To ensure that the system does not appear unresponsive to the user's input while enforcing this constraint, we allow a certain amount of “overshoot” beyond the allowed extent. In this range, further mouse movement away from the modeled environment causes an exponentially declining increase in rotation and visual feedback in the form of an increasingly intense blue gradient along the respective screen border (FIG. 3). Once the mouse button is released, the panning quickly snaps back to the allowed range. Thus, the movement appears to be constrained by a (nonlinear) spring rather a hard wall”. Note that the viewpoints is mapped to the directions); and sending to the XR wearable device an indication of the direction (See Gauglitz: Fig. 1, and [0076], “Currently, the 3D model is available only on the remote user's side. Thus, annotations can be correctly occluded by the physical scene by the remote user's renderer, but not by the local user's renderer. While other depth cues (most notably, parallax) still indicate the annotation's location, it would be desirable to enable occlusion. The remote system can send either the model geometry or alternatively local visibility information per annotation back to the local device in order to enable occlusion on the local user's side”. Note that the navigation information and world-stabilized annotation include the positional and directional information). Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Gauglitz, etc. (US 20160358383 A1) in view of Parkinson, etc. (US 20150220506 A1), further in view of Gavriliuc, etc. (US 20160371885 A1) and Song, etc. (US 20180246328 A1). Regarding claim 9, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. However, Gauglitz, modified by Parkinson and Gavriliuc, fails to explicitly disclose that the computing device of claim 1, wherein the display is a first display and wherein displaying the image further comprises: accessing parameters of a second display of the XR wearable device, the parameters comprising an aspect ratio of the second display; and displaying the image on the first display of the computing device in accordance with the parameters of the second display. However, Song teaches that the computing device of claim 1, wherein the display is a first display and wherein displaying the image further comprises: accessing parameters of a second display of the XR wearable device, the parameters comprising an aspect ratio of the second display (See Song: Fig. 1, and [0042], “Referring to FIG. 1, a management environment for an electronic device 100 according to an embodiment may include a head mounted display (HMD) device 200 and at least one external device 400 and/or 500. The electronic device 100 may share contents selected in response to control of the user or contents that is being currently reproduced with the at least one external device 400 and/or 500. In this operation, the electronic device 100 may identify an attribute of the at least one external device 400 and/or 500, and may determine a format of contents that will be shared with the at least one external device 400 and/or 500 according to the identified attribute”; and [0048], “In an embodiment, the formats (e.g., a monocular mode, a binocular mode, a resolution, or a screen ratio) of the contents that are output (or reproduced) from the first electronic device 100 and the at least one external device 400 and/or 500 may be the same or different. For example, the first electronic device 100 and the second electronic device 400 may output monocular contents or binocular contents, and the third electronic device 500 (e.g., a TV) may output only monocular contents. Further, even if the first electronic device 100 and the second electronic device 400 output binocular contents in the same way, the resolutions or screen ratios may be different according to the performance of a device or the interacting HMD device. In this regard, the first electronic device 100 may determine the format of the shared contents based on attribute information (e.g., whether binocular content(s) are managed, the screen ratio, the resolution, or whether a sound is supported) of the at least one external device 400 and/or 500. Hereinafter, various embodiments related to control of the format of the contents that are to be shared and functional operations of the elements that realize the embodiments will be described below”); and displaying the image on the first display of the computing device in accordance with the parameters of the second display See Song: Fig. 1, and [0057], “] The processor 140 may be electrically or functionally connected to at least one element of the electronic device 100 to perform control of the element, communication calculation, or data processing. For example, the processor 140 may share (transmit) at least some of the contents reproduced through the display 130 or the contents stored in the memory 120 with (to) the at least one external device 400 and/or 500 connected via the network 600 or the specific communication channel. In this regard, in an embodiment, the processor 140 may authenticate the at least one external device 400 and/or 500 with which the contents are to be shared, based on identification information (e.g., a type of the device, a device unique identifier (DUID), allocated code information, communication channel information, or network subscription information) of the at least one external device 400 and/or 500 included in the database. The processor 140 may collect attribute information (e.g., whether binocular contents are managed, a screen ratio, a resolution, or whether a sound is supported) on the at least one authenticated external device 400 and/or 500. The processor 140 may convert the format (e.g., a monocular mode or a binocular mode, a resolution, or a screen ratio) of the contents in response to control of the user in consideration of the attribute information of the at least one external device 400 and/or 500, and may share (or transmit) the converted format with (or to) a specific external device. In various embodiments, when the contents are related to security, the processor 140 may perform specific authentication or re-authentication on the specific external device in an operation of sharing contents”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gauglitz to have the computing device of claim 1, wherein the display is a first display and wherein displaying the image further comprises: accessing parameters of a second display of the XR wearable device, the parameters comprising an aspect ratio of the second display; and displaying the image on the first display of the computing device in accordance with the parameters of the second display as taught by Song in order to deliver received user input to the device (See Song: Fig. 2, and [0058], “The HMD device 200 that interacts with the electronic device 100 may support reproduction of virtual reality (VR) or augmented reality (AR) contents in relation to watching of contents of the user, and may receive a user input related to control of reproduction of contents and deliver the received user input to the electronic device 100. In an embodiment, the HMD device 200 may include a communication interface 210 and an input/output interface 220. The communication interface 210 may perform communication with the electronic device 100 or the at least one external device 400 and/or 500 based on wired communication or wireless communication. In an embodiment, the communication interface 210 may include a connector or a port. The input/output interface 220 (e.g., a touchpad, a keypad, a joystick, or a wheel) may deliver a signal or data for an input applied by the user to the electronic device 100 by using the communication interface 210”). Gauglitz teaches a method and system that may provide for an augmented shared visual space for live mobile remote collaboration on physical tasks with live image, 3D/pose data, remote annotation, un-project to world coordinates, and return the world-stabilized annotation to a local device for displaying; while Song teaches a system and method that may interact with the HMD and share visual content with other device, specifically by accessing display attributes information including screen ratio, aspect ratio, resolution of a device in the HMD system and converting or adjusting the content format so that it is properly displayed according to those parameters. Therefore, it is obvious for one of ordinary skills in the arts to modify Gauglitz by Song to obtain the second display parameters such as aspect ratio and adjust the image or video format according to those parameters for proper displaying the content. The motivation to modify Gauglitz by Song is “Use of known technique to improve similar devices (methods, or products) in the same way”. Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Gauglitz, etc. (US 20160358383 A1) in view of Parkinson, etc. (US 20150220506 A1), further in view of Gavriliuc, etc. (US 20160371885 A1) and Zhu, etc. (US 20090179895 A1). Regarding claim 11, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 1 as outlined above. However, Gauglitz, modified by Parkinson and Gavriliuc, fails to explicitly disclose that the computing device of claim 10, further comprising before the determining the 3D coordinates: determining a 3D coordinate of the 3D coordinates based on a pixel of the augmentation being closest, within the image, to the 3D coordinate of the plurality of 3D coordinates. However, Zhu teaches that the computing device of claim 10, further comprising before the determining the 3D coordinates: determining a 3D coordinate of the 3D coordinates based on a pixel of the augmentation being closest, within the image, to the 3D coordinate of the plurality of 3D coordinates (See Zhu: Figs. 2A-B, and [0030], “FIG. 2B is a diagram that depicts an example 250 of displaying the annotation created in example 200 on another image such as, for example, image 104. Image 104 is an image showing building 110 and tree 112. The perspective of image 104 contains the location determined in example 200, and as such, image 104 can be said to correspond to the location. As a result, the content of the annotation created in example 200 is displayed at location 108 on image 104. In embodiments, the text of the annotation may be displayed in an informational balloon pointing to the portion of image 104 corresponding to the location information associated with the annotation. In an embodiment, the portion of image 104 corresponding to the location may be outlined or highlighted. In an embodiment, the annotation created in example 200 may be displayed on a map. These examples are strictly illustrative and are not intended to limit the present invention”;; and [0050], “A match for a feature in the first image among the features in the second image may be determined, for example, as follows. First, the nearest neighbor (e.g., in 128-dimensional space) of a feature in the first image is determined from among the features in the second image. Second, the second-nearest neighbor (e.g., in 128 dimensional-space) of the feature in the first image is determined from among the features in the second image. Third, a first distance between the feature in the first image and the nearest neighboring feature in the second image is determined, and a second distance between the feature in the first image and the second nearest neighboring feature in the second image is determined. Fourth, a feature similarity ratio is calculated by dividing the first distance by the second distance. If the feature similarity ratio is below a particular threshold, there is a match between the feature in the feature in the first image and its nearest neighbor in the second image”. Note that determining the 3D coordinate for annotation based on the 2D pixel location of the annotation within the image and the pre-existing set of 3D coordinates or geometry associated with that image, and when the 3D data is discrete, and obvious mapping is a standard implementation is nearest neighbor/closest point lookup in the image space, and this is the step in the cited limitation). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gauglitz to have the computing device of claim 10, further comprising before the determining the 3D coordinates: determining a 3D coordinate of the 3D coordinates based on a pixel of the augmentation being closest, within the image, to the 3D coordinate of the plurality of 3D coordinates as taught by Zhu in order to enable users to create annotations corresponding to 3D objects while viewing two-dimensional (2D) images (See Zhu: Fig. 1, and [0004], “The present invention relates to annotating images. In an embodiment, the present invention enables users to create annotations corresponding to three-dimensional objects while viewing two-dimensional images. In one embodiment, this is achieved by projecting a selecting object (such as, for example, a bounding box) onto a three-dimensional model created from a plurality of two-dimensional images. The selecting object is input by a user while viewing a first image corresponding to a portion of the three-dimensional model. A location corresponding to the projection on the three-dimensional model is determined, and content entered by the user while viewing the first image is associated with the location. The content is stored together with the location information to form an annotation. The annotation can be retrieved and displayed together with other images corresponding to the location”). Gauglitz teaches a method and system that may provide for an augmented shared visual space for live mobile remote collaboration on physical tasks with live image, 3D/pose data, remote annotation, un-project to world coordinates, and return the world-stabilized annotation to a local device for displaying; while Zhu teaches a system and method that may create annotations corresponding to three-dimensional objects while viewing two-dimensional images, determine a location corresponding to the projection on the three-dimensional model, and associate the annotation content entered by the user with the location. Therefore, it is obvious for one of ordinary skills in the arts to modify Gauglitz by Zhu to determine a location for the augmentation within the image in the nearest neighbor position. The motivation to modify Gauglitz by Zhu is “Use of known technique to improve similar devices (methods, or products) in the same way”. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Gauglitz, etc. (US 20160358383 A1) in view of Parkinson, etc. (US 20150220506 A1), further in view of Gavriliuc, etc. (US 20160371885 A1) and Flint, etc. (US 20150043784 A1). Regarding claim 16, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 13 as outlined above However, Gauglitz, modified by Parkinson and Gavriliuc, fails to explicitly disclose that the XR wearable device of claim 13, wherein the operations further comprising: determining the plurality of 3D coordinates based on visual inertial odometry (VIO). However, Flint teaches that that the XR wearable device of claim 13, wherein the operations further comprising: determining the plurality of 3D coordinates based on visual inertial odometry (VIO) (See Flint: Figs. 1-2, and [0030], “FIG. 2 is a schematic illustrating an example of a visual-based inertial navigation device 100, such as the electronic computing device that may be used to produce the path 10 of FIG. 1. The device 100 includes multiple components that make up a visual-based inertial navigation system. For example, the device 100 includes an image sensor 102 that converts an optical image into an electronic signal, such as a digital camera. The sensor 102 may utilize any appropriate image sensing components, such as a digital charge-coupled device (CCD), complementary metal-oxide-semiconductor (CMOS) pixel sensors or infrared sensors. Alternatively, the image sensor 102 may include a depth sensor, a stereo camera pair, a flash lidar sensor, a laser sensor, or any combination of these. The image sensor 102 may be formed entirely in hardware or may also be configured to include software for modifying detected images. The device 100 also includes an inertial measurement unit 104. The IMU 104 may include several electronic hardware components, including a tri-axial gyroscope and accelerometer, for recording inertial data of the device 100. For example, the IMU 104 may measure and report on the device's six degrees of freedom (X, Y, and Z Cartesian coordinates of the device's acceleration, and roll, pitch, and yaw components of the device's angular velocity). The IMU 104 may output other inertial data, as well. Various IMUs are commercially available or are pre-installed on portable electronic computing devices”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gauglitz to have the XR wearable device of claim 13, wherein the operations further comprising: determining the plurality of 3D coordinates based on visual inertial odometry (VIO) as taught by Flint in order to reduce the computational burden of processing the information matrix (See Flint: Fig. 1, and [0041], “Thus, by reducing the total number of model parameters over which the optimization problem is solved, marginalization reduces the computational burden of the SWF module 108, but also maintains a consistent model estimate over time by summarizing and transferring information between each window and the next.”). Gauglitz teaches a method and system that may provide for an augmented shared visual space for live mobile remote collaboration on physical tasks with live image, 3D/pose data, remote annotation, un-project to world coordinates, and return the world-stabilized annotation to a local device for displaying; while Flint teaches a system and method that may receive sensor measurements from a pre-processing module including image data and inertial data for a device, and outputting a state of the device based on the transferred information including the position coordinate data. Therefore, it is obvious for one of ordinary skills in the arts to modify Gauglitz by Flint to use the well-known VIO system in the mobile wearable devices to obtain 3D coordinates for the wearable devices. The motivation to modify Gauglitz by Flint is “Use of known technique to improve similar devices (methods, or products) in the same way”. Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Gauglitz, etc. (US 20160358383 A1) in view of Parkinson, etc. (US 20150220506 A1), further in view of Gavriliuc, etc. (US 20160371885 A1) and Flint, etc. (US 20190325604 A1). Regarding claim 18, Gauglitz, Parkinson and Gavriliuc teach all the features with respect to claim 13 as outlined above. However, Gauglitz, modified by Parkinson and Gavriliuc, fails to explicitly disclose that the XR wearable device of claim 13, wherein the operations further comprise: sending the image to a host computing device with an instruction to determine the plurality of 3D coordinates for the image; and receiving the plurality of 3D coordinates from the host computing device. However, Flink teaches that the XR wearable device of claim 13, wherein the operations further comprise: sending the image to a host computing device with an instruction to determine the plurality of 3D coordinates for the image (See Flink: Figs. 1-3, and [0026], “In embodiments, consumer device 102 is adapted to process a captured scene and obtain an associated point cloud for use in AR applications. The point cloud may be derived using photogrammetric techniques, directly measured by use of a dedicated depth sensor, by another suitable algorithm, by a combination of one or more of the foregoing, or any other suitable technique that will provide useful AR information now known or later developed. Consumer device 102 may directly (e.g. locally on the device) process the captured scene to calculate the AR data, or in other embodiments may upload the scene to a remote server for processing and calculation of the AR data. Processing of the captured scene may be accomplished by dedicated hardware within consumer device 102, by software, or a combination of both”; and [0039], “In block 306, the subsequent capture and its associated 3D point cloud may be correlated with the source capture and point cloud. In some embodiments, the point cloud associated with each capture may be expressed via an x-y-z coordinate system that is referenced with respect to the capture. Correlation may allow establishing of a single or common coordinate system between captured images so that their corresponding point clouds may be properly merged. Without correlation, the point cloud of the subsequent capture may not properly align in space with the point cloud of the source capture. Further, correlation may also serve as verification that the subsequent capture is directed to the same scene as the original source capture. As with blocks 302 and 304, in some examples block 306 is carried out by the consumer device 102 executing all blocks of method 300. In other examples, block 306 may be executed by a remote server or system from an upload of a video or image and AR data captured and calculated in blocks 302 and 304, respectively. In still other examples, blocks 304 and 306 may be carried out by the remote server, with consumer device 102 uploaded the video or image it captured in block 302”); and receiving the plurality of 3D coordinates from the host computing device (See Flink: Figs. 1-3, and [0036], “The captured image may come from a variety of sources. In some examples, the original camera 104 used to capture the initial source image or video may be used to capture the subsequent image or video. In other examples, a different device or devices may be used to capture the subsequent image or video. In still other examples, the video or image may be captured (or capture of the image initiated) by a remote trigger. Specifically, a third party may remotely initiate a capture of the image or video on a user's device, such as where the user and third party are engaged in a communications session, and the third party triggers a capture of the video (and corresponding analysis to calculate AR data) from the user's device. In some such examples, the user and third party may be engaged in an AR session where the third party is placing objects in a video stream from the user's device. The user's device in such an example may already by calculating corresponding AR data; the third party may then trigger the user's device to upload the image and AR data to a cloud service. Alternatively, the cloud service may be facilitating the communications session between the user and the third party, and the third party may trigger the cloud service to record a portion of the session, comprised of at least a portion of the relayed video stream and corresponding point cloud. Regardless of the device used for the subsequent capture, the subsequently captured image need not be of the same type and format as the initial source image, viz. the source image may be a still frame, and the subsequent image may be a video, and vice-versa”; and [0059], “Computer program code for carrying out operations of the present disclosure may be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider)” Note that the off-load pattern : client sends images or scene data to the host server to compute the 3D coordinates in point clouds, and the client receive the results, and this is a common Client-server architecture). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gauglitz to have the XR wearable device of claim 13, wherein the operations further comprise: sending the image to a host computing device with an instruction to determine the plurality of 3D coordinates for the image; and receiving the plurality of 3D coordinates from the host computing device as taught by Flint in order enable triggering cloud service by third party to record a portion of the session (See Flint: Fig. 1, and [0036], “Alternatively, the cloud service may be facilitating the communications session between the user and the third party, and the third party may trigger the cloud service to record a portion of the session, comprised of at least a portion of the relayed video stream and corresponding point cloud. Regardless of the device used for the subsequent capture, the subsequently captured image need not be of the same type and format as the initial source image, viz. the source image may be a still frame, and the subsequent image may be a video, and vice-versa”). Gauglitz teaches a method and system that may provide for an augmented shared visual space for live mobile remote collaboration on physical tasks with live image, 3D/pose data, remote annotation, un-project to world coordinates, and return the world-stabilized annotation to a local device for displaying; while Flint teaches a system and method that may augment the previously captured augmented reality data with subsequently captured data with a local device to capture the scene image, upload them into the server to process the scene data and adding annotation, and return the result to the local device for displaying or viewing. Therefore, it is obvious for one of ordinary skills in the arts to modify Gauglitz by Flint to upload the scene data to the host server for processing and return the processed results back to the local client device. The motivation to modify Gauglitz by Flint is “Use of known technique to improve similar devices (methods, or products) in the same way”. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to GORDON G LIU whose telephone number is (571)270-0382. The examiner can normally be reached Monday - Friday 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona E Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GORDON G LIU/Primary Examiner, Art Unit 2618
Read full office action

Prosecution Timeline

Feb 26, 2025
Application Filed
Sep 02, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743838
PIXEL GENERATION TECHNIQUE
2y 8m to grant Granted Sep 22, 2026
Patent 12743821
RECORDING MEDIUM AND INFORMATION PROCESSING DEVICE
2y 3m to grant Granted Sep 22, 2026
Patent 12744018
IMAGE OUTPUT CONTROL DEVICE AND METHOD
2y 1m to grant Granted Sep 22, 2026
Patent 12725544
SCREEN DISPLAY DRIVING METHOD AND APPARATUS, SCREEN INFORMATION CONFIGURATION METHOD AND APPARATUS, MEDIUM AND DEVICE
2y 1m to grant Granted Sep 01, 2026
Patent 12725367
PICTURE DISPLAY METHOD, SYSTEM, AND APPARATUS, DEVICE, AND STORAGE MEDIUM
2y 2m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
83%
Grant Probability
98%
With Interview (+14.8%)
2y 2m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 701 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month