DETAILED ACTION
This action is in response to the application filed 11/30/2024. Claims 1 – 20 are pending and have
been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
The claimed invention of Claim 20 is directed to non-statutory subject matter. The claim does not fall within at least one of the four categories of patent eligible subject matter because, under its broadest reasonable interpretation, the recited “computer readable storage medium, having a program or instructions stored thereon” encompasses a computer program per se. The limitation merely describes the capability of the program, and does not positively require that the program be embodied in a statutory manufacture, such as a non-transitory computer-readable storage medium or a computer memory. Accordingly, the claim encompasses non-statutory subject matter and therefore, Claim 20 is rejected.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 – 20 are rejected under 35 U.S.C. 103 as being unpatentable over Roberts et al. (U.S. Pub. No. 2013/0325970, hereinafter “Roberts”) in view of Liu et al. (W.O. Pub. No. 2022/206624, hereinafter “Liu”).
Regarding Claim 1, Roberts teaches
A video data transmission method, applied to a first terminal (see Roberts Paragraph [0047], method, Abstract, During operation, the system establishes a real-time video-sharing session between a remote field computer and a local computer), and the method comprises:
performing a video call with a second terminal (see Roberts Paragraph [0024], In collaborative video-sharing system 100, a number of clients, such as a web-based client 106 and a field client 108, interact with each other and a real-time collaboration server 102 via a network 110, Paragraph [0027], Web-based client 106 may be an electronic device that can stream real-time video from real-time collaboration server 102, display the real-time video to a user of web-based client 106, Paragraph [0037], FIG. 3 presents a time-space diagram illustrating the real-time video-sharing process, in accordance with an embodiment of the present invention. During operation, a field client 302 and an expert client 306 both establish a user session with a collaboration server 304 (operations 308 and 310). Subsequently, field client 302 captures live video images (operation 312), and concurrently streams the live video images to collaboration server 304 (operation 314). The field client can be a head-mounted computer, a webcam hosted by a PC, a mobile computing device, or a web-enabled wired or wireless surveillance camera. Collaboration server 304 optionally processes the live video images (operation 316), and forwards the video to an expert client 306 (operation 318). For example, collaboration server 304 may cache the video for further use. An expert can view the video at expert client 306, thus being able to see what a novice is seeing at the remote site. For example, if the novice user is repairing a car engine, a camera mounted on field client 302 captures live videos showing the car engine and every movement of the novice. In one embodiment, expert client 306 can directly grab videos from field client 302, if expert client 306 is informed of the location of a video server associated with field client 302. However, in cases where field client 302 is behind a firewall or under other similar restrictions, expert client 306 cannot directly receive video from field client 302, and collaboration server 304 is needed to relay the video. Moreover, in situations where multiple other clients attempt to stream live videos from field client 302, due to its limited upstream bandwidth, field client 302 may not be able to satisfy all streaming requests from the multiple other clients, then collaboration server 304 is needed to relay the live videos to the multiple clients. Alternatively, the live videos can be distributed to the multiple other clients from filed client 302 using a multicast protocol or a content-centric networking (CCN) protocol, and Figure 3, in which field client 302 establishes session (or expert client 306), and field client streams video to server, thus to the expert client);
receiving labeled data which is sent by the second terminal, wherein the labeled data is used for indicating a target label added to a target object in a video frame shared by the first terminal (see Roberts Paragraph [0039], While viewing the live videos, an expert user using expert client 306 can send commands to collaboration server 304 instructing collaboration server 304 to extract one or more still images from the live video (operations 322 and 324). Collaboration server 304 then sends the extracted still images back to expert client 306 to allow the expert user to annotate on top of the extracted images (operations 326 and 328). Subsequently, the annotation is sent back to collaboration server 304 (operation 330). In one embodiment, an annotation can be created by the expert user as an HTML canvas element. The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During operation, the system establishes a real-time video-sharing session between a remote field computer and a local computer. During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0006], In a variation on this embodiment, forwarding the annotation involves: generating a message associated with the annotation and sending the message to the remote field computer, Paragraph [0010], In a variation on this embodiment, the annotation is a HyperText Markup Language (HTML) canvas object, Paragraph [0036], The real-time communication includes, but is not limited to: videos, images, annotations on the videos or the images, and audio conversations. In one embodiment, an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines, Paragraph [0044], the expert selects a step within a procedure (operation 402), and extracts one or more images associated with the selected step from one or more video streams taken during remote-assisted service sessions (operation 404). Subsequently, the expert annotates and tags the images (operation 406), and adds a textual description (operation 408). The annotation and the textual description together explain how to perform the selected step to a viewer. In one embodiment, the expert may create animated annotations);
determining, based on the labeled data, the target object labeled by the second terminal (see Roberts Paragraph [0039], While viewing the live videos, an expert user using expert client 306 can send commands to collaboration server 304 instructing collaboration server 304 to extract one or more still images from the live video (operations 322 and 324). Collaboration server 304 then sends the extracted still images back to expert client 306 to allow the expert user to annotate on top of the extracted images (operations 326 and 328). Subsequently, the annotation is sent back to collaboration server 304 (operation 330). In one embodiment, an annotation can be created by the expert user as an HTML canvas element. The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During operation, the system establishes a real-time video-sharing session between a remote field computer and a local computer. During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0006], In a variation on this embodiment, forwarding the annotation involves: generating a message associated with the annotation and sending the message to the remote field computer, Paragraph [0010], In a variation on this embodiment, the annotation is a HyperText Markup Language (HTML) canvas object, Paragraph [0036], The real-time communication includes, but is not limited to: videos, images, annotations on the videos or the images, and audio conversations. In one embodiment, an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines, Paragraph [0044], the expert selects a step within a procedure (operation 402), and extracts one or more images associated with the selected step from one or more video streams taken during remote-assisted service sessions (operation 404). Subsequently, the expert annotates and tags the images (operation 406), and adds a textual description (operation 408). The annotation and the textual description together explain how to perform the selected step to a viewer. In one embodiment, the expert may create animated annotations);
determining the target object in video data, and adding the target label to a position, corresponding to the target object, in the video data (see Roberts Paragraph [0039], While viewing the live videos, an expert user using expert client 306 can send commands to collaboration server 304 instructing collaboration server 304 to extract one or more still images from the live video (operations 322 and 324). Collaboration server 304 then sends the extracted still images back to expert client 306 to allow the expert user to annotate on top of the extracted images (operations 326 and 328). Subsequently, the annotation is sent back to collaboration server 304 (operation 330). In one embodiment, an annotation can be created by the expert user as an HTML canvas element. The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During operation, the system establishes a real-time video-sharing session between a remote field computer and a local computer. During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0006], In a variation on this embodiment, forwarding the annotation involves: generating a message associated with the annotation and sending the message to the remote field computer, Paragraph [0010], In a variation on this embodiment, the annotation is a HyperText Markup Language (HTML) canvas object, Paragraph [0036], The real-time communication includes, but is not limited to: videos, images, annotations on the videos or the images, and audio conversations. In one embodiment, an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines, Paragraph [0044], the expert selects a step within a procedure (operation 402), and extracts one or more images associated with the selected step from one or more video streams taken during remote-assisted service sessions (operation 404). Subsequently, the expert annotates and tags the images (operation 406), and adds a textual description (operation 408). The annotation and the textual description together explain how to perform the selected step to a viewer. In one embodiment, the expert may create animated annotations); and
transmitting, to the second terminal, the video data added with the target label (see Roberts Paragraph [0039], The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0036], The expert may select one video stream from the multiple video streams to annotate, and distributes the annotated videos to users of all field clients via real-time collaboration server 200).
Roberts does not expressively teach
A first channel to transport audio and video media, and a second channel to transport labeled data
However, Liu teaches
A first data channel for transmitting media associated with a video call such as audio and video (see Liu Page 3, media stream channel, receive the first media stream for AR communication between the first terminal device and the second terminal device through the established media stream channel), and a second data channel for transmitting auxiliary data such as object identifiers, tracking data, annotations, etc. (see Liu Page 4, auxiliary data channel, receive target object information from the first terminal device through the auxiliary data channel). Additionally, the AR media processing network element receives both channels simultaneously and then performs media augmenting processing to create an AR media stream that combines the two channels for output (see Liu Page 3, It is further configured to perform enhancement processing on the media stream according to the first AR auxiliary data to obtain a first AR media stream, and send the first AR media stream to the first terminal device through the media stream channel).
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the teaching of a system that enables real-time video sharing between a remote and local computer, allowing a user of a local computer to annotate the live video stream, with the annotations synchronized to and displayed over the corresponding portions of the video at the remote computer (as taught in Roberts), with a system that uses a first channel to transport audio and video media, and a second channel to transport labeled data (as taught in Liu), the motivation being to reduce connection waiting times and enable higher-quality voice and video communications by allowing each channel to be independently optimized for its respective data type (see Liu Page 1).
Regarding Claim 2, Roberts in view of Liu teaches
The method according to claim 1, wherein before the receiving labeled data which is sent by the second terminal through a second data channel, the method further comprises:
receiving a session establishment signaling which is sent by the second terminal (see Roberts Figure 3, expert client establishes session 310);
analyzing the session establishment signaling, and obtaining indication information which is carried in the session establishment signaling and indicates that the second data channel needs to be established (see Liu Page 12, In an example, the first terminal device may send the first request message to the IMS core network element in the process of establishing AR communication with the second terminal device. For example, the first request message may use a Session Initiation Protocol (Session Initiation Protocol, SIP) request (invite) message. The first request message may also carry address information of the first terminal device, for example, the address information of the first terminal device may include the IP address and/or port number of the first terminal device. Optionally, the first request message is further used to request the establishment of a media stream channel, where the media stream channel is used to transmit a media stream for AR communication between the first terminal device and the second terminal device. The first request message also carries a description parameter used by the first terminal device to establish the media stream channel with the AR media processing network element. For example, the first SDP information further includes a description parameter for the first terminal device to establish the media stream channel with the AR media processing network element. After receiving the first request message, the IMS core network element sends the second request message to the AR media processing network element. The second request message is used to request to establish the auxiliary data channel between the AR media processing network element and the first terminal device. Exemplarily, the second request message also carries the first SDP information); and
establishing the second data channel to the second terminal (see Liu Page 13, When the AR media processing network element receives the second request message, it sends a second response message to the IMS core network element, where the second response message is a response message corresponding to the second request message. For example, the second response message is used to indicate to the IMS core network element that the AR media processing network element confirms the establishment of the auxiliary data channel. For example, the second response message carries the second description parameter of the AR media processing network element for establishing the auxiliary data channel. The AR media processing network element may determine the second description parameter that it supports from the first description parameter. The second description parameter may adopt the SDP protocol, and may also adopt other protocols. For the convenience of distinction, the SDP information of the AR media processing network element is referred to as the second SDP information (may be referred to as the second SDP for short). The second SDP information includes a description parameter of the AR media processing network element for establishing the auxiliary data channel with the first terminal device. For example, the second SDP information includes parameters such as the port of the auxiliary data channel corresponding to the AR media processing network element, the type of media stream, and the supported codec format. The media stream types may include video (video stream), audio (audio stream), and data channel (AR auxiliary data). For example, in this case, the second SDP information may include m lines for describing AR assistance data negotiated by the AR media processing network element, and the media stream type of the m lines is AR assistance data. When the IMS core network element receives the second response message, it sends the first response message to the first terminal device, where the first response message carries the address of the AR media processing network element. The first response message carries the second SDP information of the AR media processing network element).
Regarding Claim 3, Roberts in view of Liu teaches
The method according to claim 1, wherein the labeled data comprises at least one of the following:
full image data added with the target label (see Roberts Paragraph [0023], In embodiments of the present invention, a collaborative video application run on a server-client system allows multiple clients to share and annotate real-time video streams. For example, an expert at a remote site can view the manipulation of a novice on the equipment through the live video stream as if the expert was "watching over the shoulder" of the novice. In addition, the application also allows the expert to provide assistance to the novice as needed by enabling real-time annotation, which means the expert can draw an arrow to point to or a circle to surround a component on the equipment when instructing the novice to perform certain actions on the component. For example, if the expert wants to instruct the novice to move a handle, the expert can extract a frame from the live video stream that shows the handle, draw an arrow pointing to the handle, and send the annotated frame to the novice, Paragraph [0027], Web-based client 106 may be an electronic device that can stream real-time video from real-time collaboration server 102, display the real-time video to a user of web-based client 106, and perform certain video processing tasks, such as receiving frames extracted from the real-time video from real-time collaboration server 102 and annotating the extracted frames based on user input. Note that, in one embodiment, real-time collaboration server 102 receives frame-extraction instructions from an expert user via web-based client 106. In addition, web-based client 106 is able to upload the annotated video frames back to real-time collaboration server 102, which in turn forwards the annotated video frames to a remote user to allow the remote user to view instructions embedded in the annotated video frames, Paragraph [0032], In one embodiment, the video API also includes functionalities that can capture messages associated with real-time annotations and associate these annotations with a captured video stream, Paragraph [0036], an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines. For example, to explain a maintenance operation involving a certain component, the expert user can circle the component on an image captured from the live video and write detailed operation instructions at the bottom of the image. Note that once the annotation is created, a message is created expressing the annotation);
the target label and region information labeled by the target label (see Roberts Paragraph [0023], In embodiments of the present invention, a collaborative video application run on a server-client system allows multiple clients to share and annotate real-time video streams. For example, an expert at a remote site can view the manipulation of a novice on the equipment through the live video stream as if the expert was "watching over the shoulder" of the novice. In addition, the application also allows the expert to provide assistance to the novice as needed by enabling real-time annotation, which means the expert can draw an arrow to point to or a circle to surround a component on the equipment when instructing the novice to perform certain actions on the component. For example, if the expert wants to instruct the novice to move a handle, the expert can extract a frame from the live video stream that shows the handle, draw an arrow pointing to the handle, and send the annotated frame to the novice, Paragraph [0027], Web-based client 106 may be an electronic device that can stream real-time video from real-time collaboration server 102, display the real-time video to a user of web-based client 106, and perform certain video processing tasks, such as receiving frames extracted from the real-time video from real-time collaboration server 102 and annotating the extracted frames based on user input. Note that, in one embodiment, real-time collaboration server 102 receives frame-extraction instructions from an expert user via web-based client 106. In addition, web-based client 106 is able to upload the annotated video frames back to real-time collaboration server 102, which in turn forwards the annotated video frames to a remote user to allow the remote user to view instructions embedded in the annotated video frames, Paragraph [0032], In one embodiment, the video API also includes functionalities that can capture messages associated with real-time annotations and associate these annotations with a captured video stream, Paragraph [0036], an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines. For example, to explain a maintenance operation involving a certain component, the expert user can circle the component on an image captured from the live video and write detailed operation instructions at the bottom of the image. Note that once the annotation is created, a message is created expressing the annotation).
Regarding Claim 4, Roberts in view of Liu teaches
The method according to claim 1, wherein after the transmitting, to the second terminal, the video data added with the target label, the method further comprises:
locally displaying the video data added with the target label on the first terminal (see Roberts Paragraph [0039], The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0036], The expert may select one video stream from the multiple video streams to annotate, and distributes the annotated videos to users of all field clients via real-time collaboration server 200).
Regarding Claim 5, Roberts in view of Liu teaches
The method according to claim 1, wherein before the adding the target label to a position, corresponding to the target object, in the video data, the method further comprises:
determining that the first terminal supports adding the target label in the video data (see Liu Page 10, Both terminal devices in AR communication may have AR capabilities, or one of the terminal devices may have AR capabilities, and the other terminal device may not have AR capabilities, Page 11, one of the two terminal devices in AR communication does not have AR capability. For example, as shown in FIG. 3 , the first terminal device has the AR capability, and the second terminal device does not have the AR capability).
Regarding Claim 6, Roberts in view of Liu teaches
The method according to claim 1, wherein after the receiving labeled data which is sent by the second terminal through a second data channel, the method further comprises:
in a case of determining that the first terminal supports adding the target label in the video data, transmitting the video data to the second terminal through the first data channel (see Liu Page 10, Both terminal devices in AR communication may have AR capabilities, or one of the terminal devices may have AR capabilities, and the other terminal device may not have AR capabilities, Page 11, one of the two terminal devices in AR communication does not have AR capability. For example, as shown in FIG. 3 , the first terminal device has the AR capability, and the second terminal device does not have the AR capability, Page 3, media stream channel, receive the first media stream for AR communication between the first terminal device and the second terminal device through the established media stream channel).
Regarding Claim 7, it is rejected similarly as Claim 1.
Regarding Claim 8, it is rejected similarly as Claim 1.
Regarding Claim 9, it is rejected similarly as Claim 2.
Regarding Claim 10, Roberts in view of Liu teaches
The method according to claim 7, wherein the obtaining target video data, which is added with the target label, of the first terminal comprises:
receiving first video data which is transmitted by the first terminal through the first data channel (see Roberts Paragraph [0024], In collaborative video-sharing system 100, a number of clients, such as a web-based client 106 and a field client 108, interact with each other and a real-time collaboration server 102 via a network 110, Paragraph [0027], Web-based client 106 may be an electronic device that can stream real-time video from real-time collaboration server 102, display the real-time video to a user of web-based client 106, Paragraph [0037], FIG. 3 presents a time-space diagram illustrating the real-time video-sharing process, in accordance with an embodiment of the present invention. During operation, a field client 302 and an expert client 306 both establish a user session with a collaboration server 304 (operations 308 and 310). Subsequently, field client 302 captures live video images (operation 312), and concurrently streams the live video images to collaboration server 304 (operation 314). The field client can be a head-mounted computer, a webcam hosted by a PC, a mobile computing device, or a web-enabled wired or wireless surveillance camera. Collaboration server 304 optionally processes the live video images (operation 316), and forwards the video to an expert client 306 (operation 318)); and
obtaining the target video data by adding the target label to a position, corresponding to the target object, in the first video data (see Roberts Paragraph [0039], While viewing the live videos, an expert user using expert client 306 can send commands to collaboration server 304 instructing collaboration server 304 to extract one or more still images from the live video (operations 322 and 324). Collaboration server 304 then sends the extracted still images back to expert client 306 to allow the expert user to annotate on top of the extracted images (operations 326 and 328). Subsequently, the annotation is sent back to collaboration server 304 (operation 330). In one embodiment, an annotation can be created by the expert user as an HTML canvas element. The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During operation, the system establishes a real-time video-sharing session between a remote field computer and a local computer. During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0006], In a variation on this embodiment, forwarding the annotation involves: generating a message associated with the annotation and sending the message to the remote field computer, Paragraph [0010], In a variation on this embodiment, the annotation is a HyperText Markup Language (HTML) canvas object, Paragraph [0036], The real-time communication includes, but is not limited to: videos, images, annotations on the videos or the images, and audio conversations. In one embodiment, an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines, Paragraph [0044], the expert selects a step within a procedure (operation 402), and extracts one or more images associated with the selected step from one or more video streams taken during remote-assisted service sessions (operation 404). Subsequently, the expert annotates and tags the images (operation 406), and adds a textual description (operation 408). The annotation and the textual description together explain how to perform the selected step to a viewer. In one embodiment, the expert may create animated annotations).
Regarding Claim 11, Roberts in view of Liu teaches
The method according to claim 10, wherein after the adding a target label for a target object to a video frame which is shared by the first terminal, the method further comprises:
sending labeled data to the first terminal through a second data channel to the first terminal (see Liu Page 2, the AR media processing network element sends multiple virtual object identifiers of the virtual object type to the first terminal device through the auxiliary data channel, Page 4, auxiliary data channel, receive target object information from the first terminal device through the auxiliary data channel), wherein the labeled data is used for indicating the target label added to the target object in the video frame shared by the first terminal (see Roberts Paragraph [0039], While viewing the live videos, an expert user using expert client 306 can send commands to collaboration server 304 instructing collaboration server 304 to extract one or more still images from the live video (operations 322 and 324). Collaboration server 304 then sends the extracted still images back to expert client 306 to allow the expert user to annotate on top of the extracted images (operations 326 and 328). Subsequently, the annotation is sent back to collaboration server 304 (operation 330). In one embodiment, an annotation can be created by the expert user as an HTML canvas element. The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During operation, the system establishes a real-time video-sharing session between a remote field computer and a local computer. During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0006], In a variation on this embodiment, forwarding the annotation involves: generating a message associated with the annotation and sending the message to the remote field computer, Paragraph [0010], In a variation on this embodiment, the annotation is a HyperText Markup Language (HTML) canvas object, Paragraph [0036], The real-time communication includes, but is not limited to: videos, images, annotations on the videos or the images, and audio conversations. In one embodiment, an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines, Paragraph [0044], the expert selects a step within a procedure (operation 402), and extracts one or more images associated with the selected step from one or more video streams taken during remote-assisted service sessions (operation 404). Subsequently, the expert annotates and tags the images (operation 406), and adds a textual description (operation 408). The annotation and the textual description together explain how to perform the selected step to a viewer. In one embodiment, the expert may create animated annotations).
Regarding Claim 12, Roberts in view of Liu teaches
The method according to claim 11, wherein before the obtaining the target video data by adding the target label to a position, corresponding to the target object, in the first video data, the method further comprises:
determining that the first video data does not contain the target label (see Roberts Paragraph [0039], While viewing the live videos, an expert user using expert client 306 can send commands to collaboration server 304 instructing collaboration server 304 to extract one or more still images from the live video (operations 322 and 324). Collaboration server 304 then sends the extracted still images back to expert client 306 to allow the expert user to annotate on top of the extracted images (operations 326 and 328). Subsequently, the annotation is sent back to collaboration server 304 (operation 330). In one embodiment, an annotation can be created by the expert user as an HTML canvas element. The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During operation, the system establishes a real-time video-sharing session between a remote field computer and a local computer. During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0006], In a variation on this embodiment, forwarding the annotation involves: generating a message associated with the annotation and sending the message to the remote field computer, Paragraph [0010], In a variation on this embodiment, the annotation is a HyperText Markup Language (HTML) canvas object, Paragraph [0036], The real-time communication includes, but is not limited to: videos, images, annotations on the videos or the images, and audio conversations. In one embodiment, an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines, Paragraph [0044], the expert selects a step within a procedure (operation 402), and extracts one or more images associated with the selected step from one or more video streams taken during remote-assisted service sessions (operation 404). Subsequently, the expert annotates and tags the images (operation 406), and adds a textual description (operation 408). The annotation and the textual description together explain how to perform the selected step to a viewer. In one embodiment, the expert may create animated annotations, in which original incoming video does not contain annotations or labels).
Regarding Claim 13, Roberts in view of Liu teaches
The method according to claim 7, wherein the obtaining target video data, which is added with the target label, of the first terminal comprises:
sending the labeled data to a data server through a second data channel to the data server (see Liu Page 2, the AR media processing network element obtains the first virtual object corresponding to the identifier of the first virtual object from the third-party server, and sends the first virtual object to the first terminal through the auxiliary data channel The device sends the first virtual object), wherein the labeled data is used for indicating the target label added to the target object in the video frame shared by the first terminal (see Roberts Paragraph [0039], While viewing the live videos, an expert user using expert client 306 can send commands to collaboration server 304 instructing collaboration server 304 to extract one or more still images from the live video (operations 322 and 324). Collaboration server 304 then sends the extracted still images back to expert client 306 to allow the expert user to annotate on top of the extracted images (operations 326 and 328). Subsequently, the annotation is sent back to collaboration server 304 (operation 330). In one embodiment, an annotation can be created by the expert user as an HTML canvas element. The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During operation, the system establishes a real-time video-sharing session between a remote field computer and a local computer. During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0006], In a variation on this embodiment, forwarding the annotation involves: generating a message associated with the annotation and sending the message to the remote field computer, Paragraph [0010], In a variation on this embodiment, the annotation is a HyperText Markup Language (HTML) canvas object, Paragraph [0036], The real-time communication includes, but is not limited to: videos, images, annotations on the videos or the images, and audio conversations. In one embodiment, an expert can create annotations on top of a video using an HTML canvas object, which allows the expert user to draw annotations on an image captured from the video, and the annotations can be expressed using a variety of common graphical elements, such as text, boxes, and freehand lines, Paragraph [0044], the expert selects a step within a procedure (operation 402), and extracts one or more images associated with the selected step from one or more video streams taken during remote-assisted service sessions (operation 404). Subsequently, the expert annotates and tags the images (operation 406), and adds a textual description (operation 408). The annotation and the textual description together explain how to perform the selected step to a viewer. In one embodiment, the expert may create animated annotations); and
receiving the target video data which is transmitted by the data server through the second data channel and is added with the target label (see Roberts Paragraph [0039], The creation of the annotation triggers a JSON message expressing the annotation to be created and sent to collaboration server 304. Collaboration server 304 subsequently distributes the annotation to other clients, such as field client 302 (operation 332). Field client 302 associates the annotation with corresponding video elements (operation 334), and overlays the annotation on top of the video elements, such as still images (operation 336), Abstract, During the established real-time video-sharing session, the system receives a real-time video stream from a remote field computer, forwards the real-time video stream to a local computer to allow an expert to provide an annotation to the real-time video stream, receives the annotation from the local computer, and forwards the annotation to the remote field computer, which associates the annotation with a corresponding portion of the real-time video stream and displays the annotation on top of the corresponding portion of the real-time video stream, Paragraph [0036], The expert may select one video stream from the multiple video streams to annotate, and distributes the annotated videos to users of all field clients via real-time collaboration server 200).
Regarding Claim 14, it is rejected similarly as Claim 2.
Regarding Claims 15 - 16, they are rejected similarly as Claims 1 - 2. The method applied to a data server can be found in Liu (Page 6, application server to perform the corresponding function in the method).
Regarding Claim 17, it is rejected similarly as Claim 1. The system can be found in Roberts (Abstract, system).
Regarding Claim 18, Roberts in view of Liu teaches
The system according to claim 17, further comprising:
a data server, configured to perform the method as claimed in claim 15 (see Liu Page 6, application server to perform the corresponding function in the method).
Regarding Claims 19 - 20, they are rejected similarly as Claims 1 - 2. The non-transitory electronic device can be found in Roberts (Paragraph [0046], The data structures and code described in this detailed description are typically stored on a computer-readable storage medium, which may be any device or medium that can store code and/or data for use by a computer system).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARISSA A JONES whose telephone number is (703)756-1677. The examiner can normally be reached Telework M-F 6:30 AM - 4:00 PM CT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CARISSA A JONES/Examiner, Art Unit 2691
/DUC NGUYEN/Supervisory Patent Examiner, Art Unit 2691