Prosecution Insights
Last updated: August 06, 2026
Application No. 19/044,170

ENHANCED VIDEO SUPPORT

Non-Final OA §102§103
Filed
Feb 03, 2025
Priority
Feb 02, 2024 — provisional 63/549,063
Examiner
JONES, CARISSA ANNE
Art Unit
Tech Center
Assignee
8x8 Inc.
OA Round
1 (Non-Final)
80%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
24 granted / 30 resolved
+20.0% vs TC avg
Strong +30% interview lift
Without
With
+30.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
26 currently pending
Career history
61
Total Applications
across all art units

Statute-Specific Performance

§101
3.0%
-37.0% vs TC avg
§103
79.0%
+39.0% vs TC avg
§102
11.4%
-28.6% vs TC avg
§112
3.6%
-36.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 30 resolved cases

Office Action

§102 §103
DETAILED ACTION This action is in response to the application filed 02/03/2025. Claims 1 – 20 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1, 5 – 7, 11 and 15 – 17 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Bartov et al. (U.S. Pub. No. 2025/0097569, hereinafter “Bartov”). Regarding Claim 1, Bartov teaches A computer-implemented method (see Bartov Paragraph [0004], computer-implemented method) comprising: sending a video elevation request configured to transfer an electronic communication of a software communications platform, between an end user device and an agent device, from a first state to a second state being a video call executed via the software communications platform (see Bartov Paragraph [0059], With reference to an illustrative embodiment, a collaboration conversation participant may operate a user device, such as capture device 100, to capture visual content of one or more subjects 120 and provide the visual content to one or more participant devices 102 via the network 150. In the illustrated embodiment, the capture device 100 is a handheld device, such as a smart phone or tablet computing device. The user of the capture device 100 may launch specialized application software, such as a conference and chat subsystem 110 configured to provide various collaboration functionalities described herein, such as facilitating live virtual conferences and asynchronous chat conversations); launching, via the software communications platform, the video call that comprises the end user device and the agent device (see Bartov Paragraph [0059], With reference to an illustrative embodiment, a collaboration conversation participant may operate a user device, such as capture device 100, to capture visual content of one or more subjects 120 and provide the visual content to one or more participant devices 102 via the network 150. In the illustrated embodiment, the capture device 100 is a handheld device, such as a smart phone or tablet computing device. The user of the capture device 100 may launch specialized application software, such as a conference and chat subsystem 110 configured to provide various collaboration functionalities described herein, such as facilitating live virtual conferences and asynchronous chat conversations, Paragraph [0089], FIG. 6 is a flow diagram of an illustrative routine 600 for managing aspects of a live collaboration session based on video content captured by a capture device 100. A live collaboration session may also be referred to as a live conference. The capture device 100 may capture or otherwise generate full resolution content (e.g., content at a highest resolution provided by a camera of the capture device 100), and a user of the capture device 100 may desire to share the content with one or more participant devices); providing the agent device with control over a camera of the end user device in a representation of the video call via a graphical user interface (GUI) of the software communications platform (see Bartov Paragraph [0060], The participant devices 102 may each also include specialized application software, such as the conference and chat subsystem 110 configured to provide various collaboration functionalities described herein. The conference and chat subsystem 110 may allow the participant devices 102 to present and interact with content from a capture device 100, communicate with other participant devices 102, and the like. Some or all of the participant devices may also include a control subsystem 114 that facilities remote control, by a participant device, of the capture device 100, Paragraph [0062], FIG. 2 illustrates components of—and interactions between—various devices that may implement various features described herein, including a capture device 100 and a control device 200. The control device 200 may be a particular participant device 102 that is configured to control—or is in the process of controlling—aspects of content capture and other functionality of the capture device 10, Paragraph [0063], As shown in FIG. 2, a capture device 100 may have a camera 220 to capture visual content such as video or images, a data store 224 to store content generated by the camera 220, a content streamer 210 to manage sending live stream content to other devices participating in a live collaboration session, a remote camera command processor 212 to process and apply camera control commands received from other devices participating in a live collaboration session, an orientation prompt and compliance monitor 214 to process device reorientation commands and determine reorientation compliance, a high-quality content provider 216 to provide high quality content (e.g., full resolution content 222 as generated by the camera 220) in response to requests from other devices participating in a collaboration session (e.g., control device 200, a participant device 102), and a communication and chat viewer 218 to facilitate other communications and interactions, including communications and interactions outside of a live collaboration session, Paragraph [0064], As shown in FIG. 2, a control device 200 may include a content previewer 250 to provide a view of live stream content received from the capture device 100, a camera control user interface (UI) processor 252 to respond to various UI control interactions and send camera control commands to a capture device 100, an orientation processor 254 to respond to various movements of the control device 200 and send device reorientation commands to the capture device 100, a communication and chat viewer 256 to facilitate other communications and interactions, including receipt of high quality content that is then stored in a local data store 260, and an annotation and casting processor 258 to facilitate generation of annotations and communication of the annotations to other devices, Paragraph [0067], At [A], the capture device 100 may transmit visual content to the control device 200. The visual content may be transmitted as a live stream of content that is available for presentation by the control device 200 in real time, as the content is generated by the capture device 100, Paragraph [0071], As shown in FIGS. 3 and 4, the live stream presented on the display 330 of the control device 200 reflects the content captured and presented on the display 130 of the capture device 100. In this example, the live stream is a representation of a field of view of a camera of the capture device 100, which includes a human subject 302 partially outside the field of view, Paragraph [0072], At [B], a user of the control device 200 may determine that it is desirable to reorient the capture device 100 to adjust what is visible in the field of view of the camera of the capture device 100 and therefore in the live video presented on the control device 200. Instead of—or in addition to—verbalizing commands to reorient the capture device 100, the user of the control device 200 may cause the control device 200 to enter a camera guidance mode in which motion of the control device 200 is sensed, and device movement data is generated. For example, the user of the control device 200 may activate a user interface option to enter the camera guidance mode, and deactivate the user interface option when the user is finished remotely controlling the movement of the capture device 100, Paragraph [0081], FIG. 5 illustrates additional or alternative commands that may be initiated remotely by a control device 200 to alter the capture or generation of video content by the capture device 100 over a period of time, as indicated by the timeline. As shown, at [A] the capture device 100 may provide substantially live stream video content to the control device 200, which presents the video content on a display of the control device 200. In some embodiments, the control device 200 may also present a set of user interface controls including, but not limited to, exposure setting, zoom, color temperature, flash on/off, focus point, other capture parameters, or any combination thereof); and capturing, in the representation of the video call presented via the software communications platform, annotation information for one or more objects viewable via the camera based on a receipt of one or more actions initiated by the agent device that are received via the GUI of the software communications platform while the agent device has control over the camera of the end user device (see Bartov Paragraph [0052], To maintain separation between the annotations and the underlying content such that the annotations are non-destructive, the annotations may be defined or referenced in annotation metadata that may be provided with the underlying content while remaining physically separate (e.g., in a separate physical file, in a header of the content file separated from the content itself, etc.). Annotation metadata may represent coordinates, colors, shapes, and other properties of on-screen annotations and the timestamps at which they are to be displayed, coordinates and content of typed annotations and the timestamps at which they are to be displayed, audio and the timestamps at which it is the be presented, other metadata, or any combination thereof as desired. Annotations may be generated on the capture device itself, or on another device after being shared with the other device. Advantageously, the non-destructive nature of the annotations can allow recipients to view and interact with the annotated content as it was generated by the annotating user, while also allowing recipients to access the underlying non-annotated content to view it unannotated, add new annotations, alter the annotations generated by the first annotating user, or any combination thereof, Paragraph [0085], In some embodiments, instead of or in addition to providing remote control of camera functions, a user of the control device 200 may markup or otherwise modify the presentation of content to create a live annotation. The markup or other modifications may be applied to the video content on a non-destructive basis. For example, markups may be saved as metadata that may be dynamically applied to the presentation of the video content without permanently altering the video content itself, Paragraph [0086], In the example shown in FIG. 5, a user of the control device 200 may add a markup 510 to the video content at [F]. The markup 510 may be drawing (e.g., drawn with a stylus or a finger), typed textual content, or another visual augmentation added to the display of the video content. The live annotation(s) may be provided to the capture device 100 (and, in some cases, one or more other participant devices) at [G]. For example, to preserve the underlying content and not alter it on a permanent basis, annotation metadata regarding the annotations (e.g., identifications of frames or video portions to be presented, degrees of zoom to be applied, coordinates of viewports to be displayed, vector graphics instructions for drawn or otherwise added annotations, timestamps at which annotations or other display aspects are to be presented, etc.) may be generated and provided by the control device 200 to the capture device 100 (and, in some cases, other participant devices)). Regarding Claim 5, Bartov teaches The computer-implemented method of claim 1, wherein the providing comprises providing, in the representation of the video call via the GUI of the software communications platform, GUI control elements enabling control over the camera of the end user device by the agent device (see Bartov Paragraph [0064], As shown in FIG. 2, a control device 200 may include a content previewer 250 to provide a view of live stream content received from the capture device 100, a camera control user interface (UI) processor 252 to respond to various UI control interactions and send camera control commands to a capture device 100, Paragraph [0072], For example, the user of the control device 200 may activate a user interface option to enter the camera guidance mode, and deactivate the user interface option when the user is finished remotely controlling the movement of the capture device 100), and wherein the capturing further comprises receiving data indicating a selection of one or more GUI control elements by the agent device in the representation of the video call (see Bartov Paragraph [0083], a user may activate one or more of the user interface controls to adjust capture parameters of the capture device 100. For example, the user may tap on a corresponding icon. The control device 200 can send one or more commands regarding setting or modifying capture parameters to the capture device 100. For example, a command may have a name or include an identifier of the particular capture parameter to be set, the value or selection to which the capture parameter is to be set, other information, or a combination thereof), and generating the annotation information based on an analysis of the data indicating the selection of the one or more GUI control elements by the agent device (see Bartov Paragraph [0077], the capture device 100 may process the device movement data and present a prompt to the user of the capture device 100. The capture device 100 may present the prompt visually, audibly, via some other modality, or in a combination of modalities. In some embodiments, if the device movement data represents motion in three-dimensional space to be applied to the capture device 100, the capture device 100 may convert the device movement data into a visual representation for display on the display 130 of the capture device 100. In the illustrated example, the capture device 100 presents a prompt 310 in the form of an arrow indicating the direction in which the capture device 100 is to be moved). Regarding Claim 6, Bartov teaches The computer-implemented method of claim 5, wherein the GUI control elements are configured to enable panning of a viewable area of the camera including the one or more objects, and wherein the capturing further comprises receiving data indicating a selection of one or more GUI control elements for adjusting an angle and a direction of a view of the camera relative to the one or more objects (see Bartov Paragraph [0064], As shown in FIG. 2, a control device 200 may include a content previewer 250 to provide a view of live stream content received from the capture device 100, a camera control user interface (UI) processor 252 to respond to various UI control interactions and send camera control commands to a capture device 100, Paragraph [0072], For example, the user of the control device 200 may activate a user interface option to enter the camera guidance mode, and deactivate the user interface option when the user is finished remotely controlling the movement of the capture device 100, Paragraph [0050], Additional aspects of the present disclosure relate to providing recipients of video content with the capability to remotely control aspects of content capture. In some embodiments, a recipient may physically change the orientation of—or otherwise manipulate—the device on which they are viewing a live video stream. An instruction to perform a corresponding reorientation or manipulation may be presented on the capturing device. For example, a recipient may wish to pan in a particular direction to adjust the field of view captured in the video stream. When the recipient performs such a movement with the recipient device, the recipient device may generate and send an instruction to the capture device regarding the desired movement to be made with the capture device. The capture device may then present, to a user of the capture device, a prompt regarding the desired movement, such as by displaying an arrow in the direction in which the capture device is to be moved or reoriented. In addition to the particular type of desired movement, data regarding the magnitude of the movement may also be sent to the capture device so that when the capture device has been adjusted to provide the desired field of view, an indication may be presented to the user of the capture device. For example, haptic feedback may be provided, a visual prompt to move the capture device may be removed from the display screen, or a visual confirmation may be displayed. In this way, the user of the capture device can be prompted to perform a particular type and magnitude of action, and be informed when the desired action has been completed (e.g., so the user doesn't overmanipulate the device), Paragraph [0081], In some embodiments, the control device 200 may also present a set of user interface controls including, but not limited to, exposure setting, zoom, color temperature, flash on/off, focus point, other capture parameters, or any combination thereof). Regarding Claim 7, Bartov teaches The computer-implemented method of claim 6, wherein the GUI control elements further comprise GUI elements control configured to enable the agent device, via the representation of the video call, to generate one or more of: a screenshot of the one or more objects, a video clip of the one or more objects; a description tag for the one or more objects, and a geo-locational marker for the one or more objects (see Bartov Paragraph [0064], As shown in FIG. 2, a control device 200 may include a content previewer 250 to provide a view of live stream content received from the capture device 100, a camera control user interface (UI) processor 252 to respond to various UI control interactions and send camera control commands to a capture device 100, Paragraph [0072], For example, the user of the control device 200 may activate a user interface option to enter the camera guidance mode, and deactivate the user interface option when the user is finished remotely controlling the movement of the capture device 100, Paragraph [0081], In some embodiments, the control device 200 may also present a set of user interface controls including, but not limited to, exposure setting, zoom, color temperature, flash on/off, focus point, other capture parameters, or any combination thereof, Paragraph [0102], For example, the user may activate one or more user interface controls to zoom, rotate, edit, manage playback, create annotations (e.g., snapshots, narrations), move or position a cursor 702, and the like, Paragraph [0113], a user may choose to generate the annotation as a snapshot. To generate a snapshot, the device on which the snapshot is being generated may create an annotation file 806 linked to the original media file 804 that can be used to provide an annotated image. The annotation file 806 may include information to reproduce the current view of the media and its overlaid annotation. For example, the annotation file 806 may include or define one or more of the following: an identifier of the image that is to serve as the underlying content for the annotation, an identifier of the current frame in the case that the underlying media file 804 is a video media file (in which case the snapshot functionality is only available when the video is paused); the pan and zoom applied to the original image or video frame (viewport); and all the annotation overlays the user added (e.g., defined as vector graphics), Paragraph [0051], Some annotations are composed of a single static image (e.g., a frame of video) to which various markups or modifications have been applied. Such annotations may be referred to as “snapshots.” Some annotations are dynamic in the sense that presentation (e.g., audio and/or video output) changes over a period of time from start to end. Such annotations may be referred to as “narrations” or “video-based annotations.” For example, a video-based annotation may be composed of one or more images, portions of video, or a combination thereof presented in a predetermined user-defined sequence. Video-based annotations may also include markups, viewing manipulations, audio tracks (e.g., user speech), and the like, Paragraph [0052], Annotation metadata may represent coordinates, colors, shapes, and other properties of on-screen annotations and the timestamps at which they are to be displayed, coordinates and content of typed annotations and the timestamps at which they are to be displayed). Regarding Claims 11 and 15 - 17, they are rejected similarly as Claims 1 and 5 - 7, respectively. The system can be found in Bartov (Paragraph [0195], system). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2, 3, 12 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Bartov et al. (U.S. Pub. No. 2025/0097569, hereinafter “Bartov”) in view of Rathod (U.S. Pub. No. 2025/0379933). Regarding Claim 2, Bartov teaches all the limitations of claim 1, but does not expressively teach The computer-implemented method of claim 1, wherein the first state of the electronic communication is a voice call between the end user device and the agent device. However, Rathod teaches The computer-implemented method of claim 1, wherein the first state of the electronic communication is a voice call between the end user device and the agent device (see Rathod Paragraph [0391], users can make phone calls to a particular contact's selected phone number and can switch phone calls during calling or during establishment of phone call session to text call by tapping or clicking on text call icon 2335. In another embodiment user can make text call to particular contact's selected phone number and can switch text call to phone call 2330 by tapping or clicking on phone call icon 2303 or switch to video call 2332 or switch to SMS 2334 or switch to default or pre-set instant messenger application, Paragraph [0488], Agent [John] can switch camera to front camera or back camera via icon 4931 and also can select or set or switch communication medium including change video chatting to voice chatting or select voice 4933 to provide response via voice or voice message or change video chatting to text chatting or select message (instant message or SMS) 4934 to provide response via text message including instant messaging or short message service (SMS), Paragraph [0556], In another embodiment user can select or switch one or more communication mediums or channels or interfaces for live chat or live communicating with available agent of user selected or user visited particular website, wherein communication mediums or channels or interfaces comprises chat, video, voice or VOIP, phone or video call and SMS communication. In an embodiment when user select chat 5691 then user can send and receive messages with connected and available agent of selected or currently viewing website. In an embodiment when user select video message 5692 then user can send and receive video messages with connected and available agent of selected or currently viewing website. In another embodiment if user select audio message 5693 then user can send and receive audio messages with connected and available agent of selected or currently viewing website. In another embodiment if user select voice call including cellular phone call or Voice over Internet Protocol (VoIP) 5694 then user anonymously connect and converse with currently available agent of selected or currently viewing website. In another embodiment if user select video call 5695 including cellular video call or video call over internet protocol then user anonymously connect and converse with currently available agent of selected or currently viewing website, Paragraph [0573], user can select or switch one or more communication mediums or channels or interfaces for live chat or live communicating with contacts, wherein communication mediums or channels or interfaces comprises chat 5891/5991, video communication or messages or call 5892/5895/5992/5995, voice or VOIP or phone communication or messages or call 5893/5894/5993/5994 and SMS 5896/5996 communication). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a system for remote video support where an agent can control a user’s camera during a video call and annotate the live camera feed (as taught in Bartov), with transitioning from a voice call to a video call (as taught in Rathod), the motivation being to provide faster, clearer and more effective communication by allowing a continuous and seamless conversation while adding visual context (see Rathod Paragraph [0488]). Regarding Claim 3, Bartov in view of Rathod teaches The computer-implemented method of claim 1, wherein the electronic communication is between the end user device and the agent device, and the first state of the electronic communication is one of an electronic chat, an electronic message, or an email (see Rathod Paragraph [0391], users can make phone calls to a particular contact's selected phone number and can switch phone calls during calling or during establishment of phone call session to text call by tapping or clicking on text call icon 2335. In another embodiment user can make text call to particular contact's selected phone number and can switch text call to phone call 2330 by tapping or clicking on phone call icon 2303 or switch to video call 2332 or switch to SMS 2334 or switch to default or pre-set instant messenger application, Paragraph [0488], Agent [John] can switch camera to front camera or back camera via icon 4931 and also can select or set or switch communication medium including change video chatting to voice chatting or select voice 4933 to provide response via voice or voice message or change video chatting to text chatting or select message (instant message or SMS) 4934 to provide response via text message including instant messaging or short message service (SMS), Paragraph [0556], In another embodiment user can select or switch one or more communication mediums or channels or interfaces for live chat or live communicating with available agent of user selected or user visited particular website, wherein communication mediums or channels or interfaces comprises chat, video, voice or VOIP, phone or video call and SMS communication. In an embodiment when user select chat 5691 then user can send and receive messages with connected and available agent of selected or currently viewing website. In an embodiment when user select video message 5692 then user can send and receive video messages with connected and available agent of selected or currently viewing website. In another embodiment if user select audio message 5693 then user can send and receive audio messages with connected and available agent of selected or currently viewing website. In another embodiment if user select voice call including cellular phone call or Voice over Internet Protocol (VoIP) 5694 then user anonymously connect and converse with currently available agent of selected or currently viewing website. In another embodiment if user select video call 5695 including cellular video call or video call over internet protocol then user anonymously connect and converse with currently available agent of selected or currently viewing website, Paragraph [0573], user can select or switch one or more communication mediums or channels or interfaces for live chat or live communicating with contacts, wherein communication mediums or channels or interfaces comprises chat 5891/5991, video communication or messages or call 5892/5895/5992/5995, voice or VOIP or phone communication or messages or call 5893/5894/5993/5994 and SMS 5896/5996 communication). Regarding Claims 12 and 13, they are rejected similarly as Claims 2 and 3, respectively. The system can be found in Bartov (Paragraph [0195], system). Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Bartov et al. (U.S. Pub. No. 2025/0097569, hereinafter “Bartov”) in view of Mann et al. (U.S. Pub. No. 2018/0013982, hereinafter “Mann”). Regarding Claim 4, Bartov teaches all the limitations of claim 1, but does not expressively teach The computer-implemented method of claim 1, wherein the providing of the agent device with control over the camera of the end user device during the video call comprises transmitting a permission request to the end user device to enable the end user device to accept a permission for the agent device to control the camera during the video call, in response to receiving acceptance of the permission, displaying a view from the camera in the representation of the video call and activating, in the representation of the video call, GUI feature functionality for the agent device to control the camera. However, Mann teaches The computer-implemented method of claim 1, wherein the providing of the agent device with control over the camera of the end user device during the video call comprises transmitting a permission request to the end user device to enable the end user device to accept a permission for the agent device to control the camera during the video call, in response to receiving acceptance of the permission, displaying a view from the camera in the representation of the video call and activating, in the representation of the video call, GUI feature functionality for the agent device to control the camera (see Mann Paragraph [0016], In addition to requesting access to a video footage from a video camera, a remote participant may request to have physical remote control of the video camera itself. For example, if granted by the computing device 110, the remote participant may choose to control the direction, zoom, contrast, or other setting of a video camera, Paragraph [0014], As a result, instead of transmitting the same multimedia stream to the remote participants 120a-120n, each remote participant has control over the audio and/or video streams they receive on their respective devices. As will be further described, control of the media transfers predominantly to the hands of each remote participant, allowing the remote participants 120a-120n to better control their respective experiences and, thus, their participation in the meeting in the collaborative workspace, Paragraph [0023], the computing device 110 may perform the algorithmic work on manipulating the various audio and/or video feeds to optimize the experience for each remote participant, and permit only a single mixed audio and/or video stream to be transmitted to each remote participant, Paragraph [0030], At 310, the computing device (e.g., computing device 110) receives streams of content (e.g., audio and/or video feeds) from devices located in the conference room. Referring back to FIG. 1, the computing device 110 may receive audio feeds from the microphones 102a-102f and video feeds from the video cameras 103a-103c, Paragraph [0031], At 320, the computing device receives requests from users (e.g., participants remote from the conference room). As an example, each request is for accessing a subset of the streams of content (e.g., the audio and/or video feeds), based on participants in the room that each user desires to follow. As a result, instead of transmitting the same multimedia stream to all users remote from the room, each user has control over the audio and/or video streams they receive on their respective devices. However, the requests from users to access streams of content from devices located in the room may be overridden by participants in the room or an administrator. As an example, each user request for accessing a subset of the streams of content may correspond to streams of content focused on a participant in the room. As a result, the streams of content accessible to the user may change based on movement of the participant in the room, Paragraph [0032], In addition to receiving requests for accessing the streams of content, the computing device may receive requests from users to control devices located in the room. As an example, the computing device may grant a request by a first user to control devices located in the room that correspond to streams of content requested by the first user for access, Paragraph [0033], At 330, the computing device, for each user request, joins the requested subset of streams of content into a single multimedia stream. For the participants in the room that each user desires to follow, the computing device may perform the algorithmic work on manipulating the various audio and/or video feeds to optimize the experience for each user, and permit only a single mixed audio and/or video stream to be transmitted to each user. By consolidating the various feeds into a single multimedia stream, and transmitting only the feeds requested by a user, the bandwidth required for transmitting the feeds to the user may be reduced. At 340, the computing device transmits each single multimedia stream to devices of respective users. As a result, any number of individually customized feeds may be sent to the users remote from the room). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a system for remote video support where an agent can control a user’s camera during a video call and annotate the live camera feed (as taught in Bartov), with transmitting a permission request to control another user’s camera (as taught in Mann), the motivation being to improve user privacy by requiring explicit permission prior to a user controlling a component of another person’s device (see Mann Paragraph [0022]). Regarding Claim 14, it is rejected similarly as Claim 4. The system can be found in Bartov (Paragraph [0195], system). Claims 8 – 10 and 18 - 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bartov et al. (U.S. Pub. No. 2025/0097569, hereinafter “Bartov”) in view of Salowitz et al. (U.S. Pub. No. 2025/0005072, hereinafter “Salowitz”). Regarding Claim 8, Bartov teaches all the limitations of claim 5, but does not expressively teach The computer-implemented method of claim 5, further comprising: applying one or more trained artificial intelligence models to generate one or more data insights for the agent device, including suggestions for capture of the annotation information, and providing the agent device with the one or more data insights, wherein to generate the one or more data insights, the one or more trained artificial intelligence models are trained to analyze receipt of action initiated by the agent device during the video call, and additional contextual information comprising one or more of: a transcription of the video call between the end user device and the agent device, and historical contextual information from previous use of the communications software platform by an end user entity associated with the end user device, and wherein the receiving of the data indicating a selection of one or more GUI control elements by the agent device occurs after the data insights are provided to the agent device. However, Salowitz teaches The computer-implemented method of claim 5, further comprising: applying one or more trained artificial intelligence models to generate one or more data insights for the agent device, including suggestions for capture of the annotation information, and providing the agent device with the one or more data insights, wherein to generate the one or more data insights, the one or more trained artificial intelligence models are trained to analyze receipt of action initiated by the agent device during the video call, and additional contextual information comprising one or more of: a transcription of the video call between the end user device and the agent device, and historical contextual information from previous use of the communications software platform by an end user entity associated with the end user device, and wherein the receiving of the data indicating a selection of one or more GUI control elements by the agent device occurs after the data insights are provided to the agent device (see Salowitz Paragraph [0032], In addition to identifying documents, the query may also identify a previous meeting as being relevant to the current meeting. This determination may be based on the meetings having a shared topic. In some configurations, a shared topic may be determined based on an analysis of a transcript of the previous meeting and an analysis of a transcript of the current meeting, although participants, title, shared screen content, time of day, and other factors may also affect whether embedding vectors of the two meetings are close enough in the embedding space to be relevant. For example, the transcript of the previous meeting may have included a conversation in which one participant promised to provide a document to another participant. The content suggestion engine may remind the user of this promise, or even propose a document that fulfils the promise, Paragraph [0036], In some configurations, interaction embeddings may be analyzed to identify patterns in user behavior. These patterns may be used to suggest documents, websites, meetings, tasks, and the like, Paragraph [0045], Chat interface 150 enables a conversational or chatbot style interaction with AI explorer 102. In some configurations, chat interface 150 is integrated into search bar 110 or vice-versa. A user may supply prompt 152 to chat interface 150. Prompt 152 may include text that is provided to a machine learning model, such as a large language model or multi-modal generative model. Prompt 152 may be augmented with additional information derived from the current context, such as the applications that are currently open, conversations or meetings that are currently active and their participants, documents that are open, content that is visible on the screen, etc. The output generated by the machine learning model may be displayed inline in the chat interface. Additionally, or alternatively, responses from the machine learning model may be used to generate user interface components that respond to the prompt, such as displaying a list of files, a list of applications, a list of people, or other suggestions that are particular to the user interface of a computing device, Paragraph [0024], Interaction data represents what the computing device was receiving as input or generating as output, such as a screenshot, an audio stream, and/or user input events such as key presses, mouse movements, voice commands, gestures, and/or any other suitable user input. Interaction data may be generated during any type of user task, such as browsing the web, participating in a meeting, playing a game, authoring a document, etc. Screenshots that capture user interaction data may be taken continuously, periodically, or at particular points in time. Pieces of interaction data are stored as entries in a timeline, which maintains a history of user interactions with the computing device, Paragraph [0074], a user operating the online meeting 400 may initiate a discovery mode that proposes suggested operations for the current context). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a system for remote video support where an agent can control a user’s camera during a video call and annotate the live camera feed (as taught in Bartov), with using an AI model to analyze video call activity, transcriptions, and historical context to generate and display contextual insights and recommendations to a user in a video call (as taught in Salowitz), the motivation being to provide a system that proactively provides suggestions to a user to help them, without having to make an explicit request (see Salowitz Paragraphs [0022] – [0023]). Regarding Claim 9, Bartov in view of Salowitz teaches The computer-implemented method of claim 5, wherein the capturing further comprises applying one or more trained artificial intelligence models to generate the annotation information based on the receipt of one or more actions initiated by the agent device, during the video call, via the GUI of the software communications platform (see Salowitz Paragraph [0032], In addition to identifying documents, the query may also identify a previous meeting as being relevant to the current meeting. This determination may be based on the meetings having a shared topic. In some configurations, a shared topic may be determined based on an analysis of a transcript of the previous meeting and an analysis of a transcript of the current meeting, although participants, title, shared screen content, time of day, and other factors may also affect whether embedding vectors of the two meetings are close enough in the embedding space to be relevant. For example, the transcript of the previous meeting may have included a conversation in which one participant promised to provide a document to another participant. The content suggestion engine may remind the user of this promise, or even propose a document that fulfils the promise, Paragraph [0036], In some configurations, interaction embeddings may be analyzed to identify patterns in user behavior. These patterns may be used to suggest documents, websites, meetings, tasks, and the like, Paragraph [0045], Chat interface 150 enables a conversational or chatbot style interaction with AI explorer 102. In some configurations, chat interface 150 is integrated into search bar 110 or vice-versa. A user may supply prompt 152 to chat interface 150. Prompt 152 may include text that is provided to a machine learning model, such as a large language model or multi-modal generative model. Prompt 152 may be augmented with additional information derived from the current context, such as the applications that are currently open, conversations or meetings that are currently active and their participants, documents that are open, content that is visible on the screen, etc. The output generated by the machine learning model may be displayed inline in the chat interface. Additionally, or alternatively, responses from the machine learning model may be used to generate user interface components that respond to the prompt, such as displaying a list of files, a list of applications, a list of people, or other suggestions that are particular to the user interface of a computing device, Paragraph [0024], Interaction data represents what the computing device was receiving as input or generating as output, such as a screenshot, an audio stream, and/or user input events such as key presses, mouse movements, voice commands, gestures, and/or any other suitable user input. Interaction data may be generated during any type of user task, such as browsing the web, participating in a meeting, playing a game, authoring a document, etc. Screenshots that capture user interaction data may be taken continuously, periodically, or at particular points in time. Pieces of interaction data are stored as entries in a timeline, which maintains a history of user interactions with the computing device, Paragraph [0074], a user operating the online meeting 400 may initiate a discovery mode that proposes suggested operations for the current context). Regarding Claim 10, Bartov in view of Salowitz teaches The computer-implemented method of claim 9, wherein the one or more trained artificial intelligence models are further trained to evaluate additional contextual information provided via the software communications platform to generate the annotation information, and wherein the additional contextual information comprises a transcription of the video call between the end user device and the agent device, or historical contextual information from previous use of the communications software platform by an end user entity associated with the end user device (see Salowitz Paragraph [0032], In addition to identifying documents, the query may also identify a previous meeting as being relevant to the current meeting. This determination may be based on the meetings having a shared topic. In some configurations, a shared topic may be determined based on an analysis of a transcript of the previous meeting and an analysis of a transcript of the current meeting, although participants, title, shared screen content, time of day, and other factors may also affect whether embedding vectors of the two meetings are close enough in the embedding space to be relevant. For example, the transcript of the previous meeting may have included a conversation in which one participant promised to provide a document to another participant. The content suggestion engine may remind the user of this promise, or even propose a document that fulfils the promise, Paragraph [0036], In some configurations, interaction embeddings may be analyzed to identify patterns in user behavior. These patterns may be used to suggest documents, websites, meetings, tasks, and the like, Paragraph [0045], Chat interface 150 enables a conversational or chatbot style interaction with AI explorer 102. In some configurations, chat interface 150 is integrated into search bar 110 or vice-versa. A user may supply prompt 152 to chat interface 150. Prompt 152 may include text that is provided to a machine learning model, such as a large language model or multi-modal generative model. Prompt 152 may be augmented with additional information derived from the current context, such as the applications that are currently open, conversations or meetings that are currently active and their participants, documents that are open, content that is visible on the screen, etc. The output generated by the machine learning model may be displayed inline in the chat interface. Additionally, or alternatively, responses from the machine learning model may be used to generate user interface components that respond to the prompt, such as displaying a list of files, a list of applications, a list of people, or other suggestions that are particular to the user interface of a computing device, Paragraph [0024], Interaction data represents what the computing device was receiving as input or generating as output, such as a screenshot, an audio stream, and/or user input events such as key presses, mouse movements, voice commands, gestures, and/or any other suitable user input. Interaction data may be generated during any type of user task, such as browsing the web, participating in a meeting, playing a game, authoring a document, etc. Screenshots that capture user interaction data may be taken continuously, periodically, or at particular points in time. Pieces of interaction data are stored as entries in a timeline, which maintains a history of user interactions with the computing device, Paragraph [0074], a user operating the online meeting 400 may initiate a discovery mode that proposes suggested operations for the current context). Regarding Claims 18 - 20, they are rejected similarly as Claims 8 - 10, respectively. The system can be found in Bartov (Paragraph [0195], system). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARISSA A JONES whose telephone number is (703)756-1677. The examiner can normally be reached Telework M-F 6:30 AM - 4:00 PM CT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CARISSA A JONES/ Examiner, Art Unit 2691 /DUC NGUYEN/ Supervisory Patent Examiner, Art Unit 2691
Read full office action

Prosecution Timeline

Feb 03, 2025
Application Filed
Jul 27, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699454
TERMINAL APPARATUS, COMMUNICATION SYSTEM
2y 11m to grant Granted Aug 04, 2026
Patent 12666140
CONTRIBUTION-BASED CLOSE-UP CONTROL
2y 3m to grant Granted Jun 23, 2026
Patent 12656991
SYSTEMS, METHODS, APPARATUSES, AND DEVICES FOR MANAGING A GAZE OF A USER ENGAGED IN A VISUAL COMMUNICATION SESSION WITH AN AUDIENCE
2y 11m to grant Granted Jun 16, 2026
Patent 12657863
System and Method for Fewer or No Non-Participant Framing and Tracking
3y 2m to grant Granted Jun 16, 2026
Patent 12659596
METHODS AND SYSTEMS FOR REDUCING SCREEN REFLECTIONS ON EYEGLASSES IN LIVE STREAMING
2y 6m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+30.0%)
2y 7m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 30 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month