Prosecution Insights
Last updated: October 02, 2026
Application No. 18/850,927

AUDIENCE CONFIGURATIONS OF AUDIOVISUAL SIGNALS

Non-Final OA §103§112
Filed
Sep 25, 2024
Priority
Apr 01, 2022 — nonprovisional of PCTUS2022023094
Examiner
JONES, CARISSA ANNE
Art Unit
2691
Tech Center
2600 — Communications
Assignee
Hewlett-Packard Development Company, L.P.
OA Round
2 (Non-Final)
77%
Grant Probability
Favorable
2-3
OA Rounds
7m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
27 granted / 35 resolved
+15.1% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
19 currently pending
Career history
62
Total Applications
across all art units

Statute-Specific Performance

§101
2.8%
-37.2% vs TC avg
§103
80.0%
+40.0% vs TC avg
§102
11.2%
-28.8% vs TC avg
§112
3.7%
-36.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 35 resolved cases

Office Action

§103 §112
DETAILED ACTION This action is in response to the remarks filed 06/25/2026. Claims 1 - 21 are pending and have been examined. Claim 4 has been cancelled. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments with respect to claims 1, 6 and 11 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Response to Amendment Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claim 1 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The claim recites “the second enhanced region”; however, there is no antecedent basis for “the second enhanced region” in the claim. The claim only previously recites “a second region”, and discusses generating a first enhanced region, but not “the second enhanced region”. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 2, 16, 20 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Ong et al. (U.S. Patent No. 11,165,992, hereinafter “Ong”) in view of O’Leary et al. (U.S. Pub. No. 2022/0247919, hereinafter “O’Leary”). Regarding Claim 1, Ong teaches An electronic device (see Ong Abstract, device), comprising: an image sensor (see Ong Figure 1, camera); and a controller (see Ong Column 1, lines 49 – 54, A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to generate a composited video layout of facial images in a video conference) to: receive a first audiovisual signal via the image sensor, the first audiovisual signal depicting multiple audience members (see Ong Column 1, lines 56 – 59, The computer-implemented method includes receiving a video frame from a video source, where the video frame includes faces of individuals engaged in a video conference, and Column 2, lines 66 – 67 and Column 3, lines 1 – 2, The single camera is typically positioned to encompass a wide enough field of view to capture video of multiple individuals participating in the conference); determine a plurality of regions of the first audiovisual signal that depict the multiple audience members (see Ong Column 4, lines 21 – 33, The video layout system 118 may also include a multi-face discovery engine 122. In certain embodiments, the multi-face discovery engine 122 is configured to detect faces in one or more video frames in the video frame buffer 120. Images of the detected faces may be used by the multi-face discovery engine 122 to generate facial images for each participant detected in the video frame. In certain embodiments, the facial images generated by the multi-face discovery engine 122 are provided to a windowed image generation engine 124. The windowed image generation engine 124 may be configured to generate a windowed image for each facial image generated by the multi-face discovery engine 122), wherein a first region of the plurality of regions depicts a first audience member and a second region of the plurality of regions depicts a second audience member (see Ong Column 5, lines 65 – 67, and Column 6, line 1, multi-face layout engine 404 is used to generate a composited video frame 406 that includes the scaled facial image 318, 320, and 322 in respective window images 408, 410, and 412, in which a first participant 318 is depicted in first region 408, second participant 320 is depicted in second region 410, and third participant 322 is depicted in third region 412); Ong does not expressively teach enhance the first region of the first audience member to generate a first enhanced region, wherein enhancing the first region includes warping a perspective of the first region, increasing an angle of vision of the first audience member, or adjusting a resolution of the first region; generate a second audiovisual signal that includes the first enhanced region and the second enhanced region; and cause display, transmission, or a combination thereof, of the second audiovisual signal. However, O’Leary teaches enhance the first region of the first audience member to generate a first enhanced region, wherein enhancing the first region includes warping a perspective of the first region, increasing an angle of vision of the first audience member, or adjusting a resolution of the first region (see O’Leary Paragraph [0315], In FIG. 8I, device 600 detects input 832 on multi-person option 830-2. In response, device 600 switches from the single-person framing mode setting to the multi-person framing mode setting. When the multi-person framing mode setting is enabled, device 600 automatically adjusts the video feed field-of-view to include additional subjects detected in field-of-view 620 (or a subset thereof). For example, when device 600 switches to the multi-person framing mode setting, device 600 expands the video feed field-of-view to include representations 628-1 and 622-1 of both Jack and Jane, as depicted in camera preview 606 in FIG. 8J. Accordingly, portion 625 in FIG. 8J represents the expanded video feed field-of-view resulting from enabling the multi-person framing mode setting); generate a second audiovisual signal that includes the first enhanced region and the second enhanced region (see O’Leary Paragraph [0372], In FIG. 10C, device 600 detects that Jane 622 has moved away from Jack 628 (she is separate from Jack by a threshold amount). In response, device 600 transitions the outputted video feed field-of-view from the continuous field-of-view depicted in FIG. 10B to a split field-of-view, as depicted in FIG. 10C. For example, device 600 displays the camera preview with first preview portion 1006-1 separated from second preview portion 1006-2 by line 1015. In some embodiments, line 1015 is depicted in video conference interface 1004, but is not included in the output video feed field-of-view. In some embodiments, line 1015 is included in the output video feed field-of-view); and cause display, transmission, or a combination thereof, of the second audiovisual signal (see O’Leary Paragraph [0372], In FIG. 10C, device 600 detects that Jane 622 has moved away from Jack 628 (she is separate from Jack by a threshold amount). In response, device 600 transitions the outputted video feed field-of-view from the continuous field-of-view depicted in FIG. 10B to a split field-of-view, as depicted in FIG. 10C. For example, device 600 displays the camera preview with first preview portion 1006-1 separated from second preview portion 1006-2 by line 1015. In some embodiments, line 1015 is depicted in video conference interface 1004, but is not included in the output video feed field-of-view. In some embodiments, line 1015 is included in the output video feed field-of-view). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member (as taught in Ong), with enhancing each audience member’s region by warping its perspective, widening the viewing angle, or adjusting its resolution, then generating a second audiovisual signal containing the enhanced regions for display or transmission (as taught in O’Leary), the motivation being to provide an electronic device with a faster and more efficient way to manage a live video communication session by automatically identifying and enhancing individual audience members, reducing the user’s cognitive burden, and providing a more efficient human-machine interface (see O’Leary Paragraph [0005]). Regarding Claim 2, Ong in view of O’Leary teaches The electronic device of claim 1, wherein a region of the regions depicts a face of an audience member of the multiple audience members (see Ong Column 4, lines 21 – 33, The video layout system 118 may also include a multi-face discovery engine 122. In certain embodiments, the multi-face discovery engine 122 is configured to detect faces in one or more video frames in the video frame buffer 120. Images of the detected faces may be used by the multi-face discovery engine 122 to generate facial images for each participant detected in the video frame. In certain embodiments, the facial images generated by the multi-face discovery engine 122 are provided to a windowed image generation engine 124. The windowed image generation engine 124 may be configured to generate a windowed image for each facial image generated by the multi-face discovery engine 122). Regarding Claim 16, Ong in view of O’Leary teaches The electronic device of claim 1, wherein enhancing the first region of the first audience member includes increasing an angle of vision of the first audience member (see O’Leary Paragraph [0315], In FIG. 8I, device 600 detects input 832 on multi-person option 830-2. In response, device 600 switches from the single-person framing mode setting to the multi-person framing mode setting. When the multi-person framing mode setting is enabled, device 600 automatically adjusts the video feed field-of-view to include additional subjects detected in field-of-view 620 (or a subset thereof). For example, when device 600 switches to the multi-person framing mode setting, device 600 expands the video feed field-of-view to include representations 628-1 and 622-1 of both Jack and Jane, as depicted in camera preview 606 in FIG. 8J. Accordingly, portion 625 in FIG. 8J represents the expanded video feed field-of-view resulting from enabling the multi-person framing mode setting). Regarding Claim 20, Ong in view of O’Leary teaches The electronic device of claim 1, wherein the controller is further to: receive a request from an application to access the image sensor (see O’Leary Figure 12H, camera can be turned on to request access from a video conference application to access the camera); in response to receiving the request, receive the first audiovisual signal (see O’Leary Paragraph [0434], input 1250 on camera option 1213, in which turns camera on for transmission of video signal); before sharing the first audiovisual signal with the application, intercepting the first audiovisual signal to enhance the first region and the second region (see O’Leary Figure 10H, prior to transmission to the video conference application, participants are framed, and zoomed in to give each participant their own video stream); share the second audiovisual signal with the application (see O’Leary Figure 10H, participants are individually framed and displayed on video conference application on devices, in which they are cropped and zoomed into in order to provide separate outputs for each participant). Regarding Claim 21, Ong in view of O’Leary teaches The electronic device of claim 1, wherein the location of the first region in the first audiovisual signal relative to the location of the second region in the first audiovisual signal corresponds to the physical location of the first audience member relative to the second audience member in an environment that includes the image sensor (see O’Leary Paragraph [0372], In FIG. 10C, device 600 detects that Jane 622 has moved away from Jack 628 (she is separate from Jack by a threshold amount). In response, device 600 transitions the outputted video feed field-of-view from the continuous field-of-view depicted in FIG. 10B to a split field-of-view, as depicted in FIG. 10C. For example, device 600 displays the camera preview with first preview portion 1006-1 separated from second preview portion 1006-2 by line 1015. In some embodiments, line 1015 is depicted in video conference interface 1004, but is not included in the output video feed field-of-view. In some embodiments, line 1015 is included in the output video feed field-of-view). Claims 3, 5, 11 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Ong et al. (U.S. Patent No. 11,165,992, hereinafter “Ong”) in view of O’Leary et al. (U.S. Pub. No. 2022/0247919, hereinafter “O’Leary”) and Hoang et al. (U.S. Pub. No. 2023/0081717, hereinafter “Hoang”). Regarding Claim 3, Ong in view of O’Leary teach all the limitations of claim 1, but do not expressively teach The electronic device of claim 1, wherein a region of the regions includes a marker to indicate an audience member depicted within the region is speaking. However, Hoang teaches The electronic device of claim 1, wherein a region of the regions includes a marker to indicate an audience member depicted within the region is speaking (see Hoang Paragraph [0023], The size of a UI tile may depend on one or more factors including the view style set for the conferencing software UI at a given time and whether the one or more conference participants represented by the UI tile are active speakers at a given time. A speaker view in which one or more UI tiles for active speakers are enlarged and arranged in a center position of the conferencing software UI while the UI tiles for other conference participants are reduced in size and arranged near an edge of the conferencing software UI, Paragraph [0096], The subset 804 is arranged in a group so that all of the UI tiles associated with the conference participants within the conference room are grouped together and also in a specified order according to the determined relative locations of those conference participants. A large UI tile 806 represents an active speaker at a given time during the conference, in which the speaker’s video is displayed larger than others to visually indicate they are speaking). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member, enhancing each audience member’s region by warping its perspective, widening the viewing angle, or adjusting its resolution, then generating a second audiovisual signal containing the enhanced regions for display or transmission (as taught in Ong in view of O’Leary), with a participant’s video region including a marker to indicate the participant is speaking (as taught in Hoang), the motivation being to provide a video conference interface with a speaker front and center to improve focus and improve the clarity of the conference (see Hoang Paragraph [0019]). Regarding Claim 5, Ong in view of O’Leary and Hoang teaches The electronic device of claim 1, wherein a region of the regions of the second audiovisual signal depicts the first audiovisual signal (see Hoang Abstract, User interface (UI) tiles of conference participants are arranged in a UI of conferencing software according to relative locations of those conference participants within a conference room. Positional information and video data are obtained from one or more video capture devices located within a conference room. Relative locations of conference participants within the conference room are determined based on the positional information, such as by defining coordinates for the conference participants within a coordinate system based on the positional information and the video data and determining the relative locations based on the coordinates. Output configured to cause a client application to arrange UI tiles associated with the conference participants is generated according to the relative locations. The output is then transmitted to one or more client devices to cause a UI of conferencing software at each of those client devices to display the UI tiles in the specified arrangement, and Paragraph [0070], where a field of view of a video capture device 400 includes multiple conference participants, a stream of video data from that video capture device can be processed to determine regions of interest corresponding to those conference participants within the conference room 402 based on that video data. For example, the conference intelligence software 406 can include functionality for determining multiple regions of interest within a field of view of a video capture device 400 and for initializing output video streams for rendering within separate UI tiles of the conferencing software 408 for each of those regions of interest. The client application 410 then renders the output video streams within the respective UI tiles for viewing at the client device 412, in which the first image includes multiple participants, and the second image would include one of the multiple participants in the first image). Regarding Claim 11, Ong teaches A non-transitory machine-readable medium storing machine-readable instructions which, (see Ong Column 10, lines 4 – 7, computer-readable medium may be any medium that can contain, store, communicate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device) when executed by a controller of an electronic device, cause the controller (see Ong Column 1, lines 49 – 54, A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to generate a composited video layout of facial images in a video conference) to: receive a first audiovisual signal via an image sensor, the first audiovisual signal depicting a first audience member and a second audience member (see Ong Column 1, lines 56 – 59, The computer-implemented method includes receiving a video frame from a video source, where the video frame includes faces of individuals engaged in a video conference, and Column 2, lines 66 – 67, Column 3, lines 1 – 2, The single camera is typically positioned to encompass a wide enough field of view to capture video of multiple individuals participating in the conference, Column 4, lines 21 – 33, The video layout system 118 may also include a multi-face discovery engine 122. In certain embodiments, the multi-face discovery engine 122 is configured to detect faces in one or more video frames in the video frame buffer 120. Images of the detected faces may be used by the multi-face discovery engine 122 to generate facial images for each participant detected in the video frame. In certain embodiments, the facial images generated by the multi-face discovery engine 122 are provided to a windowed image generation engine 124. The windowed image generation engine 124 may be configured to generate a windowed image for each facial image generated by the multi-face discovery engine 122), identify regions of the first audiovisual signal that depict the first audience member and the second audience member (see Ong Column 4, lines 21 – 33, The video layout system 118 may also include a multi-face discovery engine 122. In certain embodiments, the multi-face discovery engine 122 is configured to detect faces in one or more video frames in the video frame buffer 120. Images of the detected faces may be used by the multi-face discovery engine 122 to generate facial images for each participant detected in the video frame. In certain embodiments, the facial images generated by the multi-face discovery engine 122 are provided to a windowed image generation engine 124. The windowed image generation engine 124 may be configured to generate a windowed image for each facial image generated by the multi-face discovery engine 122); Ong does not expressively teach the first audience member stationary and the second audience member in motion; generate a second audiovisual signal that includes the regions having a specified configuration, the specified configuration including a first region depicting the first audience member stationary and a second region depicting the second audience member in motion; and cause display, transmission, or a combination thereof, of the second audiovisual signal. However, O’Leary teaches the first audience member stationary and the second audience member in motion (see O’Leary Paragraph [0277], In some embodiments, adjusting the representation of the field-of-view of the one or more cameras (e.g., 606) during the live video communication session based on the detected change in the number of subjects detected in the scene is based on a determination of whether a subject in the field-of-view (e.g., 620) is stationary (e.g., relatively stationary; not moving more than a threshold amount of movement in the field-of-view of the one or more cameras), Paragraph [0385], In response to detecting Pam 1032 moving from her position next to Jack 628 to her position on the couch); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member (as taught in Ong), with determining if audience members are moving or stationary (as taught in O’Leary), the motivation being the ability to detect movement to ensure participants are within frame and subsequently, adjust the boundaries of a camera to include a participant that is moving (see O’Leary Paragraph [0277]). Ong in view of O’Leary does not expressively teach generate a second audiovisual signal that includes the regions having a specified configuration, the specified configuration including a first region depicting the first audience member stationary and a second region depicting the second audience member in motion; and cause display, transmission, or a combination thereof, of the second audiovisual signal. However, Hoang teaches generate a second audiovisual signal that includes the regions having a specified configuration, the specified configuration including a first region depicting the first audience member stationary and a second region depicting the second audience member in motion (see Hoang Abstract, User interface (UI) tiles of conference participants are arranged in a UI of conferencing software according to relative locations of those conference participants within a conference room. Positional information and video data are obtained from one or more video capture devices located within a conference room. Relative locations of conference participants within the conference room are determined based on the positional information, such as by defining coordinates for the conference participants within a coordinate system based on the positional information and the video data and determining the relative locations based on the coordinates. Output configured to cause a client application to arrange UI tiles associated with the conference participants is generated according to the relative locations. The output is then transmitted to one or more client devices to cause a UI of conferencing software at each of those client devices to display the UI tiles in the specified arrangement, Figure 6, an illustration of an example of a conference room within which conference participants are located with video capture device(s), Paragraph [0112], At 1012, the output, which may be considered as second output in some cases, is transmitted to one or more client devices, and Figure 8, showing a grid layout of video outputs, and Paragraph [0109], At 1006, new coordinates for each of the conference participants are defined based on the detected change. In particular, in response to the detected change, coordinates previously defined for conference participants who remain detected within the video data obtained from the video capture devices despite the detected change can be re-used, however new coordinates will be defined for conference participants for whom coordinates were not previously defined or for whom previously defined coordinates are no longer valid (i.e., because the conference participants left the conference room or moved to another location within the conference room)); and cause display, transmission, or a combination thereof, of the second audiovisual signal (see Hoang Abstract, User interface (UI) tiles of conference participants are arranged in a UI of conferencing software according to relative locations of those conference participants within a conference room. Positional information and video data are obtained from one or more video capture devices located within a conference room. Relative locations of conference participants within the conference room are determined based on the positional information, such as by defining coordinates for the conference participants within a coordinate system based on the positional information and the video data and determining the relative locations based on the coordinates. Output configured to cause a client application to arrange UI tiles associated with the conference participants is generated according to the relative locations. The output is then transmitted to one or more client devices to cause a UI of conferencing software at each of those client devices to display the UI tiles in the specified arrangement, Figure 6, an illustration of an example of a conference room within which conference participants are located with video capture device(s), and Paragraph [0112], At 1012, the output, which may be considered as second output in some cases, is transmitted to one or more client devices). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member (as taught in Ong), with determining if audience members are moving or stationary (as taught in O’Leary), the motivation being the ability to detect movement to ensure participants are within frame and subsequently, adjust the boundaries of a camera to include a participant that is moving (see O’Leary Paragraph [0277]). It would have been further obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member, that may additionally determine if audience members are stationary or moving (as taught in Ong in view of O’Leary), with generating a recomposed video in which each person is positioned in the display deliberately to mimic their real-life positioning and movement (as taught in Hoang), the motivation being to display better visibility of individuals by providing an individual video of each person while maintaining the physical layout of the conference participants (see Hoang Abstract and Paragraph [0019]). Regarding Claim 15, Ong in view of O’Leary and Hoang teaches The non-transitory machine-readable medium of claim 11, wherein the controller is to: determine that the first audience member is in motion (see O’Leary Paragraph [0277], In some embodiments, adjusting the representation of the field-of-view of the one or more cameras (e.g., 606) during the live video communication session based on the detected change in the number of subjects detected in the scene is based on a determination of whether a subject in the field-of-view (e.g., 620) is stationary (e.g., relatively stationary; not moving more than a threshold amount of movement in the field-of-view of the one or more cameras), Paragraph [0385], In response to detecting Pam 1032 moving from her position next to Jack 628 to her position on the couch); generate the second audiovisual signal that includes the regions having the specified configuration (see Hoang Abstract, User interface (UI) tiles of conference participants are arranged in a UI of conferencing software according to relative locations of those conference participants within a conference room. Positional information and video data are obtained from one or more video capture devices located within a conference room. Relative locations of conference participants within the conference room are determined based on the positional information, such as by defining coordinates for the conference participants within a coordinate system based on the positional information and the video data and determining the relative locations based on the coordinates. Output configured to cause a client application to arrange UI tiles associated with the conference participants is generated according to the relative locations. The output is then transmitted to one or more client devices to cause a UI of conferencing software at each of those client devices to display the UI tiles in the specified arrangement, Figure 6, an illustration of an example of a conference room within which conference participants are located with video capture device(s), and Paragraph [0112], At 1012, the output, which may be considered as second output in some cases, is transmitted to one or more client devices), the specified configuration including the first region depicting the first audience member in motion (see Hoang Paragraph [0109], At 1006, new coordinates for each of the conference participants are defined based on the detected change. In particular, in response to the detected change, coordinates previously defined for conference participants who remain detected within the video data obtained from the video capture devices despite the detected change can be re-used, however new coordinates will be defined for conference participants for whom coordinates were not previously defined or for whom previously defined coordinates are no longer valid (i.e., because the conference participants left the conference room or moved to another location within the conference room)); use post-processing techniques to increase an angle of vision of an audience member, a full profile of the audience member, or a combination thereof (see O’Leary Paragraph [0315], In FIG. 8I, device 600 detects input 832 on multi-person option 830-2. In response, device 600 switches from the single-person framing mode setting to the multi-person framing mode setting. When the multi-person framing mode setting is enabled, device 600 automatically adjusts the video feed field-of-view to include additional subjects detected in field-of-view 620 (or a subset thereof). For example, when device 600 switches to the multi-person framing mode setting, device 600 expands the video feed field-of-view to include representations 628-1 and 622-1 of both Jack and Jane, as depicted in camera preview 606 in FIG. 8J. Accordingly, portion 625 in FIG. 8J represents the expanded video feed field-of-view resulting from enabling the multi-person framing mode setting); and cause the display, the transmission, or the combination thereof, of the second audiovisual signal (see Hoang Abstract, User interface (UI) tiles of conference participants are arranged in a UI of conferencing software according to relative locations of those conference participants within a conference room. Positional information and video data are obtained from one or more video capture devices located within a conference room. Relative locations of conference participants within the conference room are determined based on the positional information, such as by defining coordinates for the conference participants within a coordinate system based on the positional information and the video data and determining the relative locations based on the coordinates. Output configured to cause a client application to arrange UI tiles associated with the conference participants is generated according to the relative locations. The output is then transmitted to one or more client devices to cause a UI of conferencing software at each of those client devices to display the UI tiles in the specified arrangement, and Paragraph [0112], At 1012, the output, which may be considered as second output in some cases, is transmitted to one or more client devices). Claims 17 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Ong et al. (U.S. Patent No. 11,165,992, hereinafter “Ong”) in view of O’Leary et al. (U.S. Pub. No. 2022/0247919, hereinafter “O’Leary”) and Lewis et al. (U.S. Patent No. 7,126,627, hereinafter “Lewis”). Regarding Claim 17, Ong in view of O’Leary teach all the limitations of claim 16, but do not expressively teach The electronic device of claim 16, wherein the first region of the first audience member depicts the first audience member having a first angle of vision relative to an optical axis of the image sensor; and wherein adjusting the angle of vision of the first audience member changes the first angle of vision. However, Lewis teaches The electronic device of claim 16, wherein the first region of the first audience member depicts the first audience member having a first angle of vision relative to an optical axis of the image sensor (see Lewis Column 3, lines 60 – 67, Column 4, line 1, Conferee 10 is ideally seated from 2 to 8 feet from video monitor 20 and particularly, its front face carrying image field 21 displaying the image of remotely located conferee 10 a. To maximize direct eye contact, FIG. 2, shows camera 25 on or proximate to image field 21 at the eye level of conferees 10 and 10A. Stated differently, video camera 25 is placed in front of monitor 20 within image field 21 whereby its optical axis coincides with the sight line 22 drawn between the eyes of the respective conferees); and wherein adjusting the angle of vision of the first audience member changes the first angle of vision (see Lewis Column 6, lines 25 – 31, the positioning of the video camera often requires that it be tilted downward slightly and aimed at the subject's eyes as best depicted in FIG. 3. However, if the video conferee moves to one side, the camera must also be panned to center the image. Camera support devices have been developed with the capabilities needed to carry out this repositioning and aiming). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member, enhancing each audience member’s region by warping its perspective, widening the viewing angle, or adjusting its resolution, then generating a second audiovisual signal containing the enhanced regions for display or transmission (as taught in Ong in view of O’Leary), with adjusting a first audience member’s viewing angle relative to the camera’s optical axis (as taught in Lewis), the motivation being to provide a method to promote eye contact (and face capturing) in a video conference (see Lewis Column 3, line 13). Regarding Claim 18, Ong in view of O’Leary and Lewis teaches The electronic device of claim 17, wherein the first angle of vision is a central axis of the first audience member (see Lewis Column 3, lines 60 – 67, Column 4, line 1, Conferee 10 is ideally seated from 2 to 8 feet from video monitor 20 and particularly, its front face carrying image field 21 displaying the image of remotely located conferee 10 a. To maximize direct eye contact, FIG. 2, shows camera 25 on or proximate to image field 21 at the eye level of conferees 10 and 10A. Stated differently, video camera 25 is placed in front of monitor 20 within image field 21 whereby its optical axis coincides with the sight line 22 drawn between the eyes of the respective conferees); and wherein the first angle of vision is adjusted so that the adjusted first angle appears to intersect an optical axis of the image sensor (see Lewis Column 6, lines 25 – 31, the positioning of the video camera often requires that it be tilted downward slightly and aimed at the subject's eyes as best depicted in FIG. 3. However, if the video conferee moves to one side, the camera must also be panned to center the image. Camera support devices have been developed with the capabilities needed to carry out this repositioning and aiming (to intersect line of sight)). Claim 19 is rejected under 35 U.S.C. 103 as being unpatentable over Ong et al. (U.S. Patent No. 11,165,992, hereinafter “Ong”) in view of O’Leary et al. (U.S. Pub. No. 2022/0247919, hereinafter “O’Leary”) and Ostap et al. (U.S. Patent No. 11,350,029, hereinafter “Ostap”). Regarding Claim 19, Ong in view of O’Leary teach all the limitations of claim 1, but do not expressively teach The electronic device of claim 1, wherein the controller is further to: receive a third audiovisual signal, via the image sensor; determine a number of regions of the third audiovisual signal that depict multiple audience members; and in response to the number of regions being greater than a threshold count, release the third audiovisual signal for display, transmission, or a combination thereof, without modifying the third audiovisual signal for replacement or enhancing facial features of the multiple audience members of the third audiovisual signal. However, Ostap teaches The electronic device of claim 1, wherein the controller is further to: receive a third audiovisual signal, via the image sensor (see Ostap Figure 2, initiate video session); determine a number of regions of the third audiovisual signal that depict multiple audience members (see Ostap Figure 2, detect conference participants); and in response to the number of regions being greater than a threshold count, release the third audiovisual signal for display, transmission, or a combination thereof, without modifying the third audiovisual signal for replacement or enhancing facial features of the multiple audience members of the third audiovisual signal (see Ostap Column 12, lines 7 – 18, FIG. 3D illustrates a fourth conference scenario 300D, e.g., a “crowd scenario” where, due to the number and locations of participants, no particular grouping of conference participants is likely to provide an improved conference view over any other grouping. At block 206C, the method 200 includes determining whether the number of participants is large enough to exceed a predetermined threshold N, e.g., 8 or more participants, 9 or more, 10 or more, 11 or more, or 12 or more participants. If the number of participants meets or exceeds the threshold N, the method 200 continues to block 206D (does not individually frame)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member, enhancing each audience member’s region by warping its perspective, widening the viewing angle, or adjusting its resolution, then generating a second audiovisual signal containing the enhanced regions for display or transmission (as taught in Ong in view of O’Leary), with detecting the number of audience member regions in an audiovisual signal and, when the number exceeds a threshold, displaying or transmitting the signal without modification (as taught in Ostap), the motivation being to reduce burden when, due to the number of participants, no particular grouping of conference participants is likely to provide an improved conference view over any other grouping (see Ostap Column 12, lines 7 – 18). Claims 6, 9 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Ong et al. (U.S. Patent No. 11,165,992, hereinafter “Ong”) in view of Hoang et al. (U.S. Pub. No. 2023/0081717, hereinafter “Hoang”) and Kurtz et al. (U.S. Pub. No. 2008/0297588, hereinafter “Kurtz”). Regarding Claim 6, Ong teaches An electronic device (see Ong Abstract, device), comprising: an image sensor (see Ong Figure 1, camera); and a controller (see Ong Column 1, lines 49 – 54, A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to generate a composited video layout of facial images in a video conference) to: receive a first frame via the image sensor (see Ong Column 1, lines 56 – 59, The computer-implemented method includes receiving a video frame from a video source, where the video frame includes faces of individuals engaged in a video conference, and Column 2, lines 66 – 67 and Column 3, lines 1 – 2, The single camera is typically positioned to encompass a wide enough field of view to capture video of multiple individuals participating in the conference); in response to a determination that the first frame depicts a first audience member and a second audience member, determine a first region of the first audience member and a second region of the second audience member (see Ong Column 4, lines 21 – 33, The video layout system 118 may also include a multi-face discovery engine 122. In certain embodiments, the multi-face discovery engine 122 is configured to detect faces in one or more video frames in the video frame buffer 120. Images of the detected faces may be used by the multi-face discovery engine 122 to generate facial images for each participant detected in the video frame. In certain embodiments, the facial images generated by the multi-face discovery engine 122 are provided to a windowed image generation engine 124. The windowed image generation engine 124 may be configured to generate a windowed image for each facial image generated by the multi-face discovery engine 122); Ong does not expressively teach generate a second frame that includes the first region and the second region having a specified configuration, the specified configuration being a grid layout; cause display, transmission, or a combination thereof, of the second frame; receive a third frame via the image sensor; and in response to a determination that the first audience member is not in the third frame, replace the first region of the first audience member of a static image of the first audience member. However, Hoang teaches generate a second frame that includes the first region and the second region having a specified configuration, the specified configuration being a grid layout (see Hoang Abstract, User interface (UI) tiles of conference participants are arranged in a UI of conferencing software according to relative locations of those conference participants within a conference room. Positional information and video data are obtained from one or more video capture devices located within a conference room. Relative locations of conference participants within the conference room are determined based on the positional information, such as by defining coordinates for the conference participants within a coordinate system based on the positional information and the video data and determining the relative locations based on the coordinates. Output configured to cause a client application to arrange UI tiles associated with the conference participants is generated according to the relative locations. The output is then transmitted to one or more client devices to cause a UI of conferencing software at each of those client devices to display the UI tiles in the specified arrangement, Figure 6, an illustration of an example of a conference room within which conference participants are located with video capture device(s), Paragraph [0112], At 1012, the output, which may be considered as second output in some cases, is transmitted to one or more client devices, and Figure 8, showing a grid layout of video outputs); cause display, transmission, or a combination thereof, of the second frame (see Hoang Abstract, User interface (UI) tiles of conference participants are arranged in a UI of conferencing software according to relative locations of those conference participants within a conference room. Positional information and video data are obtained from one or more video capture devices located within a conference room. Relative locations of conference participants within the conference room are determined based on the positional information, such as by defining coordinates for the conference participants within a coordinate system based on the positional information and the video data and determining the relative locations based on the coordinates. Output configured to cause a client application to arrange UI tiles associated with the conference participants is generated according to the relative locations. The output is then transmitted to one or more client devices to cause a UI of conferencing software at each of those client devices to display the UI tiles in the specified arrangement, Figure 6, an illustration of an example of a conference room within which conference participants are located with video capture device(s), and Paragraph [0112], At 1012, the output, which may be considered as second output in some cases, is transmitted to one or more client devices); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member (as taught in Ong), with arranging framed audience members in a grid layout (as taught in Hoang), the motivation being to display better visibility of individuals by providing an individual video of each person in an organized manner (see Hoang Abstract and Paragraph [0019]). Ong in view of Hoang does not expressively teach receive a third frame via the image sensor; and in response to a determination that the first audience member is not in the third frame, replace the first region of the first audience member of a static image of the first audience member. However, Kurtz teaches receive a third frame via the image sensor (see Kurtz Paragraph [0073], The scene analysis algorithm can evaluate imagery for current video frames from one or more cameras 120, seeking correlation with prior video frames to improve the analysis process and results); and in response to a determination that the first audience member is not in the third frame, replace the first region of the first audience member of a static image of the first audience member (see Kurtz Paragraph [0077], Further complications arise when an individual enters (or leaves) the field of view of a local environment 415 during an ongoing communication event. In the instance that a local user 10 (particularly the only local user at the moment) leaves the local environment, the device 300 can adapt to this transition in the local image content. The device 300 can transmit paused imagery or alternate imagery of other than the local environment 415, until the local user 10 returns). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member (as taught in Ong), with arranging framed audience members in a grid layout (as taught in Hoang), the motivation being to display better visibility of individuals by providing an individual video of each person in an organized manner (see Hoang Abstract and Paragraph [0019]). It would have been further obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member, in which the framed audience members are arranged in a grid layout (as taught in Ong in view of Hoang), with replacing an audience member’s video stream with a static image when it detects that the person is no longer in the camera frame (as taught in Kurtz), the motivation being to provide a solution when a user is no longer physically present in the meeting, while maintaining privacy (see Kurtz Paragraphs [0018] and [0077]). Regarding Claim 9, Ong in view of Hoang and Kurtz teaches The electronic device of claim 6, wherein the grid layout emulates placement of the first and the second audience members within an environment that includes the image sensor (see Hoang Abstract, User interface (UI) tiles of conference participants are arranged in a UI of conferencing software according to relative locations of those conference participants within a conference room. Positional information and video data are obtained from one or more video capture devices located within a conference room. Relative locations of conference participants within the conference room are determined based on the positional information, such as by defining coordinates for the conference participants within a coordinate system based on the positional information and the video data and determining the relative locations based on the coordinates. Output configured to cause a client application to arrange UI tiles associated with the conference participants is generated according to the relative locations. The output is then transmitted to one or more client devices to cause a UI of conferencing software at each of those client devices to display the UI tiles in the specified arrangement, Figure 6, an illustration of an example of a conference room within which conference participants are located with video capture device(s)). Regarding Claim 10, Ong in view of Hoang and Kurtz teaches The electronic device of claim 6, wherein the controller is to use post-processing techniques to generate an enhanced view of the first audience member, the second audience member, or a combination thereof (see Ong Column 5, lines 5 – 31, In certain embodiments, the faces within the bounding boxes 230, 232, and 234 are detected by the face detection engine 224. The face detection engine 224 may be configured to extract the facial features of faces within the bounding boxes 230, 232, and 234. The facial features extracted by the face detection engine 224 may be provided to a facial image construction engine 236. The facial image construction engine 236, in turn, may be configured to generate a facial image for each face within the bounding boxes 230, 232, and 234. In this example, the facial image construction engine 236 has generated a facial image 238 for participant 208, a facial image 240 for participant 210, and a facial image 242 for participant 212. FIG. 3 depicts an exemplary embodiment of a windowed image generation engine 302 that may be used in certain embodiments of the video layout system 118. In this example, the windowed image generation engine 302 receives facial images 238, 240, and 242 to execute further processing operations. Here, the facial images 238, 240, and 242 are provided to a localized exposure and skin tone adjustment engine 304. The localized exposure and skin tone adjustment engine 304 may execute exposure and skin tone processes to generate enhanced facial images 306, 308, and 310 for each of the facial images 238, 240, and 242. The parameters used in executing the exposure and skin tone adjustment operations may be selected so that the enhanced facial images 306, 308, and 310 have a consistent quality). Claims 7 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Ong et al. (U.S. Patent No. 11,165,992, hereinafter “Ong”) in view of Hoang et al. (U.S. Pub. No. 2023/0081717, hereinafter “Hoang”) and Kurtz et al. (U.S. Pub. No. 2008/0297588, hereinafter “Kurtz”) and O’Leary et al. (U.S. Pub. No. 2022/0247919, hereinafter “O’Leary”). Regarding Claim 7, Ong in view of Hoang and Kurtz teach all the limitations of claim 6, but do not expressively teach The electronic device of claim 6, wherein the second audience member is stationary. However, O’Leary teaches The electronic device of claim 6, wherein the second audience member is stationary (see O’Leary Paragraph [0277], In some embodiments, adjusting the representation of the field-of-view of the one or more cameras (e.g., 606) during the live video communication session based on the detected change in the number of subjects detected in the scene is based on a determination of whether a subject in the field-of-view (e.g., 620) is stationary (e.g., relatively stationary; not moving more than a threshold amount of movement in the field-of-view of the one or more cameras)). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member, in which the framed audience members are arranged in a grid layout and may replace an audience member’s video stream with a static image when it detects that the person is no longer in the camera frame (as taught in Ong in view of Hoang and Kurtz), with determining if audience members are stationary (as taught in O’Leary), the motivation being the ability to detect movement to ensure participants are within frame (see O’Leary Paragraph [0277]). Regarding Claim 8, Ong in view of Hoang, Kurtz and O’Leary teaches The electronic device of claim 6, wherein the first audience member is in motion and the second audience member is in stationary (see O’Leary Paragraph [0277], In some embodiments, adjusting the representation of the field-of-view of the one or more cameras (e.g., 606) during the live video communication session based on the detected change in the number of subjects detected in the scene is based on a determination of whether a subject in the field-of-view (e.g., 620) is stationary (e.g., relatively stationary; not moving more than a threshold amount of movement in the field-of-view of the one or more cameras), Paragraph [0385], In response to detecting Pam 1032 moving from her position next to Jack 628 to her position on the couch). Claims 12 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Ong et al. (U.S. Patent No. 11,165,992, hereinafter “Ong”) in view of O’Leary et al. (U.S. Pub. No. 2022/0247919, hereinafter “O’Leary”), Hoang et al. (U.S. Pub. No. 2023/0081717, hereinafter “Hoang”) and Chen et al. (U.S. Pub. No. 2016/0284354, hereinafter “Chen”). Regarding Claim 12, Ong in view of O’Leary and Hoang teach all the limitations of claim 11, but do not expressively teach The non-transitory machine-readable medium of claim 11, wherein the controller is to: determine that the first audience member is speaking; generate a voice print for the first audience member; and associate the voice print with the first region. However, Chen teaches The non-transitory machine-readable medium of claim 11, wherein the controller is to: determine that the first audience member is speaking (see Chen Abstract, A computer receives audio and video components from a video conference. The computer determines which participant is speaking based on comparing images of the participants with template images of speaking and non-speaking faces); generate a voice print for the first audience member (see Chen Abstract, The computer determines the voiceprint of the speaking participant by applying a Hidden Markov Model to a brief recording of the voice waveform of the participant and associates the determined voiceprint with the face of the speaking participant); and associate the voice print with the first region (see Chen Abstract, The computer recognizes and transcribes the content of statements made by the speaker, determines the key points, and displays them over the face of the participant in the video conference). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member, that may additionally determine if audience members are stationary or moving, and generates a recomposed video in which each person is positioned in the display deliberately to mimic their real-life positioning and movement (as taught in Ong in view of O’Leary and Hoang), with generating a voiceprint for an audience member that is speaking and associate that voiceprint with their respective region (as taught in Chen), the motivation being to transcribe the content of statements made by the speaker, determine the key points, and display them over the face of the participant in the video conference to avoid the dilemmas associated with language barriers, unrecognizable accents, fast speaking, or the chance that attendees arrive late to a multi-person conference and miss what was previously discussed (see Chen Paragraphs [0002] – [0003]). Regarding Claim 13, Ong in view of O’Leary, Hoang and Chen teaches The non-transitory machine-readable medium of claim 12, wherein the controller is to: determine that the second audience member is speaking (see Chen Abstract, A computer receives audio and video components from a video conference. The computer determines which participant is speaking based on comparing images of the participants with template images of speaking and non-speaking faces); generate a second voice print for the second audience member (see Chen Abstract, The computer determines the voiceprint of the speaking participant by applying a Hidden Markov Model to a brief recording of the voice waveform of the participant and associates the determined voiceprint with the face of the speaking participant); and associate the second voice print with the second region (see Chen Abstract, The computer recognizes and transcribes the content of statements made by the speaker, determines the key points, and displays them over the face of the participant in the video conference). Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Ong et al. (U.S. Patent No. 11,165,992, hereinafter “Ong”) in view of O’Leary et al. (U.S. Pub. No. 2022/0247919, hereinafter “O’Leary”), Hoang et al. (U.S. Pub. No. 2023/0081717, hereinafter “Hoang”), Chen et al. (U.S. Pub. No. 2016/0284354, hereinafter “Chen”) and Kurtz et al. (U.S. Pub. No. 2008/0297588, hereinafter “Kurtz”). Regarding Claim 14, Ong in view of O’Leary, Hoang and Chen teaches The non-transitory machine-readable medium of claim 12, wherein the controller is to: determine that the second audience member is speaking (see Chen Abstract, A computer receives audio and video components from a video conference. The computer determines which participant is speaking based on comparing images of the participants with template images of speaking and non-speaking faces); Ong in view of O’Leary, Hoang and Chen does not expressively teach populate the second region with a static image of the second audience member in response to a determination that the second audience member is outside a field of view of the image sensor. However, Kurtz teaches populate the second region with a static image of the second audience member in response to a determination that the second audience member is outside a field of view of the image sensor (see Kurtz Paragraph [0077], Further complications arise when an individual enters (or leaves) the field of view of a local environment 415 during an ongoing communication event. In the instance that a local user 10 (particularly the only local user at the moment) leaves the local environment, the device 300 can adapt to this transition in the local image content. The device 300 can transmit paused imagery or alternate imagery of other than the local environment 415, until the local user 10 returns). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of an electronic device that uses a camera to capture multiple audience members and a controller to identify separate regions of the video corresponding to each audience member, that may additionally determine if audience members are stationary or moving, and generates a recomposed video in which each person is positioned in the display deliberately to mimic their real-life positioning and movement, and generates a voiceprint for an audience member that is speaking and associate that voiceprint with their respective region (as taught in Ong in view of O’Leary, Hoang and Chen), with replacing an audience member’s video stream with a static image when it detects that the person is no longer in the camera frame (as taught in Kurtz), the motivation being to provide a solution when a user is no longer physically present in the meeting, while maintaining privacy (see Kurtz Paragraphs [0018] and [0077]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARISSA A JONES whose telephone number is (703)756-1677. The examiner can normally be reached Telework M-F 6:30 AM - 4:00 PM CT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CARISSA A JONES/Examiner, Art Unit 2691 /DUC NGUYEN/Supervisory Patent Examiner, Art Unit 2691
Read full office action

Prosecution Timeline

Sep 25, 2024
Application Filed
Apr 01, 2026
Non-Final Rejection mailed — §103, §112
Jun 18, 2026
Applicant Interview (Telephonic)
Jun 18, 2026
Examiner Interview Summary
Jun 25, 2026
Response Filed
Sep 01, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744981
METHODS, SYSTEMS, APPARATUSES, AND DEVICES FOR FACILITATING MANAGING THE BODY LANGUAGE OF A USER DURING A VIDEO COMMUNICATION SESSION
2y 8m to grant Granted Sep 22, 2026
Patent 12737846
ITERATIVE BACKGROUND GENERATION FOR VIDEO STREAMS
3y 0m to grant Granted Sep 15, 2026
Patent 12699454
TERMINAL APPARATUS, COMMUNICATION SYSTEM
2y 11m to grant Granted Aug 04, 2026
Patent 12666140
CONTRIBUTION-BASED CLOSE-UP CONTROL
2y 3m to grant Granted Jun 23, 2026
Patent 12656991
SYSTEMS, METHODS, APPARATUSES, AND DEVICES FOR MANAGING A GAZE OF A USER ENGAGED IN A VISUAL COMMUNICATION SESSION WITH AN AUDIENCE
2y 11m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+33.3%)
2y 7m (~7m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 35 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month