DETAILED ACTION
This action is in response to the application filed 12/17/2024. Claims 1 - 15 are pending and have
been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
The claimed invention of Claim 15 is directed to non-statutory subject matter. The claim does not fall within at least one of the four categories of patent eligible subject matter because, under its broadest reasonable interpretation, the recited “computer program loadable into an internal memory” encompasses a computer program per se. The limitation “loadable into” merely describes the capability of the program, and does not positively require that the program be embodied in a statutory manufacture, such as a non-transitory computer-readable storage medium or a computer memory. Accordingly, the claim encompasses non-statutory subject matter and therefore, Claim 15 is rejected.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 4, 5, 10, 11 and 15 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Tangeland et al. (U.S. Pub. No. 2023/0300295, hereinafter “Tangeland”).
Regarding Claim 1, Tangeland teaches
A computer implemented method for processing at least one video feed, the video feed comprising images of a human presenter (see Tangeland Paragraph [0042], a method 500 of identifying a presenter of shared content during an online meeting for displaying video of the presenter on top of the shared content), the method comprising the steps of:
acquiring, with one or more audio-visual input devices, of a first User Equipment at least a first source video feed (see Tangeland Paragraph [0020], At 152, video endpoint device 120 may receive video from camera 124 and at 154, video endpoint device 120 may receive audio data from microphone 126. The video and audio data may include video and audio of one or more users participating in the online meeting via video endpoint device 120. For example, the video and audio data may include video of the users in the meeting/conference room and audio of the user or users presenting or describing the shared content);
segmenting the first source video feed, thereby extracting a first presenter video feed showing a human presenter (see Tangeland Paragraph [0021], Video endpoint device 120 may detect the participants in the video of the meeting/conference room and identify which participant or participants is/are presenting the shared content. To detect the participants in the room, video endpoint device 120 may apply a machine learning-based segmentation model to separate the foreground (people) from the background (room), Paragraph [0028], To identify a participant in the room as a presenter of the shared content, video endpoint device 120 may utilize position and shape information from a foreground/background segmentation tool to create a user interface to present to the participants, Paragraph [0044], At 520, one of the multiple users is identified as a presenter for the shared content. For example, video endpoint device 120 may use a segmentation model to separate the participants from the background in a video stream of the participants in the conference room. The video endpoint device 120 may additionally generate silhouettes that define areas in the video stream that contain the participants. In one embodiment, video endpoint device 120 may identify the presenter by receiving a selection of the presenter, as described above with respect to FIGS. 2A and 2B. In another embodiment, video endpoint device 120 may automatically identify the presenter based on detecting an active speaker, as described above with respect to FIGS. 3A and 3B. In another embodiment, video endpoint device 120 may automatically identify the presenter using facial recognition to identify a participant who had been assigned a presenter role. In other embodiments, the presenter may be identified in different ways. Video endpoint device 120 may identify a silhouette that corresponds to the identified presenter);
modifying the first presenter video feed, thereby creating a modified first presenter video feed (see Tangeland Paragraph [0021], Silhouettes or masks indicating locations of the different participants in the meeting room may be added to the video of the meeting/conference room. Each silhouette/mask defines an area in the video that contains a participant in the meeting room, Paragraph [0022], when a user has been identified as a presenter, video endpoint device 120 may overlay video of the presenter defined by the silhouette/mask on the shared content and transmit the video of the presenter overlaid on the shared content to meeting server(s) 110 . In another embodiment, video endpoint device 120 may transmit information associated with the silhouette/mask surrounding the presenter to the meeting server(s) 110 as metadata with the video stream of the meeting/conference room and the shared content so that the meeting server(s) 110 or receiver devices (e.g., end devices 160-1 to 160-N) may identify the presenter from the video stream, extract video of the presenter, and place the video of the presenter on top of the shared content);
retrieving or generating one or more background images or background video feeds (see Tangeland Paragraph [0019], The meeting server 110 and the video endpoint device 120 are configured to support immersive sharing in which videos of one or more users are placed on top of shared content during online meetings. In the example illustrated in FIG. 1, video endpoint device 120 may be in a meeting or conference room that includes multiple users participating in an online meeting via the video endpoint device 120. In the example described in FIG. 1, video endpoint device 120 may receive a selection of an option to begin an immersive sharing session. For example, video endpoint device 120 may receive an input from a user via display 122 (e.g., when display 122 is a touch screen) or via input device 128 indicating that the users would like to start an immersive sharing session. In one embodiment, at 150, video endpoint device 120 may receive shared content from user device 140. The shared content may include a presentation, a document, one or more images, or other content to be shared with end devices 160-1 to 160-N during the online meeting. In another embodiment, video endpoint device 120 may directly open content to share during the online meeting instead of receiving the content from user device 140, Paragraph [0029], Video endpoint device 120 may receive the selection of participant 204 and may additionally obtain shared content 214 (e.g., video endpoint device 120 may directly open shared content 214 or may receive shared content 214 from user device 140));
compositing the modified first presenter video feed and the one or more background images or background video feeds, thereby creating an output video feed (see Tangeland Paragraph [0022], when a user has been identified as a presenter, video endpoint device 120 may overlay video of the presenter defined by the silhouette/mask on the shared content and transmit the video of the presenter overlaid on the shared content to meeting server(s) 110. In another embodiment, video endpoint device 120 may transmit information associated with the silhouette/mask surrounding the presenter to the meeting server(s) 110 as metadata with the video stream of the meeting/conference room and the shared content so that the meeting server(s) 110 or receiver devices (e.g., end devices 160-1 to 160-N) may identify the presenter from the video stream, extract video of the presenter, and place the video of the presenter on top of the shared content, Paragraph [0025], When the presenter has been identified from the group of participants, video endpoint device 120 may transmit information to meeting server(s) 110 via a content channel for the immersive sharing session. In one embodiment, video endpoint device 120 may place the video of the presenter(s) on top of (overlaying) the shared content and transmit the shared content with the video of the presenter(s) overlaying the shared content to the meeting server(s) 110. In another embodiment, video endpoint device 120 may transmit the shared content, the video stream of the participants, and metadata identifying the silhouette(s)/mask(s) of the presenter(s) to meeting server(s) 110. Meeting server(s) 110 or end devices 160-1 to 160-N may extract the video of the presenter(s) identified by the metadata and place the video of the presenter(s) on top of (overlaying) the shared content and for displaying to participants in the online meeting);
outputting the output video feed as a synthetic camera feed to a videoconferencing system (see Tangeland Paragraph [0025], Meeting server(s) 110 or end devices 160-1 to 160-N may extract the video of the presenter(s) identified by the metadata and place the video of the presenter(s) on top of (overlaying) the shared content and for displaying to participants in the online meeting, Paragraph [0030], Meeting server(s) 110 may place the video of participant 204 on top of (overlaying) the shared content 214 and may transmit the video of participant 204 on top of the shared content 214 to end devices 160-1 to 160-N for display, meeting server(s) 110 may transmit the shared content, the video stream of participants 202-210, and the metadata identifying the silhouette of participant 204 to end devices 160-1 to 160-N and end devices 160-1 to 160-N may place the video of participant 204 on top of shared content 214 for display on end devices 160-1 to 160-N, Paragraph [0034], FIG. 3B illustrates an example in which video of participant 308 has been placed on top of shared content 312. In one embodiment, video endpoint device 120 may place the video of participant 308 on top of (overlaying) the shared content and transmit the shared content with the video of participant 308 overlaying the shared content to the meeting server(s) 110. In another embodiment, video endpoint device 120 may transmit the shared content, the video of participants 302-310, and the metadata identifying the silhouette of participant 308 to meeting server(s) 110. Meeting server(s) 110 may place the video of participant 308 on top of the shared content 312 and may transmit the video of participant 308 on top of the shared content 312 to end devices 160-1 to 160-N for display. Alternatively, meeting server(s) 110 may transmit the shared content, the video of participants 302-310, and the metadata identifying the silhouette of participant 308 to end devices 160-1 to 160-N and end devices 160-1 to 160-N may place the video of participant 308 on top of shared content 312 for display on end devices 160-1 to 160-N);
wherein modifying the first presenter video feed comprises the steps of:
detecting, at least in the first source video feed, a trigger situation (see Tangeland Paragraph [0024], In other embodiments described below with respect to FIGS. 2A and 2B, video endpoint device 120 may display a picture or video of the participants in the meeting/conference room. The picture or video includes the silhouettes/masks of the participants and one of the participants may manually select the participant(s) who is/are the presenter, Paragraph [0028], A user may manually select which participant shown in the user interface is the presenter. In one embodiment, video endpoint device 120 may display the self-view of the conference room on display 122 and, when display 122 is a touch screen, a participant may select one of participants 202-210 as the presenter by touching an image of the presenter on the touch screen. In another embodiment, video endpoint device 120 may display a selection tool 212 (e.g., a cursor, an arrow, a finger, etc.) to allow a user to move about within the displayed view to select which one of the participants 202-210 is the presenter using, for example, a mouse or other input device, such as input device 128 (not illustrated in FIG. 2A). In the example illustrated in FIG. 2A, participant 204 has been selected as the presenter);
depending on the trigger situation detected, modifying the first presenter video feed (see Tangeland Paragraph [0024], In other embodiments described below with respect to FIGS. 2A and 2B, video endpoint device 120 may display a picture or video of the participants in the meeting/conference room. The picture or video includes the silhouettes/masks of the participants and one of the participants may manually select the participant(s) who is/are the presenter, Paragraph [0028], A user may manually select which participant shown in the user interface is the presenter. In one embodiment, video endpoint device 120 may display the self-view of the conference room on display 122 and, when display 122 is a touch screen, a participant may select one of participants 202-210 as the presenter by touching an image of the presenter on the touch screen. In another embodiment, video endpoint device 120 may display a selection tool 212 (e.g., a cursor, an arrow, a finger, etc.) to allow a user to move about within the displayed view to select which one of the participants 202-210 is the presenter using, for example, a mouse or other input device, such as input device 128 (not illustrated in FIG. 2A). In the example illustrated in FIG. 2A, participant 204 has been selected as the presenter, Paragraph [0025], When the presenter has been identified from the group of participants, video endpoint device 120 may transmit information to meeting server(s) 110 via a content channel for the immersive sharing session. In one embodiment, video endpoint device 120 may place the video of the presenter(s) on top of (overlaying) the shared content and transmit the shared content with the video of the presenter(s) overlaying the shared content to the meeting server(s) 110. In another embodiment, video endpoint device 120 may transmit the shared content, the video stream of the participants, and metadata identifying the silhouette(s)/mask(s) of the presenter(s) to meeting server(s) 110. Meeting server(s) 110 or end devices 160-1 to 160-N may extract the video of the presenter(s) identified by the metadata and place the video of the presenter(s) on top of (overlaying) the shared content and for displaying to participants in the online meeting).
Regarding Claim 4, Tangeland teaches
The method of claim 1, wherein the step of retrieving or generating the background image or background video feed (see Tangeland Paragraph [0019], The meeting server 110 and the video endpoint device 120 are configured to support immersive sharing in which videos of one or more users are placed on top of shared content during online meetings. In the example illustrated in FIG. 1, video endpoint device 120 may be in a meeting or conference room that includes multiple users participating in an online meeting via the video endpoint device 120. In the example described in FIG. 1, video endpoint device 120 may receive a selection of an option to begin an immersive sharing session. For example, video endpoint device 120 may receive an input from a user via display 122 (e.g., when display 122 is a touch screen) or via input device 128 indicating that the users would like to start an immersive sharing session. In one embodiment, at 150, video endpoint device 120 may receive shared content from user device 140. The shared content may include a presentation, a document, one or more images, or other content to be shared with end devices 160-1 to 160-N during the online meeting. In another embodiment, video endpoint device 120 may directly open content to share during the online meeting instead of receiving the content from user device 140, Paragraph [0029], Video endpoint device 120 may receive the selection of participant 204 and may additionally obtain shared content 214 (e.g., video endpoint device 120 may directly open shared content 214 or may receive shared content 214 from user device 140)) comprises retrieving them from a storybook dataset comprising an ordered sequence of background images and/or background video feeds (see Tangeland Paragraph [0019], the shared content may include a presentation, a document, one or more images, or other content to be shared with end devices 160-1 to 160-N during the online meeting. The background is changed based on a pre-existing slide presentation, in which a dataset comprises an ordered sequence of background images).
Regarding Claim 5, Tangeland teaches
The method of claim 1, wherein a step of switching from a current background image or background video feed to a subsequent background image or background video feed in the sequence is triggered by detection of a trigger situation (see Tangeland Paragraph [0019], the shared content may include a presentation, a document, one or more images, or other content to be shared with end devices 160-1 to 160-N during the online meeting. The background is changed based on a pre-existing slide presentation, in which a dataset comprises an ordered sequence of background images).
Regarding Claim 10, Tangeland teaches
The method of claim 1, wherein the steps recited are performed by the first User Equipment, and wherein the first User Equipment is a personal computing device (see Tangeland Paragraph [0015], The video endpoint device 120 may be a videoconference endpoint designed for personal use (e.g., a desk device used by a single user)).
Regarding Claim 11, Tangeland teaches
The method of claim 1, further comprising the steps of:
acquiring, with the first User Equipment a second source video feed (see Tangeland Paragraph [0040], In some embodiments, multiple users in different locations may be designated as presenters. For example, a host of the online meeting (or another participant) may designate a first participant who is participating in the online meeting via video endpoint device 120 as a presenter and may additionally designate a second participant who is participating in the online meeting via end device 160-1 as a presenter);
segmenting the second source video feed, thereby extracting a second presenter video feed showing the human presenter (see Tangeland Paragraph [0040], In these embodiments, video endpoint device 120 transmits video and metadata including information identifying the silhouette of the first participant to meeting server(s) 110 and end device 160-1 (or a meeting application associated with end device 160-1) transmits video and metadata including information identifying the silhouette of the second participant to meeting server(s) 110);
modifying the second presenter video feed, thereby creating a modified second presenter video feed (see Tangeland Paragraph [0040], In these embodiments, video endpoint device 120 transmits video and metadata including information identifying the silhouette of the first participant to meeting server(s) 110 and end device 160-1 (or a meeting application associated with end device 160-1) transmits video and metadata including information identifying the silhouette of the second participant to meeting server(s) 110)
selectively compositing the modified second presenter video feed instead of the modified first presenter video feed when creating the output video feed depending on a detected trigger situation (see Tangeland Paragraph [0041], placement of the videos of participants 404 and 408 may dynamically change (e.g., based on a user selection, which participant is talking, where the participants are looking, if the physical locations of the participants change, etc.), Paragraph [0004], FIGS. 2A and 2B show examples of manually selecting a participant as a presenter of shared content, according to an example embodiment, Paragraph [0005], FIGS. 3A and 3B show examples of automatically selecting a participant as a presenter based on voice detection, according to an example embodiment, therefore presenters can be manually or automatically chosen, therefore if a different presenter is chosen automatically or manually, they would be the presenter).
Regarding Claim 15, Tangeland teaches
A computer program loadable into an internal memory of a personal computing device comprising a camera, microphone and an input device, the computer program comprising computer program code to make, when said computer program is loaded in the personal computing device, the personal computing device execute the method according to claim 1 (see Tangeland Paragraph [0056], In various embodiments, entities as described herein may store data/information in any suitable volatile and/or non-volatile memory item (e.g., magnetic hard disk drive, solid state hard drive, semiconductor storage device, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), application specific integrated circuit (ASIC), etc.), software, logic (fixed logic, hardware logic, programmable logic, analog logic, digital logic), hardware, and/or in any other suitable component, device, element, and/or object as may be appropriate, Paragraph [0016], Video endpoint device 120 may include display 122, camera 124, and microphone 126. In one embodiment, display 122, camera 124, and/or microphone 126 may be integrated with video endpoint device 120. In another embodiment, display 122, camera 124, and/or microphone 126 may be separate devices connected to video endpoint device 120 via a wired or wireless connection. Display 122 may include a touch screen display configured to receive an input from a user. Video endpoint device 120 may further include an input device 128, such as a keyboard or a mouse, that may be integrated in or connected to video endpoint device 120, Paragraph [0057], Note that in certain example implementations, operations as set forth herein may be implemented by logic encoded in one or more tangible media that is capable of storing instructions and/or digital information and may be inclusive of non-transitory tangible media and/or non-transitory computer readable storage media (e.g., embedded logic provided in: an ASIC, digital signal processing (DSP) instructions, software [potentially inclusive of object code and source code], etc.) for execution by one or more processor(s), and/or other similar machine, etc. Generally, memory element(s) 604 and/or storage 606 can store data, software, code, instructions (e.g., processor instructions), logic, parameters, combinations thereof, and/or the like used for operations described herein. This includes memory element(s) 604 and/or storage 606 being able to store data, software, code, instructions (e.g., processor instructions), logic, parameters, combinations thereof, or the like that are executed to carry out operations in accordance with teachings of the present disclosure).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Tangeland et al. (U.S. Pub. No. 2023/0300295, hereinafter “Tangeland”) in view of Aiba (U.S. Pub. No. 2017/0041556).
Regarding Claim 2, Tangeland teaches all the limitations of claim 1, but does not expressively teach
The method of claim 1, wherein the step of detecting trigger situations comprises detecting a video trigger situation in a video feed, typically in a source video feed or in a presenter video feed,
wherein the video trigger situation is at least one of:
recognition of a hand gesture performed by the presenter;
a change of the degree in which the presenter's movements are animated;
a facial expression of the presenter;
the head of the presenter turning;
the presenter moving as a whole.
However, Aiba teaches
The method of claim 1, wherein the step of detecting trigger situations comprises detecting a video trigger situation in a video feed, typically in a source video feed or in a presenter video feed,
wherein the video trigger situation is at least one of:
recognition of a hand gesture performed by the presenter;
a change of the degree in which the presenter's movements are animated;
a facial expression of the presenter (see Aiba Paragraph [0032], The speaker identification unit 102 identifies a current speaker from among multiple persons included in the video and the video-related information acquired by the video acquisition unit 101. The speaker identification unit 102 may identify the current speaker, for example, by reading a change of facial expression);
the head of the presenter turning;
the presenter moving as a whole.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the teaching of a method that detects a presenter in a video, modifies the presenter based on detected trigger conditions, composites them with a selected or generated background, and outputs the result as a synthetic camera feed for a videoconferencing system (as taught in Tangeland), with identifying a speaker in a video conference based on facial expressions (as taught in Aiba), the motivation being to improve the accuracy of determination as to which participant should be the primary focus, such as the speaker (see Aiba Paragraph [0078]).
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Tangeland et al. (U.S. Pub. No. 2023/0300295, hereinafter “Tangeland”) in view of Daredia et al. (U.S. Pub. No. 2020/0403816, hereinafter “Daredia”).
Regarding Claim 3, Tangeland teaches all the limitations of claim 1, but does not expressively teach
The method of claim 1, wherein the step of detecting trigger situations comprises detecting a sound trigger situation in at least one sound track, typically in a sound track of the source video feed or presenter video feed,
wherein the sound trigger situation is at least one of:
recognition of a keyword or key phrase in the presenter's speech;
an increase or decrease in average loudness of the presenter's speech;
a pause in the presenter's speech.
However, Daredia teaches
The method of claim 1, wherein the step of detecting trigger situations comprises detecting a sound trigger situation in at least one sound track, typically in a sound track of the source video feed or presenter video feed,
wherein the sound trigger situation is at least one of:
recognition of a keyword or key phrase in the presenter's speech;
an increase or decrease in average loudness of the presenter's speech (see Daredia Abstract, the disclosed systems can attribute segments of audio associated with a meeting to meeting attendees (i.e., identify a meeting attendee as the speaker of one or more audio segments) based on speaking volumes captured by the audio. For example, the audio of a meeting can include speech from a plurality of meeting attendees, where the speech associated with each meeting attendee corresponds to a particular speaking volume. The disclosed systems can use the speaking volumes to map speakers to meeting attendees, therefore associating the meeting attendees with particular segments of speech, Paragraph [0022], One or more embodiments described herein include a speaker attribution system that utilizes flexible volume-based speaker attribution to accurately identify speakers in a meeting and associate digital meeting content (i.e., digital meeting items) with those speakers. In particular, the speaker attribution system can analyze speaking volumes captured within audio associated with a meeting to attribute segments of the audio to meeting attendees (i.e., identify one of the meeting attendees as the speaker of one or more audio segments). For example, the audio of a meeting can include speech from a plurality of meeting attendees, where the speech of each meeting attendee corresponds to a particular speaking volume);
a pause in the presenter's speech.
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the teaching of a method that detects a presenter in a video, modifies the presenter based on detected trigger conditions, composites them with a selected or generated background, and outputs the result as a synthetic camera feed for a videoconferencing system (as taught in Tangeland), with identifying a speaker in a video conference based on the speaking volume of participants (as taught in Daredia), the motivation being to flexibly identify meeting attendees as speakers and accurately associate digital meeting content with those meeting attendees (see Daredia Paragraph [0007]).
Claims 6 – 9, 13 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Tangeland et al. (U.S. Pub. No. 2023/0300295, hereinafter “Tangeland”) in view of Krol et al. (U.S. Pub. No. 2022/0124284, hereinafter “Krol”).
Regarding Claim 6, Tangeland teaches all the limitations of claim 4, but does not expressively teach
The method of claim 4, wherein the storybook dataset comprises at least one definition of a virtual 3D scene, and wherein at least one background image or background video feed is generated by rendering the virtual 3D scene in a virtual camera.
However, Krol teaches
The method of claim 4, wherein the storybook dataset comprises at least one definition of a virtual 3D scene, and wherein at least one background image or background video feed is generated by rendering the virtual 3D scene in a virtual camera (see Krol Paragraph [0039] and Figure 1, The virtual environment rendered in interface 100 includes background image 120 and a three-dimensional model 118 of an arena. The arena may be a venue or building in which the videoconference should take place. The arena may include a floor area bounded by walls, Paragraph [0040], In addition to the arena, the virtual environment can include various other three-dimensional models that illustrate different components of the environment. For example, the three-dimensional environment can include a decorative model 114, a speaker model 116, and a presentation screen model 122).
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the teaching of a method that detects a presenter in a video, modifies the presenter based on detected trigger conditions, composites them with a selected or generated background, and outputs the result as a synthetic camera feed for a videoconferencing system (as taught in Tangeland), with a selection of 3D scenes to select for a video conference background by viewing it through a virtual camera (as taught in Krol), the motivation being to address the issue in video conferencing technology that causes loss of a sense of place in a virtual environment, with a goal of improving realism and enhancing social connections (see Krol Paragraph [0007]).
Regarding Claim 7, Tangeland in view of Krol teaches
The method of claim 6, wherein the step of compositing comprises placing further background images or background video feeds in the virtual 3D environment and rendering them as part of the virtual 3D scene in the virtual camera (see Krol Abstract, The system has a presented mode that allows for a presentation stream to be texture mapped to a presenter screen situated within the virtual environment, Paragraph [0041], Presentation screen model 122 can serve to provide an outlet to present a presentation. Video of the presenter or a presentation screen share may be texture mapped onto presentation screen model 122, Paragraph [0118], FIG. 11 illustrates an interface 1100 with a presentation screen share in a three-dimensional virtual environment used for videoconferencing. As described above with respect to FIG. 1, interface 1100 may be displayed to a user who can navigate around the virtual environment. As illustrated in interface 1100, the virtual environment includes an avatar 1104 and a presentation screen 1106, Paragraph [0119], In this embodiment, a presentation stream from a device of a participant in the conference is received. The presentation stream is texture mapped onto a three-dimensional model of a presentation screen 1106. In one embodiment, the presentation stream may be a video stream from a camera on user's device. In another embodiment, the presentation stream may be a screen share from the user's device, where a monitor or window is shared. Through screen share or otherwise, the presentation video and audio stream could also be from an external source, for example a livestream of an event. When the user enables presenter mode, the presentation stream (and audio stream) of the user is published to the server tagged with the name of the screen the user wants to use. Other clients are notified that a new stream is available).
Regarding Claim 8, Tangeland in view of Krol teaches
The method of claim 6, wherein the step of compositing comprises placing the modified first presenter video feed in the virtual 3D environment and rendering it as part of the virtual 3D scene in the virtual camera (see Krol Paragraph [0036], Interface 100 includes avatars 102A and B, which each represent different participants to the videoconference. Avatars 102A and B, respectively, have texture mapped video streams 104A and B from devices of the first and second participant. A texture map is an image applied (mapped) to the surface of a shape or polygon. Here, the images are respective frames of the video. The camera devices capturing video streams 104A and B are positioned to capture faces of the respective participants. In this way, the avatars have texture mapped thereon, moving images of faces as participants in the meeting talk and listen, Paragraph [0041], Presentation screen model 122 can serve to provide an outlet to present a presentation. Video of the presenter or a presentation screen share may be texture mapped onto presentation screen model 122).
Regarding Claim 9, Tangeland in view of Krol teaches
The method of claim 1, wherein the step of modifying the first presenter video feed comprises at least one of:
gradually zooming in onto the presenter's face, or gradually zooming away;
rapid switching to a view with different zoom level;
upscaling or downscaling;
rendering a virtual view of the presenter from a perspective other than that provided by physically existing video cameras (see Krol Paragraph [0146], Renderer 1518 renders, from a perspective of a virtual camera of the user of device 306A, for output to display 1510 the three-dimensional virtual space including the texture-mapped three-dimensional models of the avatars for respective participants located at the received, corresponding position and oriented at the direction, Paragraph [0078], As device 306A receives video stream 424B, device 306A texture maps frames from video stream 424A on to an avatar corresponding to device 306B. That texture mapped avatar is re-rendered within the three-dimensional virtual space and presented to a user of device 306A);
modifying a soundtrack that is incorporated in the output video feed.
Regarding Claim 13, Tangeland in view of Krol teaches
The method of claim 1, further comprising the steps of
acquiring, with one or more audio-visual input devices of a second User Equipment at least a further source video feed (see Tangeland Paragraph [0040], In some embodiments, multiple users in different locations may be designated as presenters. For example, a host of the online meeting (or another participant) may designate a first participant who is participating in the online meeting via video endpoint device 120 as a presenter and may additionally designate a second participant who is participating in the online meeting via end device 160-1 as a presenter);
segmenting the further source video feed, thereby extracting a further presenter video feed showing a human presenter (see Tangeland Paragraph [0040], In these embodiments, video endpoint device 120 transmits video and metadata including information identifying the silhouette of the first participant to meeting server(s) 110 and end device 160-1 (or a meeting application associated with end device 160-1) transmits video and metadata including information identifying the silhouette of the second participant to meeting server(s) 110);
modifying the further presenter video feed, thereby creating a modified further presenter video feed (see Tangeland Paragraph [0040], In these embodiments, video endpoint device 120 transmits video and metadata including information identifying the silhouette of the first participant to meeting server(s) 110 and end device 160-1 (or a meeting application associated with end device 160-1) transmits video and metadata including information identifying the silhouette of the second participant to meeting server(s) 110)
creating the output video feed by compositing the modified further presenter video feed in addition to the modified first presenter video feed and the background image or background video feed (see Tangeland Paragraph [0040], video endpoint device 120 or end device 160-1 transmits shared content to meeting server(s) 110 (e.g., based on where the shared content is stored). When the shared content is shared during the online meeting, meeting server(s) 110 or receiver endpoints (e.g., end device 160-N) use the metadata identifying the silhouettes of the first and second participants to extract the videos of the first and second participants/presenters and place the videos on top of the shared content so the videos of the first and second participants/presenters are displayed on top of the shared content at the same time);
wherein the step of compositing comprises placing the modified first presenter video feed and the modified further presenter video feed in the same virtual 3D environment (see Krol Paragraph [0036], Interface 100 includes avatars 102A and B, which each represent different participants to the videoconference. Avatars 102A and B, respectively, have texture mapped video streams 104A and B from devices of the first and second participant. A texture map is an image applied (mapped) to the surface of a shape or polygon. Here, the images are respective frames of the video. The camera devices capturing video streams 104A and B are positioned to capture faces of the respective participants. In this way, the avatars have texture mapped thereon, moving images of faces as participants in the meeting talk and listen, Paragraph [0041], Presentation screen model 122 can serve to provide an outlet to present a presentation. Video of the presenter or a presentation screen share may be texture mapped onto presentation screen model 122, and Figure 1, displaying two video streams of participants).
Regarding Claim 14, Tangeland in view of Krol teaches
The method of claim 13, comprising performing the step compositing in a processing unit, the processing unit being implemented in the first User Equipment, or the processing unit being implemented by a remote, cloud-based service (see Tangeland Paragraph [0048], the computing device 600 may include one or more processor(s) 602, one or more memory element(s) 604, storage 606, a bus 608, one or more network processor unit(s) 610 interconnected with one or more network input/output (I/O) interface(s) 612, one or more I/O interface(s) 614, and control logic 620. In various embodiments, instructions associated with logic for computing device 600 can overlap in any manner and are not limited to the specific allocation of instructions and/or operations described herein, Paragraph [0052], network processor unit(s) 610 may enable communication between computing device 600 and other systems, entities, etc., via network I/O interface(s) 612 (wired and/or wireless) to facilitate operations discussed for various embodiments described herein. Examples of wireless communication capabilities include short-range wireless communication (e.g., Bluetooth), wide area wireless communication (e.g., 4G, 5G, etc.). In various embodiments, network processor unit(s) 610 can be configured as a combination of hardware and/or software, such as one or more Ethernet driver(s) and/or controller(s) or interface cards, Fibre Channel (e.g., optical) driver(s) and/or controller(s), wireless receivers/transmitters/transceivers, baseband processor(s)/modem(s), and/or other similar network interface driver(s) and/or controller(s) now known or hereafter developed to enable communications between computing device 600 and other systems, entities, etc. to facilitate operations for various embodiments described herein. In various embodiments, network I/O interface(s) 612 can be configured as one or more Ethernet port(s), Fibre Channel ports, any other I/O port(s), and/or antenna(s)/antenna array(s) now known or hereafter developed. Thus, the network processor unit(s) 610 and/or network I/O interface(s) 612 may include suitable interfaces for receiving, transmitting, and/or otherwise communicating data and/or information in a network environment).
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Tangeland et al. (U.S. Pub. No. 2023/0300295, hereinafter “Tangeland”) in view of Kerr et al. (U.S. Pub. No. 2021/0074014, hereinafter “Kerr”).
Regarding Claim 12, Tangeland teaches all the limitations of claim 1, but does not expressively teach
The method of claim 11, wherein the step of compositing comprises, when switching between the modified first and second presenter video feed, adapting the background image or background video feed according to a relative pose of a first camera and a second camera that generate the first source video feed and second source video feed, respectively;
and/or wherein the method comprises a camera registration step for determining the relative pose based upon the first source video feed and the second source video feed.
However, Kerr teaches
The method of claim 11, wherein the step of compositing comprises, when switching between the modified first and second presenter video feed, adapting the background image or background video feed according to a relative pose of a first camera and a second camera that generate the first source video feed and second source video feed, respectively;
and/or wherein the method comprises a camera registration step for determining the relative pose based upon the first source video feed and the second source video feed (see Kerr Abstract, A device for positional synchronization of virtual and physical cameras may include a processor configured to determine a first position of a physical camera relative to another electronic device in a physical environment. The processor may be configured to initiate positioning of a virtual camera in a second position within a computer-generated environment, wherein the second position relative to a representation of the person in the computer-generated environment coincides with the first position. The processor may be configured to receive an image frame captured by the physical camera and a virtual image frame generated by the virtual camera. The processor may be configured to generate a computer-generated reality image frame that includes at least a portion of the image frame composited with at least a portion of the virtual image frame ).
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the teaching of a method that detects a presenter in a video, modifies the presenter based on detected trigger conditions, composites them with a selected or generated background, and outputs the result as a synthetic camera feed for a videoconferencing system (as taught in Tangeland), with adapting background imagery according to relative poses of a plurality of cameras that generate a video feed (as taught in Kerr), the motivation being to bridge a gap between computer-generated environments and a physical environment by providing an enhanced physical environment that is augmented with electronic information (see Kerr Paragraph [0003]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARISSA A JONES whose telephone number is (703)756-1677. The examiner can normally be reached Telework M-F 6:30 AM - 4:00 PM CT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CARISSA A JONES/Examiner, Art Unit 2691
/DUC NGUYEN/Supervisory Patent Examiner, Art Unit 2691