DETAILED ACTION
This action is in response to the remarks filed 06/30/2026. Claims 1 – 22 are pending and have
been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claims 1 - 22 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Response to Amendment
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 8 – 10, 13, 14, 16, 21 and 22 are rejected under 35 U.S.C. 103 as being unpatentable over Sommerlade et al. (U.S. Pub. No. 2022/0400228, hereinafter “Sommerlade”) in view of Agrawal et al. (U.S. Pub. No. 2023/0276115, hereinafter “Agrawal”).
Regarding Claim 1, Sommerlade teaches
A network video communication device (see Sommerlade Paragraph [0037], one or more computing systems 102, 104, and/or 106), comprising:
a transmission circuit (see Sommerlade Paragraph [0045], The computing system 200 may include one or more communication connections 216 allowing communications with other computing devices/systems 250. Examples of suitable communication connections 216 include, but are not limited to, radio frequency (RF) transmitter, receiver, and/or transceiver circuitry; universal serial bus (USB), parallel, and/or serial ports, and Paragraph [0038], the computing system 200 includes one or more of the computing systems 102, 104, and 106 described with respect to the network environment 100 of FIG. 1);
a display (see Sommerlade Paragraph [0033], each computing system 102, 104, and 106 includes a display 116, 118, and 120, respectively, and a camera 122, 124, and 126, respectively. The camera may be a built-in component of the computing system, such as the camera 122 corresponding to the computing system 102, which is a tablet computer, and the camera 126 corresponding to the computing system 106, which is a laptop computer. Alternatively, the camera may be an external component of the computing system, such as the camera 124 corresponding to the computing system 104, which is a desktop computer. Moreover, it is to be understood that the computing systems 102, 104, and/or 106 can take various other forms, such as, for example, that of a mobile phone (e.g., smartphone), wearable computing system, television (e.g., smart TV), set-top box, and/or gaming console. Furthermore, the specific embodiment of the display device and/or camera may be tailored to each particular type of computing system);
an image capture circuit, for capturing a first real-time image (see Sommerlade Paragraph [0033], each computing system 102, 104, and 106 includes a display 116, 118, and 120, respectively, and a camera 122, 124, and 126, respectively. The camera may be a built-in component of the computing system, such as the camera 122 corresponding to the computing system 102, which is a tablet computer, and the camera 126 corresponding to the computing system 106, which is a laptop computer. Alternatively, the camera may be an external component of the computing system, such as the camera 124 corresponding to the computing system 104, which is a desktop computer. Moreover, it is to be understood that the computing systems 102, 104, and/or 106 can take various other forms, such as, for example, that of a mobile phone (e.g., smartphone), wearable computing system, television (e.g., smart TV), set-top box, and/or gaming console. Furthermore, the specific embodiment of the display device and/or camera may be tailored to each particular type of computing system); and
a processing circuit (see Sommerlade Paragraph [0038], processing unit 202), coupled to the transmission circuit, the image capture circuit and the display (see Sommerlade Figure 2, computing device consists of processing unit 202, communication connections 216, camera 218, Paragraph [0045], The computing system 200 may also have one or more input device(s) 212 such as a keyboard, a mouse, a pen, a sound or voice input device, a touch, or swipe input device, and/or the camera 218, etc. The one or more input device 212 may include an image sensor, such as one or more image sensors included in a camera 124, 126, 122, or 218 for example. The output device(s) 214 may include those components such as, but not limited to, a display, speakers, a printer, etc. The aforementioned devices are examples and others may be used. The computing system 200 may include one or more communication connections 216 allowing communications with other computing devices/systems 250. Examples of suitable communication connections 216 include, but are not limited to, radio frequency (RF) transmitter, receiver, and/or transceiver circuitry; universal serial bus (USB), parallel, and/or serial ports); wherein the network video communication device performs following steps:
controlling, by the processing circuit, the image capture circuit to capture the first real-time image of a first user local to the network video communication device (see Sommerlade Paragraph [0067], The method starts and flow proceeds to 602. At 602, the method may capture, via a camera of a computing system, a video stream comprising images of a user of the computing system. For example, a camera, such as camera 126 (FIG. 1) may acquire a video stream of user 112, Paragraph [0066], aspects of the method 600 are performed by one or more processing devices, such as a computing system or server. Further, the method 600 can be performed by gates or circuits associated with a processor, Application Specific Integrated Circuit (ASIC), a field programmable gate array (FPGA), a system on chip (SOC), a neural processing unit, or other hardware device. Hereinafter, the method 600 shall be explained with reference to the systems, components, modules, software, data structures, user interfaces, etc. described in conjunction with FIGS. 1-5, Paragraph [0043], a number of program modules and data files may be stored in the system memory 204. While executing on the processing unit 202, the program modules 206 (e.g., software applications 220) may perform processes including, but not limited to, the aspects, as described herein);
receiving, by the transmission circuit, a second video signal, wherein the second video signal comprises a second real-time image captured by a second video communication device which is another one of the participants of the online video conference (see Sommerlade Paragraph [0049], As depicted in FIG. 3A, a first participant 302 may be viewing a conferencing session at a display device 304. The display device 304 may include an integrated camera 306 for acquiring a video stream, an image, or a stream of images of the participant 302. A gaze tracker, for example the gaze tracker 221, may utilize one or more of the images to determine that the gaze 308 of the participant 302 is directed to an image or video representation of the participant 310. A second participant 310 may be viewing the same conferencing session at a display device 312. The display device 312 may include an integrated camera 314 for acquiring a video stream, an image, or a stream of images of the participant 310, Paragraph [0045], The computing system 200 may include one or more communication connections 216 allowing communications with other computing devices/systems 250. Examples of suitable communication connections 216 include, but are not limited to, radio frequency (RF) transmitter, receiver, and/or transceiver circuitry; universal serial bus (USB), parallel, and/or serial ports);
controlling, by the processing circuit, the display to display a video conference screen, wherein the video conference screen comprises the second real-time image (see Sommerlade Paragraph [0049], As depicted in FIG. 3A, a first participant 302 may be viewing a conferencing session at a display device 304. The display device 304 may include an integrated camera 306 for acquiring a video stream, an image, or a stream of images of the participant 302. A gaze tracker, for example the gaze tracker 221, may utilize one or more of the images to determine that the gaze 308 of the participant 302 is directed to an image or video representation of the participant 310. A second participant 310 may be viewing the same conferencing session at a display device 312. The display device 312 may include an integrated camera 314 for acquiring a video stream, an image, or a stream of images of the participant 310, Figures 3A – 3D, in which displays show a video conference interface including video streams of multiple participants, Paragraph [0066], aspects of the method 600 are performed by one or more processing devices, such as a computing system or server. Further, the method 600 can be performed by gates or circuits associated with a processor, Application Specific Integrated Circuit (ASIC), a field programmable gate array (FPGA), a system on chip (SOC), a neural processing unit, or other hardware device. Hereinafter, the method 600 shall be explained with reference to the systems, components, modules, software, data structures, user interfaces, etc. described in conjunction with FIGS. 1-5, Paragraph [0043], a number of program modules and data files may be stored in the system memory 204. While executing on the processing unit 202, the program modules 206 (e.g., software applications 220) may perform processes including, but not limited to, the aspects, as described herein);
determining, by the processing circuit, an identity or a status of the first user corresponding to the online video conference (see Sommerlade Paragraph [0039], The system memory 204 may include an operating system 205 and one or more program modules 206 suitable for running software application 220, such as one or more components supported by the systems described herein. The operating system 205, for example, may be suitable for controlling the operation of the computing system 102, 104, or 106 for example. As examples, system memory 204 may include a gaze tracker 221, a gaze coordinator 222, a compositor 223, and a gaze adjuster 224. The gaze tracker 221 may identify gaze and head pose of each participant that shares or does not share a video stream of their respective device. In some examples, the gaze tracker 221 may map gaze direction and head pose to an on-screen location for each user, or an “off screen” location if the user looks away. The gaze tracker 221 may provide an indication including a source (e.g., an identity of the user) and a target (e.g., a location of the directed gaze and/or an identity of the participant to which the gaze is directed) to the gaze coordinator 222, Paragraph [0040], The gaze coordinator 222 receives the information from the gaze tracker 221; as described above, the gaze coordinator 222 may receive source-target pairs identifying an identity of the user and an identity of the participant to which the user's gaze is directed, Paragraph [0062], a compositor 432 residing at a same computing system as the gaze tracker 404 for example, may provide an identity of a participant to which the gaze information is directed and/or otherwise indicate that the gaze information for a participant is away from a display device. Accordingly, the gaze information may include source-target pair information, Paragraph [0069], the compositor may provide a gaze detector an identity of the participant);
processing, by the processing circuit, the first real-time image of the first user to generate a first user information, wherein the first user information comprises an eye gaze direction of the first user if the identity or the status of the first user meets a requirement (see Sommerlade Paragraph [0007], providing gaze information to a gaze coordinator, the gaze information including an identifier associated with the user and an identifier associated with the participant, Paragraph [0062], The gaze tracker 404 may receive one or more images 420 from an image sensor of a camera, for example camera 126. The gaze tracker 404 may take the received one or more images 420, and extract one or more features from the image 420 using the feature extractor 424. For example the feature extractor 424 may determine and/or detect a user's face and extract feature information such as, but not limited to, a location of a user's, eyes, pupils, nose, chin, ears etc. In examples, the extracted information may be provided to a neural network model 428, where the neural network model may provide gaze information as an output. In examples, the neural network model may include but is not limited to a transformer model, a convolutional neural network model, and/or a support vector machine model. The gaze information may include coordinates, (e.g., x,y,z coordinates) of a participant's gaze in relation to an origin point on a display associated with a computing device. In examples, a compositor 432 residing at a same computing system as the gaze tracker 404 for example, may provide an identity of a participant to which the gaze information is directed and/or otherwise indicate that the gaze information for a participant is away from a display device. Accordingly, the gaze information may include source-target pair information, Paragraph [0070], In examples, the gaze coordinator may determine what gaze information is to be provided to which computing device based at least on a user associated with the computing device. Thus, the gaze coordinator may provide gaze information for a first plurality of participants displayed at a layout of a computing system for a first user, and information for a second plurality of participants displayed at a layout of a computing system for a second user);
controlling, by the processing circuit, the transmission circuit to transmit the first user information to the second video communication device (see Sommerlade Paragraph [0070], In examples, the gaze coordinator may determine what gaze information is to be provided to which computing device based at least on a user associated with the computing device. Thus, the gaze coordinator may provide gaze information for a first plurality of participants displayed at a layout of a computing system for a first user, and information for a second plurality of participants displayed at a layout of a computing system for a second user, Paragraph [0063], The gaze coordinator 408 may be the same as or similar to the gaze coordinator 222 of FIG. 2. The gaze coordinator 408 may receive a plurality of gaze information 436A, 436B. A gaze configurator 440 may determine specific gaze parameters for each participant, store such information in a storage 444, and then provide said information as stream specific adjustment parameters to a gaze adjustor 414. In examples, the specific parameters for each participant may include but is not limited to source-pair information, gaze adjustment features information (e.g., eye gaze, pose, position, body movement etc.) and/or whether a gaze associated with a participant requires adjustment. In some examples, a gaze associated with one participant displayed at a first display device may not need adjusting while a gaze associated with the same participant but displayed at a second different device may need adjusting. Accordingly, the adjustment parameters may be specific to each stream);
receiving, by the transmission circuit, the second video signal comprising the adjusted real-time image with an adjusted field of view, wherein the adjusted field of view is directed to a focus area of the video conference screen corresponding to the eye gaze direction of the first user (see Sommerlade Paragraph [0024], The present techniques provide real-time video modification to adjust participants' gaze during video communications. “Gaze” as used herein is the video representation of a user to show directional viewing of that user. Consequently, and more specifically, the present techniques adjust the video representations of a user's eye gaze, head pose, body portions of participants in real-time such that participants may appear to be attentive and engaged during the video communication. As a result, such techniques increase the quality of human communication that can be achieved via digital live and/or recorded video sessions, Paragraph [0025], In various examples, the gaze adjustment techniques described herein involve capturing a video stream of a user's face and making adjustments to the images within the video stream such that the direction of the user's eye gaze is adjusted and/or the user's head pose and/or body is adjusted. In some examples, the gaze adjustments described herein are provided, at least in part, by modifying the images within the video stream to synthesize images including specific eye gaze locations and/or specific head poses and/or body movements, Paragraph [0034], the presenter (e.g., 108) may perceive the adjusted eye gaze, head pose, upper and/or full body adjustments as a more natural representation of the users 110 and 112; further, body language in eye gaze, head pose, upper and/or full body positions may be communicated to the user 108. More specifically, as user 110 is focusing on or otherwise exhibiting an eye gaze in the direction of the video representation of user 112, and user 112 is focusing on or otherwise exhibiting an eye gaze in the direction of the video representation of user 110; accordingly, the video representation of users 110 and 112 as displayed by the display device 116 of computing system 102, may be adjusted such that the eyes, head pose, and/or other body positions are reflective of the user's attention and/or gaze. That is, the images of the user 110 and 112 may be adjusted such that users 110 and 112 appear to be looking at one another, Paragraph [0040], The gaze coordinator 222 may identify those video streams needing adjustments and may send to each participant, adjustment information including details for how the incoming video streams will need to be adjusted on the respective participant's device. The gaze coordinator 222 may also identify video streams that do not need adjustment or need adjustment in the same manner for all or a sufficient majority of participants. The gaze adjuster 224 can then perform this adjustment for the corresponding set of participants. In some examples, the gaze coordinator 222 adjusts an image of a participant who is not looking at the screen (or who has their video window obscured, on a second screen or hidden), such that the other participants do or do not become aware that said user is not looking directly ahead at a camera for instance, Paragraph [0041], the gaze adjuster 224 may adjust an image of participants such that the gaze of the participants appear to be directed to another participant as provided by the source-target pair, Paragraph [0054], as the gaze adjustment may also adjust pose and/or upper/total body position, the rotation of one or more body features of the image at the location 336 depicting the second participant 310 provides additional non-verbal information to the third participant viewing the conferencing session that the first participant 302 appears to have the attention of the second participant 310, Figure 7, generate gaze-adjusted images based on the desired eye gaze direction of the first participant);
controlling, by the processing circuit, the display to display the second real-time image with the adjusted field of view in the video conference screen (see Sommerlade Figure 7, generate gaze-adjusted images based on the desired eye gaze direction of the first participant, and replace the images within the video stream with the gaze-adjusted images, Paragraph [0072], The method 700 may then proceed to 710 and generate gaze-adjusted images based on the desired eye gaze direction of the first participant. In examples, an image from an original video stream (e.g., a non-gaze adjusted video stream) may be received and provided to a neural network. For example, the neural network model 456 may receive the image and generate a gaze-adjusted image. At 712, the generated gaze adjusted image may replace the original image(s) within the video stream);
wherein the second real-time image with the adjusted field of view is transmitted to the participants of the online video conference (see Sommerlade Figure 7, generate gaze-adjusted images based on the desired eye gaze direction of the first participant, and replace the images within the video stream with the gaze-adjusted images, Paragraph [0072], The method 700 may then proceed to 710 and generate gaze-adjusted images based on the desired eye gaze direction of the first participant. In examples, an image from an original video stream (e.g., a non-gaze adjusted video stream) may be received and provided to a neural network. For example, the neural network model 456 may receive the image and generate a gaze-adjusted image. At 712, the generated gaze adjusted image may replace the original image(s) within the video stream, Paragraph [0064], A neural network model 456 may receive the location information and adjustment parameters 452 and perform a gaze adjustment as previously described. The adjusted image may then be provided as output for display in the displayed layout at the identified location, Paragraph [0090], receiving, at computing system, image adjustment information associated with a video stream including images of a first participant; identifying, for a display layout of a communication application, a location displaying the images of the first participant; determining, based on the received image adjustment information, a location displaying images of a second participant for the display layout, the received image adjustment information indicating that an eye gaze of the first participant being directed toward the second participant; computing an eye gaze direction of the first participant based on the location displaying images of the second participant; generating gaze-adjusted images based on the eye gaze direction of the first participant, wherein the gaze-adjusted images include at least one of an adjusted eye gaze direction of the first participant or an adjusted head pose of the first participant; and replacing the images within the video stream with the gaze-adjusted images);
wherein the eye gaze direction is a direction which the first user looks at the second real-time image (see Sommerlade Paragraph [0024], The present techniques provide real-time video modification to adjust participants' gaze during video communications. “Gaze” as used herein is the video representation of a user to show directional viewing of that user. Consequently, and more specifically, the present techniques adjust the video representations of a user's eye gaze, head pose, body portions of participants in real-time such that participants may appear to be attentive and engaged during the video communication. As a result, such techniques increase the quality of human communication that can be achieved via digital live and/or recorded video sessions, Paragraph [0025], In various examples, the gaze adjustment techniques described herein involve capturing a video stream of a user's face and making adjustments to the images within the video stream such that the direction of the user's eye gaze is adjusted and/or the user's head pose and/or body is adjusted. In some examples, the gaze adjustments described herein are provided, at least in part, by modifying the images within the video stream to synthesize images including specific eye gaze locations and/or specific head poses and/or body movements, Paragraph [0034], the presenter (e.g., 108) may perceive the adjusted eye gaze, head pose, upper and/or full body adjustments as a more natural representation of the users 110 and 112; further, body language in eye gaze, head pose, upper and/or full body positions may be communicated to the user 108. More specifically, as user 110 is focusing on or otherwise exhibiting an eye gaze in the direction of the video representation of user 112, and user 112 is focusing on or otherwise exhibiting an eye gaze in the direction of the video representation of user 110; accordingly, the video representation of users 110 and 112 as displayed by the display device 116 of computing system 102, may be adjusted such that the eyes, head pose, and/or other body positions are reflective of the user's attention and/or gaze. That is, the images of the user 110 and 112 may be adjusted such that users 110 and 112 appear to be looking at one another, Paragraph [0040], The gaze coordinator 222 may identify those video streams needing adjustments and may send to each participant, adjustment information including details for how the incoming video streams will need to be adjusted on the respective participant's device. The gaze coordinator 222 may also identify video streams that do not need adjustment or need adjustment in the same manner for all or a sufficient majority of participants. The gaze adjuster 224 can then perform this adjustment for the corresponding set of participants. In some examples, the gaze coordinator 222 adjusts an image of a participant who is not looking at the screen (or who has their video window obscured, on a second screen or hidden), such that the other participants do or do not become aware that said user is not looking directly ahead at a camera for instance, Paragraph [0041], the gaze adjuster 224 may adjust an image of participants such that the gaze of the participants appear to be directed to another participant as provided by the source-target pair, Paragraph [0054], as the gaze adjustment may also adjust pose and/or upper/total body position, the rotation of one or more body features of the image at the location 336 depicting the second participant 310 provides additional non-verbal information to the third participant viewing the conferencing session that the first participant 302 appears to have the attention of the second participant 310, Figure 7, generate gaze-adjusted images based on the desired eye gaze direction of the first participant, Paragraph [0090], receiving, at computing system, image adjustment information associated with a video stream including images of a first participant; identifying, for a display layout of a communication application, a location displaying the images of the first participant; determining, based on the received image adjustment information, a location displaying images of a second participant for the display layout, the received image adjustment information indicating that an eye gaze of the first participant being directed toward the second participant; computing an eye gaze direction of the first participant based on the location displaying images of the second participant; generating gaze-adjusted images based on the eye gaze direction of the first participant, wherein the gaze-adjusted images include at least one of an adjusted eye gaze direction of the first participant or an adjusted head pose of the first participant; and replacing the images within the video stream with the gaze-adjusted images).
Sommerlade does not expressively teach
controlling, by the processing circuit, the transmission circuit to connect a server and join an online video conference as one of a plurality of participants of the online video conference;
However, Agrawal teaches
controlling, by the processing circuit, the transmission circuit to connect a server and join an online video conference as one of a plurality of participants of the online video conference (see Agrawal Paragraph [0048], wireless network 150 can include one or more servers 190 that support exchange of wireless data and video and other communication between electronic device 100 and second electronic device 300, Paragraph [0041], Communication module 138 includes program code that is executed by processor 102 to enable electronic device 100 to communicate with other external devices and systems, and Paragraph [0126], Method 1000 begins at the start block 1002 (video communication begins and users join). At block 1004, processor 102 detects that video data 270 is being captured by the current active camera);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a video conferencing device that tracks a user’s gaze and uses it to adjust and display the remote participant’s video field of view in real time (as taught in Sommerlade), with joining a video conference on a device via a server under processing circuit control (as taught in Agrawal), the motivation being to provide real-time interaction and communication in which eye gaze information is determined during a conference, and allows the displayed video to update dynamically as the user’s attention changes (see Agrawal Paragraphs [0048] and [0126]).
Regarding Claim 2, Sommerlade in view of Agrawal teaches
The network video communication device of claim 1, wherein the image capture circuit comprises a plurality of photographic lenses for capturing a plurality of images, wherein the processing circuit selects at least one of the captured images to be processed and generates the first real-time image (see Agrawal Paragraph [0003], With communication devices that have multiple rear facing cameras, the rear facing cameras can have lenses that are optimized for various focal angles and distances. For example, one rear facing camera can have a wide angle lens, another rear facing camera can have a telephoto lens, and an additional rear facing camera can have a macro lens and Paragraph [0028], The electronic device further includes at least one processor that is communicatively coupled to each of the at least one front facing camera, each of the at least one rear facing camera, and to the memory. The at least one processor executes program code of the CSCM, which enables the electronic device to capture, via the at least one front facing camera, a first image stream containing a face of a first user. A first eye gaze direction of the first user is determined based on a first image retrieved from the first image stream. The first eye gaze direction corresponds to a first location where the first user is looking. In response to determining that the first user is looking away from the front surface of the electronic device and towards a direction within a field of view (FOV) of the at least one rear facing camera, the processor selects an active camera that corresponds to one of the at least one rear facing cameras with a FOV containing the first location to which the first user is looking).
Regarding Claim 8, Sommerlade in view of Agrawal teaches
The network video communication device of claim 1, wherein the identity or the status of the first user meets the requirement when the identity of the first user is a host of the online video conference or the status of the first user is current speaker of the online video conference (see Agrawal Paragraph [0084], communication input 220 (FIG. 2) can be received contemporaneously with the detection of the eye-gaze direction of the user and used to generate context identifying data 254. In another example illustrated by the first camera selecting row of context table, if the user is looking away from the display 130 in a direction (or towards a location) that is behind the electronic device, while the user's speech is analyzed by NLP/CE engine 212 to provide context identifying data 254 including context identifier 510 as “self-image” and context type 520 is “speaker”, the selected camera would remain front facing camera 132A as indicated by an “X” in the box under the front facing camera 132A in the first row of context table 252. This prevents the electronic device from incorrectly switching to a rear camera when the user becomes distracted with something in the background that the user does not intend to share with the second user. As an example of such a distraction, a child or pet coming into a room while the user is on a conference call, would cause the user's eye gaze to shift away from the device display; However, the user does not want to show to the other participants on the video conference the child or dog that has entered into the FOV of the rear camera(s), in which first user’s speech is used to identify they are a speaker, and continue to use front camera displaying the user’s face).
Regarding Claim 9, it is rejected similarly as Claim 1.
Regarding Claim 10, Sommerlade in view of Agrawal teaches
The network video communication device of claim 9, wherein the image capture circuit comprises a plurality of photographic lenses for capturing a plurality of images (see Agrawal Paragraph [0003], With communication devices that have multiple rear facing cameras, the rear facing cameras can have lenses that are optimized for various focal angles and distances. For example, one rear facing camera can have a wide angle lens, another rear facing camera can have a telephoto lens, and an additional rear facing camera can have a macro lens and Paragraph [0028], The electronic device further includes at least one processor that is communicatively coupled to each of the at least one front facing camera, each of the at least one rear facing camera, and to the memory. The at least one processor executes program code of the CSCM, which enables the electronic device to capture, via the at least one front facing camera, a first image stream containing a face of a first user. A first eye gaze direction of the first user is determined based on a first image retrieved from the first image stream. The first eye gaze direction corresponds to a first location where the first user is looking. In response to determining that the first user is looking away from the front surface of the electronic device and towards a direction within a field of view (FOV) of the at least one rear facing camera, the processor selects an active camera that corresponds to one of the at least one rear facing cameras with a FOV containing the first location to which the first user is looking), wherein the processing circuit selects at least one of the captured images according to the first user information to merge into the second real-time image (see Agrawal Paragraph [0090] Referring to FIG. 6C, electronic device 100 is further shown continuing video communication session 386 by user 610. In FIG. 6C, user 610 has looked away from front surface 176 and display 130 in a second eye gaze direction 260B towards dog 626. Dog 626 is located in second location 262B, which is offset to the rear of electronic device 100 at a distance 644 away from electronic device 100. Dog 626 is located facing rear surface 180 and is in a FOV 646 of at least one of rear facing cameras 133 A-C (FIG. 1C). A secondary display device 648 is communicatively coupled to electronic device 100. In one embodiment, electronic device 100 can transmit captured video data 270 to secondary display device 648 for presentation of images/video thereon, Paragraph [0091] CSCM 136 enables electronic device 100 to capture, via front facing camera 132A, an image stream 240 including second image 244 containing a face 614 of user 610 and determine second eye gaze direction 260B, by processor 102 processing second image 244 retrieved from image stream 240. Second eye gaze direction 260B corresponds to second location 262B where the user is looking. In one or more embodiments, CSCM 136 enables electronic device 100 to determine second location 262B at least partially based on second eye gaze direction 260B and on camera parameters 264, including distances to objects within a FOV, and Paragraph [0092], CSCM 136 further enables processor 102 of electronic device 100 to determine if the user is looking away from the front surface 176 and towards a location behind electronic device 100 (i.e., towards a location that can be captured within a FOV of at least one of the rear facing cameras) for more than a threshold amount of time. In response to processor 102 determining that the user's eye gaze is in a direction within the FOV of at least one of the rear cameras, and in part based on a context determined from processing, by NLP/CE engine, of detected speech of the user, processor 102 triggers ICD controller 134 (FIG. 1A) to select as an active camera, a corresponding one of at least one rear facing cameras 133A-133C with a FOV 646 towards second eye gaze direction 260B and containing second location 262B to which the user 610 is looking. In the example of FIG. 6C, rear facing main camera 133A can be selected as the active camera with a FOV 646 that includes dog 626 when the user says “my dog is so adorable” contemporaneously with fixing his/her gaze in direction of dog 626. CSCM 136 further enables electronic device 100 to activate rear facing main camera 133A and to capture images (e.g., image/video data 270) within the FOV 646 of the activated camera).
Regarding Claim 13, Sommerlade in view of Agrawal teaches
The network video communication device of claim 9, wherein the first identity of the remote user is a host of the online video conference, and the first status of the remote user is current speaker in the online video conference (see Agrawal Paragraph [0084], communication input 220 (FIG. 2) can be received contemporaneously with the detection of the eye-gaze direction of the user and used to generate context identifying data 254. In another example illustrated by the first camera selecting row of context table, if the user is looking away from the display 130 in a direction (or towards a location) that is behind the electronic device, while the user's speech is analyzed by NLP/CE engine 212 to provide context identifying data 254 including context identifier 510 as “self-image” and context type 520 is “speaker”, the selected camera would remain front facing camera 132A as indicated by an “X” in the box under the front facing camera 132A in the first row of context table 252. This prevents the electronic device from incorrectly switching to a rear camera when the user becomes distracted with something in the background that the user does not intend to share with the second user. As an example of such a distraction, a child or pet coming into a room while the user is on a conference call, would cause the user's eye gaze to shift away from the device display; However, the user does not want to show to the other participants on the video conference the child or dog that has entered into the FOV of the rear camera(s), in which first user’s speech is used to identify they are a speaker, and continue to use front camera displaying the user’s face).
Regarding Claim 14, Sommerlade in view of Agrawal teaches
The network video communication device of claim 9, further comprising: a display,
wherein the transmission circuit receives a video signal including a remote real-time image and the processing circuit controls the display to display the remote real-time image in a video conference screen (see Sommerlade Paragraph [0067], The method starts and flow proceeds to 602. At 602, the method may capture, via a camera of a computing system, a video stream comprising images of a user of the computing system. For example, a camera, such as camera 126 (FIG. 1) may acquire a video stream of user 112, Paragraph [0066], aspects of the method 600 are performed by one or more processing devices, such as a computing system or server. Further, the method 600 can be performed by gates or circuits associated with a processor, Application Specific Integrated Circuit (ASIC), a field programmable gate array (FPGA), a system on chip (SOC), a neural processing unit, or other hardware device. Hereinafter, the method 600 shall be explained with reference to the systems, components, modules, software, data structures, user interfaces, etc. described in conjunction with FIGS. 1-5, Paragraph [0043], a number of program modules and data files may be stored in the system memory 204. While executing on the processing unit 202, the program modules 206 (e.g., software applications 220) may perform processes including, but not limited to, the aspects, as described herein, Paragraph [0064], Accordingly, the gaze adjuster 412 may operate in or near real-time such that a gaze of a participant may be adjusted in or near real-time without consuming resources traditionally expended by a CPU, Paragraph [0049], As depicted in FIG. 3A, a first participant 302 may be viewing a conferencing session at a display device 304. The display device 304 may include an integrated camera 306 for acquiring a video stream, an image, or a stream of images of the participant 302. A gaze tracker, for example the gaze tracker 221, may utilize one or more of the images to determine that the gaze 308 of the participant 302 is directed to an image or video representation of the participant 310. A second participant 310 may be viewing the same conferencing session at a display device 312. The display device 312 may include an integrated camera 314 for acquiring a video stream, an image, or a stream of images of the participant 310, Figures 3A – 3D, in which displays show a video conference interface including video streams of multiple participants).
Regarding Claim 16, it is rejected similarly as Claim 1. The method can be found in Agrawal (Abstract, method).
Regarding Claim 21, it is rejected similarly as Claim 8. The method can be found in Agrawal (Abstract, method).
Regarding Claim 22, Sommerlade in view of Agrawal teaches
The method of video conference image processing of claim 16, wherein the step of displaying the video conference screen including the second real-time image with the adjusted field of view comprises:
receiving a priority for the second real-time image with the adjusted field of view (see Agrawal Paragraph [0119], With specific reference to FIG. 9A, method 900 begins at the start block 902. At block 904, processor 102 detects that video data 270 is being captured by the current active camera (i.e., one of front facing cameras 132A-132B or rear facing cameras 133A-133C). Processor 102 triggers front facing camera 132A to capture an image stream 240 including second image 244 containing a face 614 of user 610 (block 906). Processor 102 receives image stream 240 including second image 244 (block 908). Processor 102 determines second eye gaze direction 260B and corresponding second location 262B based on second image 244 retrieved from the image stream 240 (block 910). Processor 102 stores second eye gaze direction 260B and second location 262B to system memory 120 (block 911). The second eye gaze direction 260B corresponds to second location 262B where the user is looking. The second location 262B is determined at least partially based on second eye gaze direction 260B, Paragraph [0120], Processor 102 retrieves eye gaze threshold time 218 and starts timer 216 (block 912). Timer 216 tracks the amount of time that a user's eye gaze has rested/remained on a specific area that is away from the display 130 at front surface 176 of electronic device 100. Processor 102 determines if user 610 is looking away from the display 130 at front surface 176 of electronic device 100 and towards a direction within a FOV of the rear facing cameras (decision block 914). In response to determining that user 610 is not looking away from the display 130, processor 102 captures video data 270 using the current active camera (block 928). Method 900 then ends at end block 930. In response to determining that user 610 is looking away from the front surface 176 and towards a direction behind the rear surface 180 of electronic device 100, processor 102 determines if the value of timer 216 is greater than eye gaze threshold time 218 (decision block 916). Eye gaze threshold time 218 is a minimum amount of time that a user is looking away from front surface 176 and display 130 to trigger a switch of active cameras to one of the rear facing cameras. In response to determining that the value of timer 216 is not greater than eye gaze threshold time 218, processor 102 returns to decision block 914 to continue determining if user 610 is looking away from the display 130, and Paragraph [0121], In response to determining that the value of timer 216 is greater than eye gaze threshold time 218, processor 102 selects to activate as an active camera, a corresponding one of the rear facing cameras 133A-133C based on second location 262B and identified characteristics of the second location 262B (block 918). The selected rear facing camera has a FOV 646 towards second eye gaze direction 260B and containing second location 262B to which the user 610 is looking. If electronic device 100 has only one rear facing camera (i.e., rear facing camera 133A), processor 102 selects the single rear facing camera. Processor 102 retrieves camera parameters 264 and camera settings 266 for the selected rear facing camera (block 920), in which when there is a gaze that exceeds the timer, it is the priority gaze, thus the priority field of view to be displayed); and
arranging a screen area in the video conference screen corresponding to the received priority (see Agrawal Figure 11, capture video data of eye gaze location using active rear camera, transmit video to 2nd electronic device).
Claims 3 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Sommerlade et al. (U.S. Pub. No. 2022/0400228, hereinafter “Sommerlade”) in view of Agrawal et al. (U.S. Pub. No. 2023/0276115, hereinafter “Agrawal”) and Geerds (U.S. Pub. No. 2014/0267596).
Regarding Claim 3, Sommerlade in view of Agrawal teaches all the limitations of claim 1, but does not expressively teach
The network video communication device of claim 1, wherein the image capture circuit comprises an upper photographic lens, a bottom photographic lens, a middle photographic lens, a left photographic lens and a right photographic lens, and the image capture circuit covers a field of view more than 130 degrees horizontally and 105 degrees vertically.
However, Geerds teaches
The network video communication device of claim 1, wherein the image capture circuit comprises an upper photographic lens, a bottom photographic lens, a middle photographic lens, a left photographic lens and a right photographic lens, and the image capture circuit covers a field of view more than 130 degrees horizontally and 105 degrees vertically (see Geerds Figure 1, image capture circuit comprising 6 lenses item 110, in which there is a top lens, bottom lens, and 4 side lenses (could be considered middle, left, or right, depending on orientation) and Paragraph [0032], geometry of the disclosed camera system makes it possible to produce a fully 360 by 180 degrees, spherical image or fully spherical video).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a video conferencing device that joins a video conference via a server, tracks a user’s gaze and uses it to adjust and display the remote participant’s video field of view in real time (as taught in Sommerlade in view of Agrawal), with a camera that has multiple lenses in various directions that captures a very large and wide field of view (as taught in Geerds), the motivation being to address the issue of a lack of field of view provided by a singular camera, and ensure areas of interest are captured (see Geerds Paragraph [0003]).
Regarding Claim 11, it is rejected similarly as Claim 3.
Claims 4, 12 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Sommerlade et al. (U.S. Pub. No. 2022/0400228, hereinafter “Sommerlade”) in view of Agrawal et al. (U.S. Pub. No. 2023/0276115, hereinafter “Agrawal”) and Teshome et al. (U.S. Pub. No. 2016/0378183, hereinafter “Teshome”).
Regarding Claim 4, Sommerlade in view of Agrawal teaches all the limitations of claim 1, but does not expressively teach
The network video communication device of claim 1, wherein the step of processing, by the processing circuit, the first real-time image of the first user to generate the first user information comprises:
performing edge detection on the first real-time image for detecting a user face position;
detecting eye characteristics for determining a user eye position corresponding to the user face position on the first real-time image; and
determining the eye gaze direction according to the user eye position.
However, Teshome teachesThe network video communication device of claim 1, wherein the step of processing, by the processing circuit, the first real-time image of the first user to generate the first user information comprises:
performing edge detection on the first real-time image for detecting a user face position (see Teshome Paragraph [0074], the image processor 130 may detect a position of an eye using a method (for example, Haar-like features) of finding object features from the image. As a result, the image processor 130 may detect the iris and the pupil from the position of the eye using an edge detection method, and detect the point of gaze of the user based on relative positions of the iris and the pupil);
detecting eye characteristics for determining a user eye position corresponding to the user face position on the first real-time image (see Teshome Paragraph [0074], the image processor 130 may detect a position of an eye using a method (for example, Haar-like features) of finding object features from the image. As a result, the image processor 130 may detect the iris and the pupil from the position of the eye using an edge detection method, and detect the point of gaze of the user based on relative positions of the iris and the pupil); and
determining the eye gaze direction according to the user eye position (see Teshome Paragraph [0074], the image processor 130 may detect a position of an eye using a method (for example, Haar-like features) of finding object features from the image. As a result, the image processor 130 may detect the iris and the pupil from the position of the eye using an edge detection method, and detect the point of gaze of the user based on relative positions of the iris and the pupil);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a video conferencing device that joins a video conference via a server, tracks a user’s gaze and uses it to adjust and display the remote participant’s video field of view in real time (as taught in Sommerlade in view of Agrawal), with determining a user’s gaze using edge detection to find positions of the user’s eyes (as taught in Teshome), the motivation being to provide a solution to creating a more realistic meeting experience in a virtual communication, by applying a process that determines a user’s gaze in order to implement a correction if needed (see Teshome Paragraphs [0005] – [0008]).
Regarding Claim 12, Sommerlade in view of Agrawal and Teshome teach
The network video communication device of claim 9, wherein the network video communication device further performs following step:
performing, by the processing circuit, edge detection on the second real-time image with the adjusted field of view (see Teshome Paragraph [0074], the image processor 130 may detect a position of an eye using a method (for example, Haar-like features) of finding object features from the image. As a result, the image processor 130 may detect the iris and the pupil from the position of the eye using an edge detection method, and detect the point of gaze of the user based on relative positions of the iris and the pupil) and determining an object-of-interest in the second real-time image with the adjusted field of view (see Agrawal Figure 10, in which process 1000 can be repeated to track the eye gaze direction of the user, Paragraph [0066], Camera settings 266 are values and characteristics that can change during the operation of cameras 132A-132B and 133A-133C to capture images by the cameras. In one embodiment, camera settings 266 can be determined by either processor 102 or by ICD controller 134. Camera settings 266 can include various settings such as aperture, shutter speed, iso level, white balance, zoom level, directional settings (i.e., region of interest (ROI)), distance settings, focus and others. Camera settings 266 can include optimal camera settings 268 that are selected to optimize the quality of the images captured by the cameras. Optimal camera settings 268 can include zoom levels, focus distance, and directional settings that allow images within a camera FOV and that are desired to be captured to be in focus, centered, and correctly sized. Optimal camera settings 268 can further include digital crop levels, focal distance of the focus module and directional audio zoom. The directional audio zoom enables microphone 108 to be tuned to receive audio primarily in the direction of a cropped FOV in a desired ROI, and Paragraph [0067], Image characteristics 272 can further include computer vision (CV), machine learning (ML) and artificial intelligence (AI) based techniques for determining an object of interest within the eye gaze area ROI); and
controlling, by the processing circuit, the image capture circuit to zoom-in and capture an enlarged image of the object-of-interest as the second real-time image with the adjusted field of view (see Agrawal Paragraph [0066], Camera settings 266 are values and characteristics that can change during the operation of cameras 132A-132B and 133A-133C to capture images by the cameras. In one embodiment, camera settings 266 can be determined by either processor 102 or by ICD controller 134. Camera settings 266 can include various settings such as aperture, shutter speed, iso level, white balance, zoom level, directional settings (i.e., region of interest (ROI)), distance settings, focus and others. Camera settings 266 can include optimal camera settings 268 that are selected to optimize the quality of the images captured by the cameras. Optimal camera settings 268 can include zoom levels, focus distance, and directional settings that allow images within a camera FOV and that are desired to be captured to be in focus, centered, and correctly sized. Optimal camera settings 268 can further include digital crop levels, focal distance of the focus module and directional audio zoom. The directional audio zoom enables microphone 108 to be tuned to receive audio primarily in the direction of a cropped FOV in a desired ROI).
Regarding Claim 17, it is rejected similarly as Claim 4. The method can be found in Agrawal (Abstract, method).
Claims 5 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Sommerlade et al. (U.S. Pub. No. 2022/0400228, hereinafter “Sommerlade”) in view of Agrawal et al. (U.S. Pub. No. 2023/0276115, hereinafter “Agrawal”), Teshome et al. (U.S. Pub. No. 2016/0378183, hereinafter “Teshome”) and Chakravarthula et al. (U.S. Pub. No. 2013/0057553, hereinafter “Chakravarthula”).
Regarding Claim 5, Sommerlade in view of Agrawal and Teshome teach all the limitations of claim 4, but does not expressively teach
The network video communication device of claim 4, wherein the step of processing, by the processing circuit, the first real-time image of the first user to generate the first user information further comprises:
calculating a user distance based on user face size information corresponding to the user face position on the first real-time image and focal length information of the image capture circuit.
However, Chakravarthula teaches
The network video communication device of claim 4, wherein the step of processing, by the processing circuit, the first real-time image of the first user to generate the first user information further comprises:
calculating a user distance based on user face size information corresponding to the user face position on the first real-time image and focal length information of the image capture circuit (see Chakravarthula Paragraph [0048], the focal length of the camera can be used to determine the distance of the user from the display, or alternatively the focal length can be combined with detected features such as the size of the face or the relative size of facial features on the user to determine the distance of the user from the display).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a video conferencing device that joins a video conference via a server, determines a user’s gaze by using edge detection to locate the user’s eyes, and uses the gaze information to adjust and display the remote participant’s video field of view in real time (as taught in Sommerlade in view of Agrawal and Teshome), with calculating user distance (as taught in Chakravarthula), the motivation being to adapt a display to a user by adjusting item size on the display according to the distance between the user and display (see Chakravarthula Paragraph [0048]).
Regarding Claim 18, it is rejected similarly as Claim 5. The method can be found in Agrawal (Abstract, method).
Claims 6 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Sommerlade et al. (U.S. Pub. No. 2022/0400228, hereinafter “Sommerlade”) in view of Agrawal et al. (U.S. Pub. No. 2023/0276115, hereinafter “Agrawal”), Teshome et al. (U.S. Pub. No. 2016/0378183, hereinafter “Teshome”) and Gavino et al. (U.S. Pub. No. 2019/0042834, hereinafter “Gavino”).
Regarding Claim 6, Sommerlade in view of Agrawal and Teshome teach all the limitations of claim 4, but does not expressively teach
The network video communication device of claim 4, wherein the step of determining the eye gaze direction according to the user eye position comprises:
determining a horizontal face midline and a vertical face midline of a user face at the user face position; and
comparing the user's eye positions with the horizontal face midline and the vertical face midline of the user face.
However, Gavino teaches
The network video communication device of claim 4, wherein the step of determining the eye gaze direction according to the user eye position comprises:
determining a horizontal face midline and a vertical face midline of a user face at the user face position (see Gavino Paragraph [0069], The example position tracker 338 identifies a horizontal line representing one-third of the vertical measure from the top bounding line of the face bounding rectangle. In examples disclosed herein, the example position tracker 338 identifies the horizontal line by finding a line that is one-third of the distance from a top edge of the face bounding rectangle and a bottom edge of the face bounding rectangle. However, any other ratio for finding a horizontal line may additionally or alternatively be used. The example position tracker 338 calculates an intersection of the vertical centerline and the horizontal line to determine the estimated eye position); and
comparing the user's eye positions with the horizontal face midline and the vertical face midline of the user face (see Gavino Paragraph [0069], The example position tracker 338 identifies a horizontal line representing one-third of the vertical measure from the top bounding line of the face bounding rectangle. In examples disclosed herein, the example position tracker 338 identifies the horizontal line by finding a line that is one-third of the distance from a top edge of the face bounding rectangle and a bottom edge of the face bounding rectangle. However, any other ratio for finding a horizontal line may additionally or alternatively be used. The example position tracker 338 calculates an intersection of the vertical centerline and the horizontal line to determine the estimated eye position).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a video conferencing device that joins a video conference via a server, determines a user’s gaze by using edge detection to locate the user’s eyes, and uses the gaze information to adjust and display the remote participant’s video field of view in real time (as taught in Sommerlade in view of Agrawal and Teshome), with mapping a horizontal and vertical line on a user’s face and comparing the positions of their eyes to determine the user’s eye positions (as taught in Gavino), the motivation being to simply identify a position of a person’s eyes in order to generate corresponding data (see Gavino Paragraph [0069]).
Regarding Claim 19, it is rejected similarly as Claim 6. The method can be found in Agrawal (Abstract, method).
Claims 7 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Sommerlade et al. (U.S. Pub. No. 2022/0400228, hereinafter “Sommerlade”) in view of Agrawal et al. (U.S. Pub. No. 2023/0276115, hereinafter “Agrawal”) and Norris et al. (U.S. Pub. No. 2015/0373477, hereinafter “Norris”).
Regarding Claim 7, Sommerlade in view of Agrawal teaches all the limitations of claim 1, but does not expressively teach
The network video communication device of claim 1, further comprising:
a microphone, wherein the microphone receives a first real-time audio local to the network video communication device, and the processing circuit processes the first real-time audio to determine a direction and the status of the first user.
However, Norris teaches
The network video communication device of claim 1, further comprising:
a microphone, wherein the microphone receives a first real-time audio local to the network video communication device, and the processing circuit processes the first real-time audio to determine a direction and the status of the first user (see Norris Paragraph [0112], an electronic device can intelligently assign locations for one or more sound localization points or virtual microphone points. Selection of the location can be based on, for example, available space near the listener, location of another person, previous assignments of sound localization or virtual microphone points, type or origin of the sound, environment in which the listener is located, objects near the person, a social status or personal characteristics of a person, a person with whom the listener is communicating, time of arrival or reservation or other time-related property, etc.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a video conferencing device that joins a video conference via a server, tracks a user’s gaze and uses it to adjust and display the remote participant’s video field of view in real time (as taught in Sommerlade in view of Agrawal), with using a user’s audio to determine the direction and status of the user (as taught in Norris), the motivation being to provide a solution to virtual meetings feeling unnatural and impersonal, and use audio data to collect information (see Norris Paragraphs [0001] and [0112]).
Regarding Claim 20, it is rejected similarly as Claim 7. The method can be found in Agrawal (Abstract, method).
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Sommerlade et al. (U.S. Pub. No. 2022/0400228, hereinafter “Sommerlade”) in view of Agrawal et al. (U.S. Pub. No. 2023/0276115, hereinafter “Agrawal”) and Cranfill et al. (U.S. Pub. No. 2011/0249074, hereinafter “Cranfill”).
Regarding Claim 15, Sommerlade in view of Agrawal teaches all the limitations of claim 9, but does not expressively teach
The network video communication device of claim 9, wherein the network video communication device further performs the following steps:
controlling, by the processing circuit, the image capture circuit to capture the first real-time image with a previous field of view and the second real-time image with the adjusted field of view at the same time; and
transmitting, by the transmission circuit, the video signal including the first real-time image with the previous field of view and the second real-time image with the adjusted field of view to the at least one of the other participants of the online video conference.
However, Cranfill teaches
The network video communication device of claim 9, wherein the network video communication device further performs the following steps:
controlling, by the processing circuit, the image capture circuit to capture the first real-time image with a previous field of view and the second real-time image with the adjusted field of view at the same time (see Cranfill Paragraph [0673], selection of the "Select L1" button 7420 would cause the UI 7475 to display only the video captured by the local device's back camera (being presented in the foreground inset display 7410). Selection of the "Select L2" button 7425 would cause the UI 7475 to display only the video captured by the local device's front camera (being presented in the foreground inset display 7405). Selection of the "Select Both" button 7430 would cause the UI 7475 to continue displaying both videos captured by both cameras on the local device, and selecting the "Cancel" button 7485 would cancel the operation); and
transmitting, by the transmission circuit, the video signal including the first real-time image with the previous field of view and the second real-time image with the adjusted field of view to the at least one of the other participants of the online video conference (see Cranfill Paragraph [0687] and Figure 75, the set of selectable UI items includes: a "Transmit L1" item 7528 (e.g. button 7528); a "Transmit L2" item 7530 (e.g. button 7530); a "Transmit Both" item 7532 (e.g. button 7532); and a "Cancel" item 7534 (e.g. button 7534). In this example, selection of the "Transmit L1" button 7528 would cause the UI 7500 to transmit only the video captured by the device's back camera to the remote device during the video conference. Selection of the "Transmit L2" button 7530 would cause the UI 7500 to transmit only the video captured by the device's front camera to the remote device during the video conference. Selection of the "Transmit Both" button 7532 would cause the UI 7500 to transmit both videos captured by the device's front and back camera to the remote user for the video conference, and selecting the "Cancel" button 7534 would cancel the operation).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a video conferencing device that joins a video conference via a server, tracks a user’s gaze and uses it to adjust and display the remote participant’s video field of view in real time (as taught in Sommerlade in view of Agrawal), with transmitting two field of views from a camera in a communication session (as taught in Cranfill), the motivation being to environment a more inclusive and informational meeting to allow a user to display multiple views or items at once (see Cranfill Paragraph [0687], Paragraph [0001] –[0002] and Figure 75).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARISSA A JONES whose telephone number is (703)756-1677. The examiner can normally be reached Telework M-F 6:30 AM - 4:00 PM CT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CARISSA A JONES/Examiner, Art Unit 2691
/DUC NGUYEN/Supervisory Patent Examiner, Art Unit 2691