DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Allowable Subject Matter
Claim 7-16 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “second video conferencing device” in claims 18-19.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1-6,17-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Prokopenya et al (US 20190019063 A1) in view of Park et al (US 20230281885 A1)
Regarding claim 1, Prokopenya discloses Video conferencing method ([0066] exemplary computer system of the present invention may be utilized for various applications which may include, but not limited to, gaming, mobile-device games, video chats, video conferences, live video streaming, video streaming),
in which first video image data are reproduced by a first video conferencing device by means of a first display device and at least a region of the head of a first user comprising the eyes is captured by a first image capture device in a position in which the first user is looking at the video image data reproduced by the first display device ([0048] imaging device of the computing device 102 may be provided via either a peripheral eye tracking camera or as an integrated a peripheral eye tracking camera in backlight system.),
the video image data reproduced by the first display device comprising at least a depiction of the eyes of a second user ([0032] As used herein, the term “user” shall have a meaning of at least one user.) captured by a second image capture device of a second video conferencing device arranged remotely from the first video conferencing device ([0033] exemplary inventive computer system of the present invention may transmit the acquired visual data to a remote server for processing, or in other implementations, may process the acquired visual data in the real-time on a computing device (e.g., mobile device, computer, etc.)., [0047] computer system environment incorporating certain embodiments of the present invention. As shown in FIG. 1, the inventive environment may include a user (101), who uses, for example, a first computing device (102) (e.g., mobile phone device, laptop, etc.), a server (103) and a second computing device (104) (e.g., mobile phone device, laptop, etc.)) ;
a processing unit receives and modifies the video image data of at least the region of the head of the first user comprising the eyes, captured by the first image capture device, and the modified video image data are transmitted to and reproduced by a second display device of the second video conferencing device ([0034] , examples of visual transformation effects may be, without limitation: [0035] turn user's face into a face of an animal, [0036] turn user's face into a face of another user, [0037] a race transformation, [0038] a gender transformation, [0039] an age transformation (making people look younger or older), [0040] choose an object which may be closest to the subject's appearance (based on the machine-learning algorithm's logic), [0041] swap parts of subject's head, [0042] make drawings on the subject's face or/and head, [0043] intentionally deform user's face, [0044] utilize dynamic masks (e.g., changing eye-gaze direction, eyebrow motion, opening mouth, etc.), and [0045] change user's appearance based on some predefined logic (e.g., changing user's appearance in a such way, as if they were poor, rich, live in the Stone Age, in future, astronauts, etc.).),
the modified video image data are generated by a Generative Adversarial Network (GAN) with a generator network and a discriminator network ([0033] utilizes the at least one conditional generative adversarial neural network which is programmed/configured to perform visual appearance transformations), and
the generator network generating modified video image data and the discriminator network evaluating a similarity between the depiction of the head of the first user in the modified video image data and the captured video image data and also evaluating a match between the direction of gaze of the first user in the modified video image data and the target direction of gaze ([0044] utilize dynamic masks (e.g., changing eye-gaze direction, eyebrow motion, opening mouth, etc.), [0068] presenting the at least one first neural network with at least one training real representation of the plurality of training real representations and one or more candidate meta-parameters of one or more latent variables of the at least one first visual effect to incorporate the at least one first visual effect into the at least one portion of the at one first real subject of the at least one training real representation to generate at least one first training photorealistic-imitating synthetic representation of the at least one portion of the at least one first real subject with the at least one first visual effect; ii) presenting the at least one second neural network with (1) the at least one first training photorealistic-imitating synthetic representation of the at least one first portion of the at least one first real subject with the at least one first visual effect and (2) the at least one training synthetic representation having the at least one visual effect applied to the at least one portion of the at least one synthetic subject to determine one or more actual meta-parameters of the one or more latent variables of the at least one first visual effect, where the one or more actual meta-parameters are meta-parameters at which the at least one second neural network has identified that the at least one first training photorealistic-imitating synthetic representation of the at least one portion of the at least one first real subject with the at least one first visual effect to be realistic).
Park discloses wherein the direction of gaze of the first user is detected during the processing of the video image data and, in the video image data, at least the reproduction of the region of the head of the first user comprising the eyes is then modified so that a target direction of gaze of the first user depicted in the modified video image data appears as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device during the processing of the video image data ([0040] imaging system identifies that a gaze of the first user as represented in the image data is directed toward a displayed representation of at least a portion of (e.g., a face) of a second user. For instance, the imaging system can determine that the gaze of the first user is on an upper-right corner of a display screen of the first user’s device while the upper-right corner of the display screen of the first user’s device is displaying an image of a second user - and therefore that the gaze of the first user is on the second user. [0065] modified image data 242 appears straightened relative to the image data 232 (e.g., no longer skewed or tilted clockwise), and the face of the user 215 in the modified image data 242 appears to face toward the camera that captured the image data)
Prokopenya and Park are combinable because they are from the same field of invention.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify computer system of Prokopenya to include wherein the direction of gaze of the first user is detected during the processing of the video image data and, in the video image data, at least the reproduction of the region of the head of the first user comprising the eyes is then modified so that a target direction of gaze of the first user depicted in the modified video image data appears as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device during the processing of the video image data as described by Park
The motivation for doing so would have been to direct gaze of the first user as represented in the image data toward a displayed representation of at least a portion (e.g., a face) of a second user (Park, [0004]).
Therefore, it would have been obvious to combine Prokopenya and Park to obtain the invention as specified in claim 1.
Regarding claim 2, Prokopenya discloses wherein it is determined by means of the detected direction of gaze of the first user whether the first user is looking at a point of the first display device, and, if it has been determined that a point of the first display device is being looked at, it is determined which object is currently being depicted at this point by the first display device ([0044] utilize dynamic masks (e.g., changing eye-gaze direction, eyebrow motion, opening mouth, etc.),).
Regarding claim 3, Prokopenya is silent to wherein if it has been determined that the object is the depiction of the face of the second user, when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data appears to be such that the first user is looking at the face of the second user depicted on the first display device.
Park discloses wherein if it has been determined that the object is the depiction of the face of the second user, when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data appears to be such that the first user is looking at the face of the second user depicted on the first display device. ([0040] imaging system identifies that a gaze of the first user as represented in the image data is directed toward a displayed representation of at least a portion of (e.g., a face) of a second user. For instance, the imaging system can determine that the gaze of the first user is on an upper-right corner of a display screen of the first user’s device while the upper-right corner of the display screen of the first user’s device is displaying an image of a second user - and therefore that the gaze of the first user is on the second user.)
Prokopenya and Park are combinable because they are from the same field of invention.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify computer system of Prokopenya to include wherein if it has been determined that the object is the depiction of the face of the second user, when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data appears to be such that the first user is looking at the face of the second user depicted on the first display device as described by Park
The motivation for doing so would have been to direct gaze of the first user as represented in the image data toward a displayed representation of at least a portion (e.g., a face) of a second user (Park, [0004]).
Therefore, it would have been obvious to combine Prokopenya and Park to obtain the invention as specified in claim 3.
Regarding claim 4, Prokopenya is silent to wherein if it has been determined that the object is the depiction of the face of the second user, but it has not been determined which region of the depiction of the face is being looked at, when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data appears such that the first user is looking at an eye of the second user depicted on the first display device.
Park discloses wherein if it has been determined that the object is the depiction of the face of the second user, but it has not been determined which region of the depiction of the face is being looked at, when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data appears such that the first user is looking at an eye of the second user depicted on the first display device. ([0040] he imaging system uses one or more trained machine learning models to generate the modified image data based on the image data, the gaze, and/or the arrangement. For instance, if the second user is viewing the display, then in the modified image data, the first user can face the position of the second user while the second user views the display.)
Prokopenya and Park are combinable because they are from the same field of invention.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify computer system of Prokopenya to include wherein if it has been determined that the object is the depiction of the face of the second user, but it has not been determined which region of the depiction of the face is being looked at, when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data appears such that the first user is looking at an eye of the second user depicted on the first display device as described by Park
The motivation for doing so would have been to direct gaze of the first user as represented in the image data toward a displayed representation of at least a portion (e.g., a face) of a second user (Park, [0004]).
Therefore, it would have been obvious to combine Prokopenya and Park to obtain the invention as specified in claim 4.
Regarding claim 5, Prokopenya is silent to wherein the video image data reproduced by the first display device comprise at least a depiction of the eyes of a plurality of second users captured by the second image capture device and/or further second image capture devices,
it is determined whether the object is a depiction of the face of a particular one of the plurality of second users,
when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data then appears as if the first image capture device were arranged on the straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the
eyes of the particular one of the plurality of second users depicted on the first display device.
Park discloses wherein the video image data reproduced by the first display device comprise at least a depiction of the eyes of a plurality of second users captured by the second image capture device and/or further second image capture devices ([0042] the scene 110 can be a scene of one or both of the user’s eyes, and/or at least a portion of the user’s face.),
it is determined whether the object is a depiction of the face of a particular one of the plurality of second users ([0039] an image of a first user may be displayed in an upper-left corner of a second user’s screen),
when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data then appears as if the first image capture device were arranged on the straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the particular one of the plurality of second users depicted on the first display device. ([0040] imaging system can determine that the gaze of the first user is on an upper-right corner of a display screen of the first user’s device while the upper-right corner of the display screen of the first user’s device is displaying an image of a second user - and therefore that the gaze of the first user is on the second user)
Prokopenya and Park are combinable because they are from the same field of invention.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify computer system of Prokopenya to include wherein the video image data reproduced by the first display device comprise at least a depiction of the eyes of a plurality of second users captured by the second image capture device and/or further second image capture devices, it is determined whether the object is a depiction of the face of a particular one of the plurality of second users, when the video image data are processed, the target direction of gaze of the first user depicted in the modified video image data then appears as if the first image capture device were arranged on the straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the particular one of the plurality of second users depicted on the first display device as described by Park
The motivation for doing so would have been to direct gaze of the first user as represented in the image data toward a displayed representation of at least a portion (e.g., a face) of a second user (Park, [0004]).
Therefore, it would have been obvious to combine Prokopenya and Park to obtain the invention as specified in claim 5.
Regarding claim 6, Prokopenya discloses wherein the video image data captured by the first image capture device comprise at least a depiction of the head of the first user, the pose of the head of the first user is determined in the captured video image data and the direction of gaze of the first user is detected from the determined pose of the head of the first user ([0047] the user (101) may interact with the computing device (102) by means of its camera, which may take one or series of frames (e.g., still images, video frames, etc.), containing one or more visual representations of the user (e.g., user's head imagery, user's half-body imagery, user's full-body imagery, etc.).).
Regarding claim 17, Prokopenya discloses Video conferencing method ([0066] exemplary computer system of the present invention may be utilized for various applications which may include, but not limited to, gaming, mobile-device games, video chats, video conferences, live video streaming, video streaming),
wherein first video image data are reproduced by a first video conferencing device by means of a first display device and at least a region of the head of a first user comprising the eyes is captured by a first image capture device in a position in which the first user is looking at the video image data reproduced by the first display device, the video image data reproduced by the first display device ([0048] imaging device of the computing device 102 may be provided via either a peripheral eye tracking camera or as an integrated a peripheral eye tracking camera in backlight system.),
comprising at least a depiction of the eyes of a second user ([0032] As used herein, the term “user” shall have a meaning of at least one user.) captured by a second image capture device of a second video conferencing device arranged remotely from the first video conferencing device ([0033] exemplary inventive computer system of the present invention may transmit the acquired visual data to a remote server for processing, or in other implementations, may process the acquired visual data in the real-time on a computing device (e.g., mobile device, computer, etc.)., [0047] computer system environment incorporating certain embodiments of the present invention. As shown in FIG. 1, the inventive environment may include a user (101), who uses, for example, a first computing device (102) (e.g., mobile phone device, laptop, etc.), a server (103) and a second computing device (104) (e.g., mobile phone device, laptop, etc.)) ;
a processing unit receives and modifies the video image data of at least the region of the head of the first user comprising the eyes, captured by the first image capture device, and the modified video image data are transmitted to and reproduced by a second display device of the second video conferencing device ([0034] , examples of visual transformation effects may be, without limitation: [0035] turn user's face into a face of an animal, [0036] turn user's face into a face of another user, [0037] a race transformation, [0038] a gender transformation, [0039] an age transformation (making people look younger or older), [0040] choose an object which may be closest to the subject's appearance (based on the machine-learning algorithm's logic), [0041] swap parts of subject's head, [0042] make drawings on the subject's face or/and head, [0043] intentionally deform user's face, [0044] utilize dynamic masks (e.g., changing eye-gaze direction, eyebrow motion, opening mouth, etc.), and [0045] change user's appearance based on some predefined logic (e.g., changing user's appearance in a such way, as if they were poor, rich, live in the Stone Age, in future, astronauts, etc.).),
Park discloses wherein the direction of gaze of the first user is detected during the processing of the video image data and, in the video image data, at least the reproduction of the region of the head of the first user comprising the eyes is then modified so that a target direction of gaze of the first user depicted in the modified video image data appears as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device ([0040] imaging system identifies that a gaze of the first user as represented in the image data is directed toward a displayed representation of at least a portion of (e.g., a face) of a second user. For instance, the imaging system can determine that the gaze of the first user is on an upper-right corner of a display screen of the first user’s device while the upper-right corner of the display screen of the first user’s device is displaying an image of a second user - and therefore that the gaze of the first user is on the second user. [0065] modified image data 242 appears straightened relative to the image data 232 (e.g., no longer skewed or tilted clockwise), and the face of the user 215 in the modified image data 242 appears to face toward the camera that captured the image data)
successive video frames are captured by the first image capture device ([0114] image data include image data captured using the image capture and processing system 100, image data captured using image sensor(s) of the sensor(s) 230, image data captured using the first camera 330A, image data captured using the second camera 330B, image data captured using the third camera 330C, image data captured using the fourth camera 330D, image data captured using the first camera 430A, image data captured using the second camera 430B, image data captured using the third camera 430C, image data captured using the fourth camera 430D, an image used as input data for the input layer 610 of the NN 600, an image captured using the input device 845, another image described herein, another set of image data described herein, or a combination thereof.)
when the direction of gaze of the first user changes, some video frames are interpolated during the processing of the video image data in such a way that the change in direction of gaze reproduced by the modified video image data is slowed down ([0122] generating the modified image data as in operation 720 includes generating intermediate image data at least in part by modifying the image data to modify at least the portion of the first user in the image data to be visually directed toward a forward direction.)
Prokopenya and Park are combinable because they are from the same field of invention.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify computer system of Prokopenya to include wherein the direction of gaze of the first user is detected during the processing of the video image data and, in the video image data, at least the reproduction of the region of the head of the first user comprising the eyes is then modified so that a target direction of gaze of the first user depicted in the modified video image data appears as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device, successive video frames are captured by the first image capture device and when the direction of gaze of the first user changes, some video frames are interpolated during the processing of the video image data in such a way that the change in direction of gaze reproduced by the modified video image data is slowed down as described by Park
The motivation for doing so would have been to direct gaze of the first user as represented in the image data toward a displayed representation of at least a portion (e.g., a face) of a second user (Park, [0004]).
Therefore, it would have been obvious to combine Prokopenya and Park to obtain the invention as specified in claim 17.
Regarding claim 18, Prokopenya discloses Video conferencing system ([0066] exemplary computer system of the present invention may be utilized for various applications which may include, but not limited to, gaming, mobile-device games, video chats, video conferences, live video streaming, video streaming) comprising
a first video conferencing device having a first display device and a first image capture device, the first image capture device being arranged to capture at least a region of the head of a first user, said region comprising the eyes, in a position in which the first user is looking at the video image data depicted by the first display device ([0048] imaging device of the computing device 102 may be provided via either a peripheral eye tracking camera or as an integrated a peripheral eye tracking camera in backlight system.), ,
a second video conferencing device remotely located from the first video conferencing device, coupled to the first video conferencing device for data exchange, and having a second display device for reproducing video image data captured by the first image capture device ([0033] exemplary inventive computer system of the present invention may transmit the acquired visual data to a remote server for processing, or in other implementations, may process the acquired visual data in the real-time on a computing device (e.g., mobile device, computer, etc.)., [0047] computer system environment incorporating certain embodiments of the present invention. As shown in FIG. 1, the inventive environment may include a user (101), who uses, for example, a first computing device (102) (e.g., mobile phone device, laptop, etc.), a server (103) and a second computing device (104) (e.g., mobile phone device, laptop, etc.)),
a processing unit which is coupled to the first image capture device and which is configured to receive and process the video image data captured by the first image capture device and to transmit the processed video image data to the second display device of the second video conferencing device ([0034] , examples of visual transformation effects may be, without limitation: [0035] turn user's face into a face of an animal, [0036] turn user's face into a face of another user, [0037] a race transformation, [0038] a gender transformation, [0039] an age transformation (making people look younger or older), [0040] choose an object which may be closest to the subject's appearance (based on the machine-learning algorithm's logic), [0041] swap parts of subject's head, [0042] make drawings on the subject's face or/and head, [0043] intentionally deform user's face, [0044] utilize dynamic masks (e.g., changing eye-gaze direction, eyebrow motion, opening mouth, etc.), and [0045] change user's appearance based on some predefined logic (e.g., changing user's appearance in a such way, as if they were poor, rich, live in the Stone Age, in future, astronauts, etc.).,
during the processing of the video image data, the modified video image data are generated by a Generative Adversarial Network (GAN) with a generator network and a discriminator network ([0033] utilizes the at least one conditional generative adversarial neural network which is programmed/configured to perform visual appearance transformations), and
the generator network generating modified video image data and the discriminator network evaluating a similarity between the depiction of the head of the first user in the modified video image data and the captured video image data and also evaluating a match between the direction of gaze of the first user in the modified video image data and the target direction of gaze ([0044] utilize dynamic masks (e.g., changing eye-gaze direction, eyebrow motion, opening mouth, etc.), [0068] presenting the at least one first neural network with at least one training real representation of the plurality of training real representations and one or more candidate meta-parameters of one or more latent variables of the at least one first visual effect to incorporate the at least one first visual effect into the at least one portion of the at one first real subject of the at least one training real representation to generate at least one first training photorealistic-imitating synthetic representation of the at least one portion of the at least one first real subject with the at least one first visual effect; ii) presenting the at least one second neural network with (1) the at least one first training photorealistic-imitating synthetic representation of the at least one first portion of the at least one first real subject with the at least one first visual effect and (2) the at least one training synthetic representation having the at least one visual effect applied to the at least one portion of the at least one synthetic subject to determine one or more actual meta-parameters of the one or more latent variables of the at least one first visual effect, where the one or more actual meta-parameters are meta-parameters at which the at least one second neural network has identified that the at least one first training photorealistic-imitating synthetic representation of the at least one portion of the at least one first real subject with the at least one first visual effect to be realistic).
.
Park discloses wherein the processing unit is configured to detect the direction of gaze of the depicted first user when processing the video image data and to modify in the video image data the reproduction at least of the region of the head of the first user comprising the eyes such that the target direction of gaze of the first user appears in the modified video image data as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device, ([0040] imaging system identifies that a gaze of the first user as represented in the image data is directed toward a displayed representation of at least a portion of (e.g., a face) of a second user. For instance, the imaging system can determine that the gaze of the first user is on an upper-right corner of a display screen of the first user’s device while the upper-right corner of the display screen of the first user’s device is displaying an image of a second user - and therefore that the gaze of the first user is on the second user. [0065] modified image data 242 appears straightened relative to the image data 232 (e.g., no longer skewed or tilted clockwise), and the face of the user 215 in the modified image data 242 appears to face toward the camera that captured the image data)
Prokopenya and Park are combinable because they are from the same field of invention.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify computer system of Prokopenya to includ wherein the processing unit is configured to detect the direction of gaze of the depicted first user when processing the video image data and to modify in the video image data the reproduction at least of the region of the head of the first user comprising the eyes such that the target direction of gaze of the first user appears in the modified video image data as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device as described by Park
The motivation for doing so would have been to direct gaze of the first user as represented in the image data toward a displayed representation of at least a portion (e.g., a face) of a second user (Park, [0004]).
Therefore, it would have been obvious to combine Prokopenya and Park to obtain the invention as specified in claim 18.
Regarding claim 19, Prokopenya discloses Video conferencing system ([0066] exemplary computer system of the present invention may be utilized for various applications which may include, but not limited to, gaming, mobile-device games, video chats, video conferences, live video streaming, video streaming),comprising
a first video conferencing device having a first display device and a first image capture device, the first image capture device being arranged to capture at least a region of the head of a first user, said region comprising the eyes, in a position in which the first user is looking at the video image data depicted by the first display device ([0048] imaging device of the computing device 102 may be provided via either a peripheral eye tracking camera or as an integrated a peripheral eye tracking camera in backlight system.),,
a second video conferencing device remotely located from the first video conferencing device, coupled to the first video conferencing device for data exchange, and having a second display device ([0032] As used herein, the term “user” shall have a meaning of at least one user.) for reproducing video image data captured by the first image capture device ([0033] exemplary inventive computer system of the present invention may transmit the acquired visual data to a remote server for processing, or in other implementations, may process the acquired visual data in the real-time on a computing device (e.g., mobile device, computer, etc.)., [0047] computer system environment incorporating certain embodiments of the present invention. As shown in FIG. 1, the inventive environment may include a user (101), who uses, for example, a first computing device (102) (e.g., mobile phone device, laptop, etc.), a server (103) and a second computing device (104) (e.g., mobile phone device, laptop, etc.)),
a processing unit which is coupled to the first image capture device and which is configured to receive and process the video image data captured by the first image capture device and to transmit the processed video image data to the second display device of the second video conferencing device ([0034] , examples of visual transformation effects may be, without limitation: [0035] turn user's face into a face of an animal, [0036] turn user's face into a face of another user, [0037] a race transformation, [0038] a gender transformation, [0039] an age transformation (making people look younger or older), [0040] choose an object which may be closest to the subject's appearance (based on the machine-learning algorithm's logic), [0041] swap parts of subject's head, [0042] make drawings on the subject's face or/and head, [0043] intentionally deform user's face, [0044] utilize dynamic masks (e.g., changing eye-gaze direction, eyebrow motion, opening mouth, etc.), and [0045] change user's appearance based on some predefined logic (e.g., changing user's appearance in a such way, as if they were poor, rich, live in the Stone Age, in future, astronauts, etc.).),
wherein the processing unit is configured to detect the direction of gaze of the depicted first user when processing the video image data and to modify in the video image data the reproduction at least of the region of the head of the first user comprising the eyes such that the target direction of gaze of the first user appears in the modified video image data as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device,
successive video frames are captured by the first image capture device and,
when the direction of gaze of the first user changes, some video frames are interpolated during the processing of the video image data in such a way that the change in direction of gaze reproduced by the modified video image data is slowed down.
Park discloses wherein the processing unit is configured to detect the direction of gaze of the depicted first user when processing the video image data and to modify in the video image data the reproduction at least of the region of the head of the first user comprising the eyes such that the target direction of gaze of the first user appears in the modified video image data as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device ([0040] imaging system identifies that a gaze of the first user as represented in the image data is directed toward a displayed representation of at least a portion of (e.g., a face) of a second user. For instance, the imaging system can determine that the gaze of the first user is on an upper-right corner of a display screen of the first user’s device while the upper-right corner of the display screen of the first user’s device is displaying an image of a second user - and therefore that the gaze of the first user is on the second user. [0065] modified image data 242 appears straightened relative to the image data 232 (e.g., no longer skewed or tilted clockwise), and the face of the user 215 in the modified image data 242 appears to face toward the camera that captured the image data)
successive video frames are captured by the first image capture device ([0114] image data include image data captured using the image capture and processing system 100, image data captured using image sensor(s) of the sensor(s) 230, image data captured using the first camera 330A, image data captured using the second camera 330B, image data captured using the third camera 330C, image data captured using the fourth camera 330D, image data captured using the first camera 430A, image data captured using the second camera 430B, image data captured using the third camera 430C, image data captured using the fourth camera 430D, an image used as input data for the input layer 610 of the NN 600, an image captured using the input device 845, another image described herein, another set of image data described herein, or a combination thereof.)
when the direction of gaze of the first user changes, some video frames are interpolated during the processing of the video image data in such a way that the change in direction of gaze reproduced by the modified video image data is slowed down ([0122] generating the modified image data as in operation 720 includes generating intermediate image data at least in part by modifying the image data to modify at least the portion of the first user in the image data to be visually directed toward a forward direction.)
Prokopenya and Park are combinable because they are from the same field of invention.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify computer system of Prokopenya to include wherein the processing unit is configured to detect the direction of gaze of the depicted first user when processing the video image data and to modify in the video image data the reproduction at least of the region of the head of the first user comprising the eyes such that the target direction of gaze of the first user appears in the modified video image data as if the first image capture device were arranged on a straight line passing through a first surrounding region of the eyes of the first user and through a second surrounding region of the eyes of the second user depicted on the first display device, successive video frames are captured by the first image capture device and, when the direction of gaze of the first user changes, some video frames are interpolated during the processing of the video image data in such a way that the change in direction of gaze reproduced by the modified video image data is slowed down as described by Park
The motivation for doing so would have been to direct gaze of the first user as represented in the image data toward a displayed representation of at least a portion (e.g., a face) of a second user (Park, [0004]).
Therefore, it would have been obvious to combine Prokopenya and Park to obtain the invention as specified in claim 19
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHIVANG I PATEL whose telephone number is (571)272-8964. The examiner can normally be reached on M-F 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alicia Harrington can be reached on (571) 272-2330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHIVANG I PATEL/Primary Examiner, Art Unit 2615