Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: an image generator, a measuring unit, an evaluation unit, an actuation unit in claims 1-3.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 11 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 11 recites “selecting a synthetic voice or synthetic visual appearance or synthetic environment from selection of a first synthetic voice and at least one second synthetic voice for further operation of a voice cloning device”. The scope of this limitation is unclear because the claim appears to require selection of a synthetic visual appearance or a synthetic environment from a group consisting of a first synthetic voice and at least one second synthetic voice. A synthetic voice is an audio-based representation, whereas the synthetic visual appearance and the synthetic environment are visual or spatial representation. Therefore, it is unclear how the synthetic visual appearance or synthetic environment can be selected from a set of synthetic voices, also the synthetic content is selected from different categories. In addition, the claim does not clearly define the relationship between the first synthetic voice, the at least one second synthetic voice, and the synthetic visual appearance or synthetic environment. For examination purpose, the limitation is interpreted as depending on the evaluating of the at least one physiological parameter for the at least one second synthetic visual appearance or the one second synthetic environment, selecting the second synthetic visual appearance or synthetic environment for further operation.
Claim Objections
Claim 3 is objected to because of the following informalities:
Claim 3 recites actuate the voice cloning device; it would be changed to actuate a voice cloning device.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2 and 4-12 are rejected under 35 U.S.C. 103 as being unpatentable over Bitouk et al. (US 2013/0336600) in view of Kacelenga (US 2020/0089321) in view of Dariush et al. (US 2011/0054870).
Regarding claim 1, Bitouk et al. (hereinafter Bitouk) discloses a system (Bitouk, [0002], “The disclosed subject matter relates to methods, systems”), the system comprising:
a deepfake device (Bitouk, [0041], “processor 1302 can be a general purpose device such as a computer or a special purpose device such as a client, a server, an image capture device (such as a camera, video recorder, scanner, mobile telephone, personal data assistant, etc.)”) comprising:
an image generator (Bitouk, [0041], “As shown, hardware 1300 includes an image processor 1302”) that is configured to replace a natural visual appearance of a person or a natural environment with a synthetic visual appearance or synthetic environment that is different than the natural visual appearance of the person or the natural environment (Bitouk, [0034], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image (using any suitable mechanism, such as an affine transformation) and output as necessary at 118”), wherein at least two synthetic visual appearances or two synthetic environments are selectable (Bitouk, [0034], “At 116, the face-swapped copies are then ranked to determine which copy is best. Any suitable mechanism for ranking the face-swapped copies can be used in some embodiments”); and
a visualization device for visualizing the synthetic visual appearance of the person or the surrounding synthetic environment (Bitouk, [0034], “the best copy can be output to a display device such as a computer monitor, camera display, mobile phone display, personal data assistant display”);
an actuation unit configured to actuate the deepfake device such that a synthetic visual appearance or synthetic background is selected or rejected depending on selection (Bitouk, [0034], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image (using any suitable mechanism, such as an affine transformation) and output as necessary at 118”. In addition, in paragraph [0041], “hardware 1300 includes an image processor 1302”);
Bitouk does not expressly disclose “monitoring system”;
Kacelenga discloses monitoring system for monitoring a user (Kacelenga, [0031], “an xR application may monitor parameters describing the environment 100 in which the xR session is being conducted…102A-B that monitor various physiological characteristics of the respective users 101A-B.”);
a measuring unit configured to record at least one physiological parameter of the user (Kacelenga, [0042], “HMD 102A may utilize a physiological tracking system 220 that includes a variety of sensors that support monitoring of physiological characteristics of the user wearing the HMD”);
an evaluation unit that is configured to evaluate the at least one measured physiological parameter of the user (Kacelenga, [0046], “the physiological sensor data, including audio, that is captured by HMD 102A may be streamed via the tether to the host IHS 103A, where this physiological state data may be used to evaluate and classify the user's physiological state”. In addition, in paragraph [0049], “IHS 103A includes one or more processors 301, such as a Central Processing Unit (CPU), to execute code retrieved from a system memory 305”);
actuate a device depending on a result of the evaluation by the evaluation unit (Kacelenga, [0073], “such as micro fans and warmers, that may be operated in response to physiological state determinations made at step 445. For instance, while the user is determined to be below a threshold stress level classification, warmers may be activated in order to enhance the user's xR experience, such as to convey suspense and excitement”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the face swapping system of Bitouk to incorporate the physiological detection and evaluation techniques of Kacelenga to monitor a user’s response to the displayed face swapped visual appearance. The motivation for doing so would have been improving operation through automated sensor based feedback.
Bitouk as modified by Kacelenga does not expressly disclose “monitoring a patient”;
Dariush et al. (hereinafter Dariush) discloses monitoring a patient (Dariush, [0078], “A subject (patient or game player) is requested to take a sequence of postures (by remembering the posture sequence). The computer software can identify which postures were taken and which postures were skipped (forgotten), how correct the sequence (order of postures) was, thus being able to rate the subject ability of re-creating a given posture sequence. This type of operation is useful in games and rehabilitation”).
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the face swapping system of Bitouk to include patient monitoring as taught by Dariush. The motivation for doing so would have been providing ability to observe the patient’s response and interaction with the generated appearance.
Regarding claim 2, Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses compare the at least one measured physiological parameter of the user with at least one threshold value (Kacelenga, [0010], “a warmer that is activated if the physiological stress level of the individual is below a first threshold”).
Bitouk as modified by Kacelenga and Dariush discloses the patient (Dariush, [0078], “A subject (patient or game player) is requested to take a sequence of postures (by remembering the posture sequence). The computer software can identify which postures were taken and which postures were skipped (forgotten), how correct the sequence (order of postures) was, thus being able to rate the subject ability of re-creating a given posture sequence. This type of operation is useful in games and rehabilitation”).
Regarding claim 4, Bitouk discloses a system comprising (Bitouk, [0002], “The disclosed subject matter relates to methods, systems”):
the deepfake device (Bitouk, [0041], “processor 1302 can be a general purpose device such as a computer or a special purpose device such as a client, a server, an image capture device (such as a camera, video recorder, scanner, mobile telephone, personal data assistant, etc.)”), wherein a first synthetic visual appearance or a first synthetic environment is set (Bitouk, [0034], “the best copy can be output to a display device such as a computer monitor, camera display, mobile phone display, personal data assistant display”);
the deepfake device using the first synthetic visual appearance or the first synthetic environment (Bitouk, [0034], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image (using any suitable mechanism, such as an affine transformation) and output as necessary at 118”);
Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses recording the at least one physiological parameter of the user during operation (Kacelenga, [0046], “the physiological sensor data, including audio, that is captured by HMD 102A may be streamed via the tether to the host IHS 103A, where this physiological state data may be used to evaluate and classify the user's physiological state”);
activating the device (Kacelenga, [0073], “warmers may be activated in order to enhance the user's xR experience”);
evaluating the at least one physiological parameter of the user (Kacelenga, [0046], “the physiological sensor data, including audio, that is captured by HMD 102A may be streamed via the tether to the host IHS 103A, where this physiological state data may be used to evaluate and classify the user's physiological state”. In addition, in paragraph [0049], “IHS 103A includes one or more processors 301, such as a Central Processing Unit (CPU), to execute code retrieved from a system memory 305”);
automatically actuating device depending on a result of the evaluating (Kacelenga, [0073], “such as micro fans and warmers, that may be operated in response to physiological state determinations made at step 445. For instance, while the user is determined to be below a threshold stress level classification, warmers may be activated in order to enhance the user's xR experience, such as to convey suspense and excitement”);
Bitouk as modified by Kacelenga and Dariush discloses the patient (Dariush, [0078], “A subject (patient or game player) is requested to take a sequence of postures (by remembering the posture sequence). The computer software can identify which postures were taken and which postures were skipped (forgotten), how correct the sequence (order of postures) was, thus being able to rate the subject ability of re-creating a given posture sequence. This type of operation is useful in games and rehabilitation”).
The remaining limitations recite in claim 4 are similar in scope to the functions recited in claim 1 and therefore are rejected under the same rationale.
Regarding claim 5, Bitouk discloses the first synthetic visual appearance or the first synthetic environment is selected or rejected for further operation of the deepfake device (Bitouk, [0034-0035], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image…a library of candidate faces can be used for face swapping”);
Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses depending on the result of the evaluating by the evaluation unit (Kacelenga, [0046], “the physiological sensor data, including audio, that is captured by HMD 102A may be streamed via the tether to the host IHS 103A, where this physiological state data may be used to evaluate and classify the user's physiological state”).
Regarding claim 6, Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses comparing the at least one physiological parameter with at least one threshold value (Kacelenga, [0010], “a warmer that is activated if the physiological stress level of the individual is below a first threshold”).
The method of claim 4, wherein evaluating the at least one physiological parameter of the patient comprises comparing the at least one physiological parameter with at least one threshold value.
Regarding claim 7, Bitouk discloses the deepfake device takes place such that the first synthetic visual appearance or the first synthetic environment is selected (Bitouk, [0034-0035], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image (using any suitable mechanism, such as an affine transformation) and output as necessary at 118…a library of candidate faces can be used for face swapping”);
that the first synthetic visual appearance or the first synthetic environment is rejected (Bitouk, [0034-0035], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image…a library of candidate faces can be used for face swapping”. The unselected faces are considered rejected);
Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses actuating the device (Kacelenga, [0073], “warmers may be activated in order to enhance the user's xR experience”);
when the evaluating of the at least one physiological parameter indicates that a threshold value for the at least one physiological parameter is not exceeded (Kacelenga, [0010], “a warmer that is activated if the physiological stress level of the individual is below a first threshold”);
when the evaluating of the at least one physiological parameter indicates that a threshold value for the at least one physiological parameter is exceeded (Kacelenga, [0080], “in which additional types of lower-confidence detections are sufficient to make a physiological state classification, such as a determination that the user's stress level is above a threshold”).
Regarding claim 8, Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses evaluating the at least one physiological parameter with respect to a state of mind of the patient (Kacelenga, [0063], “The physiological sensor data is streamed to the sensor hub of the host IHS and processed in order to determine the user's physiological attributes, such as heart rate”).
The method of claim 4, wherein evaluating the at least one physiological parameter comprises evaluating the at least one physiological parameter with respect to a state of mind of the patient.
Regarding claim 9, the actuating takes place that the first synthetic visual appearance or the first synthetic environment is selected (Bitouk, [0034-0035], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image…a library of candidate faces can be used for face swapping”);
the first synthetic visual appearance or the first synthetic environment is rejected (Bitouk, [0034-0035], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image…a library of candidate faces can be used for face swapping”. The unselected faces are considered rejected);
Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses the actuating takes place (Kacelenga, [0073], “warmers may be activated in order to enhance the user's xR experience”);
produces a positive state of mind in the user (Kacelenga, [0066], “At step 415d, the heartrate variability may be used to classify the physiological state of the user, such as the user's level of excitement and/or stress”. Excitement is considered positive state of mind);
produces a negative state of mind (Kacelenga, [0066], “At step 415d, the heartrate variability may be used to classify the physiological state of the user, such as the user's level of excitement and/or stress”. Stress is considered negative state of mind);
Bitouk as modified by Kacelenga and Dariush discloses the patient (Dariush, [0078], “A subject (patient or game player) is requested to take a sequence of postures (by remembering the posture sequence). The computer software can identify which postures were taken and which postures were skipped (forgotten), how correct the sequence (order of postures) was, thus being able to rate the subject ability of re-creating a given posture sequence. This type of operation is useful in games and rehabilitation”).
Regarding claim 10, Bitouk discloses the best is selected (Bitouk, [0034-0035], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image…a library of candidate faces can be used for face swapping”);
Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses produce positive state of mind of the user (Kacelenga, [0081], “in an xR environment that includes a rendering of a hot desert environment, warmers in the HMD may be activated as long as the user's excitement or stress level does not rise above a specified threshold”);
Bitouk as modified by Kacelenga and Dariush discloses the patient (Dariush, [0078], “A subject (patient or game player) is requested to take a sequence of postures (by remembering the posture sequence). The computer software can identify which postures were taken and which postures were skipped (forgotten), how correct the sequence (order of postures) was, thus being able to rate the subject ability of re-creating a given posture sequence. This type of operation is useful in games and rehabilitation”).
Regarding claim 11, Bitouk discloses at least one second synthetic visual appearance or one second synthetic environment (Bitouk, [0034-0035], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image…a library of candidate faces can be used for face swapping”. Repeat the process a second time subsequent to the first time, results in a second synthetic visual appearance);
selecting the second synthetic visual appearance (Bitouk, [0034], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image (using any suitable mechanism, such as an affine transformation) and output as necessary at 118”);
Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses measuring and evaluating at least one physiological parameter (Kacelenga, [0046], “the physiological sensor data, including audio, that is captured by HMD 102A may be streamed via the tether to the host IHS 103A, where this physiological state data may be used to evaluate and classify the user's physiological state”);
depending on the evaluating of the at least one physiological parameter for further operation of the device (Kacelenga, [0073], “such as micro fans and warmers, that may be operated in response to physiological state determinations made at step 445. For instance, while the user is determined to be below a threshold stress level classification, warmers may be activated in order to enhance the user's xR experience, such as to convey suspense and excitement”).
Regarding claim 12, at least one second synthetic visual appearance or one second synthetic environment in each case (Bitouk, [0034-0035], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image…a library of candidate faces can be used for face swapping”. Repeat the process a second time subsequent to the first time, results in a second synthetic visual appearance)
selecting a synthetic visual appearance or synthetic environment from the selection of the first synthetic visual appearance and the at least second synthetic visual appearance or the first synthetic environment and the at least second synthetic environment for further operation of the deepfake device (Bitouk, [0034], “The most-similar face-swapped copy can then be considered the best. This best copy can then be re-aligned to the alignment of the original input image (using any suitable mechanism, such as an affine transformation) and output as necessary at 118”);
Bitouk as modified by Kacelenga with the same motivation from claim 1 discloses measuring and evaluating at least one physiological parameter (Kacelenga, [0046], “the physiological sensor data, including audio, that is captured by HMD 102A may be streamed via the tether to the host IHS 103A, where this physiological state data may be used to evaluate and classify the user's physiological state”);
depending on the evaluating of the at least one physiological parameter (Kacelenga, [0073], “such as micro fans and warmers, that may be operated in response to physiological state determinations made at step 445. For instance, while the user is determined to be below a threshold stress level classification, warmers may be activated in order to enhance the user's xR experience, such as to convey suspense and excitement”).
Allowable Subject Matter
Claim 3 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KYLE ZHAI whose telephone number is (571)270-3740. The examiner can normally be reached 9AM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ke Xiao can be reached at (571) 272 - 7776. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KYLE ZHAI/Primary Examiner, Art Unit 2611