DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Preliminary Amendment
2. Preliminary amendment to claims filed on 12/19/2024 has been accepted and is examined below.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
3. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: sending module and receiving module in claims 9-10 and 12-19.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
4. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
5. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
6. Claims 1 and 9-12 are rejected under 35 U.S.C. 103 as being unpatentable over Albuz et al. (US Patent No. 10,755,463 B1).
7. Regarding Claim 1, Albuz discloses A method for data processing, (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The method includes processing audio signal.) comprising:
acquiring audio information, (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The audio signal corresponds to acquiring audio information.) motion feature information of a human face, (col. 14, lines 36-42 reciting “In addition, the method may further comprise, prior to processing the audio signal, accessing, from a database, animation data associated with a plurality of lip movements and facial-component movements. In particular embodiments, the animation data may be created by training animation models based on collected audio and corresponding videotaped facial-component movements.” The database of animation data (lip and facial-component movements) correspond to motion feature information of a human face.) and a target human face image; (col. 3, lines 4-10 reciting “These animation models may be used for illustrating a generic human face showing movement of facial features corresponding to audio to provide a representation of the face modeling lip movement representative of the audio to look as if the animated face was speaking the words and/or sounds of the audio (e.g., a representative animation of a person talking in virtual reality).” Generic human face corresponds to a target human face image.)
adjusting the motion feature information according to the audio information to acquire target motion feature information of the human face; and (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The database of animation is adjusted (processed) so that the lip and facial component animations that corresponds to the speech units are selected and applied.)
Albuz obviously discloses generating a target human face image sequence corresponding to the audio information according to the target motion feature information and the target human face image. (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The animated face can correspond to a target human face image sequence. This is obvious because Albuz further discloses that the face can be a generic human face and animating a generic human face makes the interaction more intuitive and personable for the end-user. The animated face (or animated generic human face) uses the lip and facial component animations determined from processed speech units of the audio. Therefore, the generic human face is animated based on both the audio information and selected sequence of lip and facial component animations corresponding to the audio information.)
00. Regarding Claim 9, Albuz discloses A video conference system, (see FIG. 7 wherein 700 displays a video conference system.) comprising a sending module, and a receiving module, wherein:
the sending module is configured to send audio information to the receiving module; (col. 6, lines 34-38 reciting “The data collected by the microphone may be processed by client system 130, or sent to network 110 for processing (e.g., may be processed by network 110, social-networking system 160, and/or third-party system 170),” Microphone corresponds to sending module that captures and sends the audio data.) and the receiving module is configured to perform a method for data processing, (col. 6, lines 34-38 reciting “The data collected by the microphone may be processed by client system 130, or sent to network 110 for processing (e.g., may be processed by network 110, social-networking system 160, and/or third-party system 170),” the client system is a receiving module that obtains the audio signal and processes the audio signal.) comprising:
acquiring the audio information, (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The audio signal corresponds to acquiring audio information.) motion feature information of a human face, (col. 14, lines 36-42 reciting “In addition, the method may further comprise, prior to processing the audio signal, accessing, from a database, animation data associated with a plurality of lip movements and facial-component movements. In particular embodiments, the animation data may be created by training animation models based on collected audio and corresponding videotaped facial-component movements.” The database of animation data (lip and facial-component movements) correspond to motion feature information of a human face.) and a target human face image; (col. 3, lines 4-10 reciting “These animation models may be used for illustrating a generic human face showing movement of facial features corresponding to audio to provide a representation of the face modeling lip movement representative of the audio to look as if the animated face was speaking the words and/or sounds of the audio (e.g., a representative animation of a person talking in virtual reality).” Generic human face corresponds to a target human face image.)
adjusting the motion feature information according to the audio information to acquire target motion feature information of the human face; (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The database of animation is adjusted (processed) so that the lip and facial component animations that corresponds to the speech units are selected and applied.)
8. Regarding Claim 10, Albuz further discloses The video conference system according to claim 9, further comprising a display module, wherein:
the display module is connected with the receiving module, and the display module is configured to display the target human face image sequence that is generated by the receiving module. (col. 4, lines 52-58 reciting “The artificial reality system that provides the artificial reality content may be implemented on various platforms, including a head-mounted display (HMD) connected to a host computer system, a standalone HMD, a mobile device or computing system, or any other hardware platform capable of providing artificial reality content to one or more viewers.” HMD or mobile devices have display screen and displays the processed generic human face with lip/facial component animations for the viewer to view.)
Albuz obviously discloses and generating a target human face image sequence corresponding to the audio information according to the target motion feature information and the target human face image. (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The animated face can correspond to a target human face image sequence. This is obvious because Albuz further discloses that the face can be a generic human face and animating a generic human face makes the interaction more intuitive and personable for the end-user. The animated face (or animated generic human face) uses the lip and facial component animations determined from processed speech units of the audio. Therefore, the generic human face is animated based on both the audio information and selected sequence of lip and facial component animations corresponding to the audio information.)
9. Regarding Claim 11, Albuz discloses A device for data processing, comprising a memory, a processor and a computer program stored in the memory and executable by the processor which, when executed by the processor causes the processor to carry out a method for data processing, comprising: (col. 0, lines 65-67 reciting “In particular embodiments, memory 1004 includes main memory for storing instructions for processor 1002 to execute or data for processor 1002 to operate on.”)
acquiring audio information, (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The audio signal corresponds to acquiring audio information.) motion feature information of a human face, (col. 14, lines 36-42 reciting “In addition, the method may further comprise, prior to processing the audio signal, accessing, from a database, animation data associated with a plurality of lip movements and facial-component movements. In particular embodiments, the animation data may be created by training animation models based on collected audio and corresponding videotaped facial-component movements.” The database of animation data (lip and facial-component movements) correspond to motion feature information of a human face.) and a target human face image; (col. 3, lines 4-10 reciting “These animation models may be used for illustrating a generic human face showing movement of facial features corresponding to audio to provide a representation of the face modeling lip movement representative of the audio to look as if the animated face was speaking the words and/or sounds of the audio (e.g., a representative animation of a person talking in virtual reality).” Generic human face corresponds to a target human face image.)
adjusting the motion feature information according to the audio information to acquire target motion feature information of the human face; (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The database of animation is adjusted (processed) so that the lip and facial component animations that corresponds to the speech units are selected and applied.)
Albuz obviously discloses and generating a target human face image sequence corresponding to the audio information according to the target motion feature information and the target human face image. (Abstract reciting “In one embodiment, a method includes receiving an audio signal comprising a plurality of speech units, processing the audio signal to associate each of the speech units with a corresponding lip animation, determining pitch information associated with each of the plurality of speech units, processing the pitch information of each of the plurality of speech units to associate at least one of the speech units with a facial-component animation, and presenting the audio signal with a displayed animation of a face, wherein the animation of the face displays the lip animation associated with each of the speech units and the facial-component animation associated with the at least one speech unit. The animation of the face may be displayed in real time with the audio signal. The facial component animation may include animation of the lips, eyebrows, eyelids, and other portion of the upper face.” The animated face can correspond to a target human face image sequence. This is obvious because Albuz further discloses that the face can be a generic human face and animating a generic human face makes the interaction more intuitive and personable for the end-user. The animated face (or animated generic human face) uses the lip and facial component animations determined from processed speech units of the audio. Therefore, the generic human face is animated based on both the audio information and selected sequence of lip and facial component animations corresponding to the audio information.)
10. Regarding Claim 12, Albuz futther discloses A non-transitory computer-readable storage medium storing a computer-executable instruction which, when executed by a processor, causes the processor to carry out the method as claimed in claim 1. (col. 25, line 64 to col. 26, line 11 reciting “Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.”)
Allowable Subject Matter
11. Claims 2-8 and 13-20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
12. The following is a statement of reasons for the indication of allowable subject matter: Claim 2 recites the limitation acquiring a human face information extraction model; selecting a target motion sequence from at least one preset motion sequence; inputting the target motion sequence into the human face information extraction model to acquire key point information of the human face; and taking the key point information of the human face as the motion feature information of the human face which is neither disclosed nor suggested by the cited references, either singly or combination.
13. Claims 3-7 depend from claim 2.
14. Claim 7 recites the limitation acquiring a motion feature adjustment model comprising a first encoding module, a second encoding module, and a decoding module; inputting the audio information into the first encoding module to acquire first feature information; inputting the motion feature information into the second encoding module to acquire second feature information; fusing the first feature information and the second feature information to acquire fused feature information; and inputting the fused feature information into the decoding module to acquire the target motion feature information which is neither disclosed nor suggested by the cited references, either singly or combination.
15. Claim 8 depends from claim 7.
16. Claim 13 recites the limitation acquiring a human face information extraction model; selecting a target motion sequence from at least one preset motion sequence; inputting the target motion sequence into the human face information extraction model to acquire key point information of the human face; and
taking the key point information of the human face as the motion feature information of the human face which is neither disclosed nor suggested by the cited references, either singly or combination.
17. Claims 14-17 depends from claim 13.
18. Claim 18 recites the limitation acquiring a motion feature adjustment model comprising a first encoding module, a second encoding module, and a decoding module; inputting the audio information into the first encoding module to acquire first feature information; inputting the motion feature information into the second encoding module to acquire second feature information; fusing the first feature information and the second feature information to acquire fused feature information; and inputting the fused feature information into the decoding module to acquire the target motion feature information which is neither disclosed nor suggested by the cited references, either singly or combination.
19. Claim 19 depends from claim 18.
20. Claim 20 recites the limitation acquiring a human face information extraction model; selecting a target motion sequence from at least one preset motion sequence; inputting the target motion sequence into the human face information extraction model to acquire key point information of the human face; and
taking the key point information of the human face as the motion feature information of the human face which is neither disclosed nor suggested by the cited references, either singly or combination.
CONTACT
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FRANK S CHEN whose telephone number is (571)270-7993. The examiner can normally be reached Mon - Fri 8-11:30 and 1:30-6.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at 5712727794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/FRANK S CHEN/Primary Examiner, Art Unit 2611