DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner’s Note
The Examiner notes that Claim 20 recites “At least one computer-readable storage medium”. Specification, paragraph [0127], recites “As used herein, "computer-readable media" (also called "computer-readable storage media") refers to tangible storage media. Tangible storage media are non-transitory and have at least one physical, structural component”.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4 and 9-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sun et al. (US 20250069259 A1), and in view of Du et al. (US 20160300379 A1).
Regarding Claim 1. Sun discloses A video animation system (ABST reciting “Real-time extraction of human poses from video data for animation of avatars.” Fig. 9 showing a computing device) comprising:
at least one processor; (Fig. 6 showing Processor 902) and
at least one computer-readable storage medium having encoded thereon instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including: (¶9 reciting “a system includes at least one processor; and a memory coupled to the at least one processor, with software instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to perform operations.”)
generating a first frame of a video including a first avatar wherein the first avatar has a first gaze direction and/or a first pose;
generating a second frame of the video, wherein the first avatar has a second gaze direction and/or a second pose; and
outputting the first and second frames.
Sun teaches a method to animate avatar as shown in Fig. 8. A filtered pose sequence is generated from block 814; and ¶133 recites “In block 816, in some implementations, the filtered pose sequence from block 814 (including joint angles and translations) is used to animate an avatar. For example, the filtered pose sequence can include parameters 206 or 316 that describe 3D human poses, as described above. The animation of the avatar corresponds to the movement of the person depicted in the input video.” Animating an avatar using a pose sequence reads on generating a first frame of the avatar in a first pose and a second frame of the avatar in a second pose. In addition, ¶76 recites “For example, in a live video feed (e.g., a videoconference or other feed), an avatar can be animated in a display in coordination with captured video frames that are received by system 200 via the feed.”
However, Sun does not explicitly disclose wherein the second gaze direction and/or second pose is based on a plurality of environmental features of the first frame, the environmental features including one or more emotional environmental features of the first frame and/or one or more social environmental features of the first frame.
Du teaches “Apparatuses, methods and storage medium associated with creating an avatar video” (ABST). More specifically, ¶47 teaches rendering a subsequent frame based on emotional environment features of the previous frame, recites “Referring now to FIG. 5, wherein two example image frames of an example generated video, according to the disclosed embodiments, are shown. As described earlier, video generator 106 may be configured to capture the animation and rendering of animation-rendering engine 104 into a number of image frames. Further, the captured image frames may be combined/stitched together to form a video. Illustrated in FIG. 5 are two example image frames 502 and 504 of an example video 500. Example image frames 502 and 504 respectively capture animation of two avatars corresponding to two characters speaking their dialogues 506.” In Fig. 5, during the animation, a second frame of an avatar is rendered based on the emotion features of a first frame of the same avatar, such as “It’s nice to meet you!”.
It would have been obvious to one with ordinary skill, before the effective filing date of the claimed invention, to modify the system (taught by Sun) to render a second frame of avatar animation based on emotional features of a first frame (taught by Du). The suggestions/motivations would have been to make avatar animation easier (¶3), and to apply a known technique to a known device (method, or product) ready for improvement to yield predictable results.
Regarding Claim 2. Sun in view of Du discloses The system of claim 1, wherein the at least one processor includes at least one graphics processing unit (GPU). (Sun, ¶50 reciting “these machine learning models can be implemented on a GPU of a device providing the online metaverse platform.”)
Regarding Claim 3. Sun in view of Du discloses The system of claim 1, wherein the video includes a sequence of frames depicting a sequence of respective states of a scene, the scene including one or more characters, the sequence of frames including one or more respective avatars of the one or more characters. (Sun, ¶76 reciting “For example, in a live video feed (e.g., a videoconference or other feed), an avatar can be animated in a display in coordination with captured video frames that are received by system 200 via the feed.”)
Regarding Claim 4. Sun in view of Du discloses The system of claim 3, wherein the sequence of frames includes the first frame and the second frame, wherein the first frame depicts a first state of the sequence of states of the scene, and wherein the second frame depicts a second state of the sequence of states of the scene. (Sun, ¶133 reciting “In block 816, in some implementations, the filtered pose sequence from block 814 (including joint angles and translations) is used to animate an avatar. . . The animation of the avatar corresponds to the movement of the person depicted in the input video.”)
Regarding Claim 9. Sun in view of Du discloses The system of claim 4, wherein the plurality of environmental features of the first frame include one or more physical environmental features of the first state of the scene, and wherein the one or more physical environmental features of the first state of the scene include one or more attributes of the one or more characters in the first state of the scene, wherein the one or more attributes of a particular character of the one or more characters include a location of the particular character, a location of a facial landmark or body landmark of the particular character, a gaze direction of the particular character, and/or a pose of the particular character. (Sun discloses generating animation of an avatar. In the animation sequence of frames, an avatar in a subsequent frame (i.e. a second frame) has a pose based on the avatar pose in the previous frame (i.e. the first frame))
Regarding Claim 10. Sun in view of Du discloses The system of claim 3, wherein the one or more emotional environmental features of the first frame include one or more emotional states of the one or more characters, a mood of a conversation between two or more of the characters, and/or an emotional context associated with the scene. (Du, Fig. 5 showing emotional states of the one or more characters. The suggestions/motivations would have been the same as that of Claim 1 rejections.)
Regarding Claim 11. Sun in view of Du discloses The system of claim 3, wherein the one or more social features of the first frame include one or more social statuses of the one or more characters, one or more social or hierarchical relationships between or among the one or more characters, and/or a cultural context associated with the one or more characters. (Du, Fig. 5 showing social relationships between two characters. The suggestions/motivations would have been the same as that of Claim 1 rejections.)
Regarding Claim 12. Sun in view of Du discloses The system of claim 1, wherein the video depicts at least a portion of a video game, movie, show, videoconference, virtual reality (VR) application, augmented reality (AR) application, metaverse, or digital assistant. (Sun, ¶76 reciting “the parameters 206 can be used to animate an avatar real time coordination with the streaming input video. For example, in a live video feed (e.g., a videoconference or other feed), an avatar can be animated in a display in coordination with captured video frames that are received by system 200 via the feed.”)
Claim 13, has similar limitations as of Claim(s) 1, therefore it is rejected under the same rationale as Claim(s) 1.
Claim 20, has similar limitations as of Claim(s) 1, therefore it is rejected under the same rationale as Claim(s) 1.
Claim 14, has similar limitations as of Claim(s) 3 and 4, therefore it is rejected under the same rationale as Claim(s) 3 and 4.
Claim 19, has similar limitations as of Claim(s) 9, therefore it is rejected under the same rationale as Claim(s) 9.
Regarding Claim 15. Sun in view of Du discloses The video animation method of claim 14, wherein the plurality of environmental features of the first frame include at least two of a physical environmental feature of the first state of the scene, an emotional environmental feature of the first state of the scene, or a social environmental feature of the first state of the scene. (Du, Fig. 5 showing emotional states of the one or more characters and social relationships between two characters. The suggestions/motivations would have been the same as that of Claim 1 rejections.)
Regarding Claim 16. Sun in view of Du discloses The video animation method of claim 13, further comprising generating character data indicating the second gaze direction and/or the second pose of the first avatar in the second frame, the generating the character data being based on a plurality of environmental features of the first frame. (Sun teaches generating character data (filtered pose sequence) in Fig. 8, block 808-814, and ¶132. Du teaches a plurality of environmental features are used to generate a second frame. The suggestions/motivations of combination would have been the same as that of Claim 1 rejections.)
Regarding Claim 17. Sun in view of Du discloses The video animation method of claim 16, wherein the first avatar has, in the first frame, a first facial expression, wherein the character data further indicate a second facial expression of the character, and wherein the first avatar has, in the second frame, the second facial expression. (Du teaches an animation of the avatar’s facial expressions. The suggestions/motivations would have been the same as that of Claim 1 rejections.)
Regarding Claim 18. Sun in view of Du discloses The video animation method of claim 16, wherein the generating the character data is further based on the first gaze direction, the first pose of the first avatar in the first frame, one or more predicted gaze directions of the first avatar for one or more frames subsequent to the second frame, and/or one or more predicted poses of the first avatar for one or more frames subsequent to the second frame. (Sun, Fig. 8, blocks 808-814 teaching to generate pose sequence that reads on a character data.)
Allowable Subject Matter
Claims 5-8 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding Claim 5. Sun in view of Du discloses The system of claim 4, wherein the operations further include generating the video, and wherein generating the video includes: the generating the first frame; and the generating the second frame.
However, the closest art fails to teach each and every limitations of the claim and its base claim(s).
Claims 6-8 depend from claim 5, and therefore also contain allowable subject matter.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YI WANG whose telephone number is (571)272-6022. The examiner can normally be reached 9am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at (571)272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YI WANG/ Primary Examiner, Art Unit 2619