Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-9 are rejected under 35 U.S.C. 103 as being unpatentable over Habib et al. (Pub. No.: US 2023/0252746) in view of Petajan et al. (Pub. No.: US 2022/0086395).
Regarding claim 1, Habib discloses a processor-implemented method, comprising: initiating, via one or more hardware processors (Fig. 8, Processor 802), a session for real-time live telepresence of a remote human presenter (Fig. 2, 202) in an environment of a human observer (Fig. 3 and paragraph [0027], “the viewer will be able to see facial expressions of the presenter as represented by the video avatar, which will display the presenter’s expressions in real time”), wherein an acquisition device (Fig. 2, camera 140 and position sensor 142) is located in the environment of the remote human presenter and the human observer comprises a visual rendering device (e.g. Fig. 1, viewer computer device 146); generating at an initial phase, via the one or more hardware processors, an initial digital avatar of the remote human presenter (Fig. 1, avatars 122 is a video avatar), using a 3-dimensional (3-D) human model (Fig. 2, the video avatar 206 is rendered in 3D space, see paragraph [0024]), through the acquisition device; transmitting at the initial phase, via the one or more hardware processors, the initial digital avatar of the remote human presenter human observer through a public cloud infrastructure (para [0017], “the application or a portion of the application may be executed on two or more client devices that are interconnected through a local server, or that are connected through a remote server or cloud computing system designed to work with the client applications. A 3D image of the object of interest is loaded into the application(s) and can be displayed on respective client devices as if users are sitting around the object, with each client device displaying the 3D object from that device’s perspective”), wherein the initial digital avatar of the remote human presenter is subsequently rendered in the visual rendering device of the human observer to obtain a rendered digital avatar of the remote human presenter presenter and the
It is noted that Habib does not specifically disclose “transmitting at the initial phase, via the one or more hardware processors, the initial digital avatar of the remote human presenter along with an audio to the visual rendering device of the human observer” and “using an encoding technique, to obtain an encoded motion information of the remote human presenter and an encoded environmental parameter information as a time-series data, wherein the encoding technique encodes and converts the temporally consistent 3-D human pose and shape motion information and the one or more environmental parameters of the environment of the remote human presenter into a data interchange format comprising one or more name-value pairs” and “decoding and feeding at the live-rendering phase in real-time”.
Petajan is cited to teach a processing system acquires a video image and voice data of a remote viewer of a live event content, and generating animation parameters relating to the video image. An avatar of the viewer is constructed based on the animation parameters (see abstract). Petajan further discloses “transmitting at the initial phase, via the one or more hardware processors, the initial digital avatar of the remote human presenter along with an audio to the visual rendering device of the human observer” (see Fig. 2A, Digital Audio 216). In addition, Petajan discloses ““using an encoding technique, to obtain an encoded motion information of the remote human presenter and an encoded environmental parameter information as a time-series data, wherein the encoding technique encodes and converts the temporally consistent 3-D human pose and shape motion information and the one or more environmental parameters of the environment of the remote human presenter into a data interchange format comprising one or more name-value pairs” (see Fig. 2C, Face animation parameters (FABs) 233 and FBA Encoder 234, also para [0037-0038]). Furthermore, Petajan discloses “decoding and feeding at the live-rendering phase in real-time” (see para [0042]). It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to have modified Habib with the feature of audio transmission along with video and further incorporating with an encoder and decoder as taught by Petajan so as to provide audio and video transmission signal through the network.
Regarding claim 2, Habib discloses the processor-implemented method of claim 1, wherein generating at the initial phase, the initial digital avatar of the remote human presenter using the 3-D human model through the acquisition device, comprises: capturing an image representation of the remote human presenter through the acquisition device located in the environment of the remote human presenter; estimating one or more normal maps from the image representation using the 3-D human model; converting the one or more normal maps into one or more partial surfaces, using the 3-D human model; and adding one or more missing geometries to the one or more partial surfaces using the 3-D human model, to generate the initial digital avatar of the remote human presenter, wherein the one or more missing geometries are associated with (i) a texture, (ii) a body shape, and (iii) one or more wearable garments (see abstract, “A presenter video image is captured. A 3D image of a 3D object is rendered on the client devices and a presenter avatar is rendered on at least the viewer client device. The presenter avatar includes at least a portion of the presenter video image. When a positional input is detected at the presenter client device, the system renders, on the viewer client device, an articulated virtual appurtenance associated with the positional input, the 3D image, and the presenter avatar. A virtual interaction between the articulated virtual appurtenance and the 3D image appear to a viewer as naturally positioned for the interaction with respect to the viewer”, the normal map is interpretated as portion of presenter being converted into a 3D avatar).
Regarding claim 3, Habit as modified by Petajan discloses the processor-implemented method of claim 1, wherein estimating at the live rendering phase in real-time, the temporally consistent 3-D human pose and shape motion information of the remote human presenter from the frame sequence obtained through the acquisition device (e.g. camera capturing the presenter in frame sequence), comprises: selecting a set of consecutive frames within a temporal window, from the frame sequence obtained through the acquisition device; extracting one or more body-aware deep features from each of the set of consecutive frames (e.g. Fig. 2D of Petajan illustrates Face Definition Parameter (FDP) feature points of a face model used to generate an animated virtual audience member); predicting one or more initial per-frame estimates comprising one or more body parameters of the remote human presenter and one or more device parameters of the acquisition device, from the associated one or more body-aware deep features (Petajan: [0042] “Feature points are shown for a full face 241, face profile 242, eyes 243-244, teeth 245, nose 246, tongue 247, and mouth 248); recovering one or more spatio-temporal features from the initial per-frame estimates, using one or more spatio-temporal feature aggregation techniques; and estimating the temporally consistent 3-D human pose and shape motion information of the remote human presenter, in real-time, from the one or more spatio-temporal features, using a motion estimation and refinement technique (Petajan: [0046, 0047])
Regarding claim 4, Habib further discloses A system, comprising: a memory storing instructions (Fig. 8, memory device 804); one or more input/output (1/O) interfaces (Fig. 8, I/O 806); one or more hardware processors coupled to the memory via the one or more 1/O interfaces (Fig. 8, Processor 806 is coupled to the memory device 804). Claim 4 is a system claim corresponding to method claim 1 above. Thus, Claim 4 is rejected for the same reason as claim 1 above.
Regarding claim 5, it is a system claim corresponding to the method claim 2. Thus, Claim 5 is rejected for the same reason as claim 2 above.
Regarding claim 6, it is a system claim corresponding to the method claim 3 above. Thus, Claim 6 is rejected for the same reason as claim 3 above.
Regarding 7, it is a CRM claim corresponding to the method claim 1 above. Thus, claim 7 is rejected for the same reason as claim 1 above.
Regarding claim 8, it is CRM claim corresponding to method claim 2 above. Thus, claim 8 is rejected for the same reason claim 2 above.
Regarding claim 9, it is a CRM claim corresponding to method claim 3 above. Thus, claim 9 is rejected for the same reason as claim 3 above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Menon et al. (Pub No.: US2024/0354996) is cited to teach a plurality of predicted images may be encoded by the autoencoder network to generate a plurality of encoded predicted images. The autoencoder network encodes a plurality of keypoint images to generate a plurality of encoded keypoint images.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIAO M WU whose telephone number is (571)272-7761. The examiner can normally be reached Monday to Friday 7:30am to 4pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexander Beck can be reached 571-272-3750. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613