DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This Office Action is in response to the Applicants’ communication filed on November 18, 2024. In virtue of this communication, claims 1-20 are currently presented in the instant application.
Drawings
The drawings submitted on November 18, 2024. These drawings are reviewed and accepted by the examiner.
Information Disclosure Statement
The information Disclosure Statement (IDS) Forms PTO-1449, filed on December 10, 2024 in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosed therein was considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 8, 10-14 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over von Cramon (US 20170154463 A1) in view of IWASE et al. “Relightable Hands: Efficient Neural Relighting of Articulated Hand Models”, Proceeding of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pages 16663-16673.
Regarding claim 1. von Cramon discloses a computer-implemented method for generating an animation sequence, the method comprising:
receiving one or more three-dimensional (3D) input meshes, wherein each input mesh includes a representation of an object included in a 3D scene (von Cramon, see at least par. [0107], FIG.2 is a schematic diagram illustrating an exemplary real-world scene 200 to be recorded with an image-capture system using novel image-capture techniques. The example real-world scene 200 is a junction of two streets in a city bordered by man-made structures such as two and three story buildings. The various structures and features of the real-world scene 200 can be defined in a three-dimensional coordinate system or three-dimensional space having an origin 201, an abscissa or X-axis 202, an ordinate or Y-axis 204, and a Z-axis 203… [0111] As further illustrated by way of a relatively small insert near a lower left-most corner of a building that faces both streets, a material used on the front of the building (e.g., concrete, granite, brick, etc.), which may include large enough surface variation to be measured by a photogrammetry engine, is represented by a localized three-dimensional polygonal mesh 215);
receiving, for each of the one or more 3D input meshes, a virtual camera position associated with the 3D input mesh and one or more virtual lighting positions associated with the 3D input mesh (von Cramon, see at least par. [0152] As further illustrated in FIG. 5, model generator 530, which may be embodied in hardware, firmware, software, or in combinations of these, is arranged to use the image information 560 to generate a model 580 of the real-world scene imaged in the photographically captured and modified images derived therefrom. The model 580, which may be a textured model, includes a render mesh 581, virtual light information 583, virtual camera information 585, isolated-specular projection map(s) 587 and diffuse-projection map(s) 589. The render mesh 581 includes digital assets, the individual members or files of which are identified by the virtual camera information to apply or overlay one or more appropriate diffuse-projection maps 589, and when so desired, one or more appropriate isolated-specular projection maps 587 on the polygonal surfaces of the render mesh to generate a three-dimensional color representation of the modeled real-world location. In some embodiments, the render mesh 581 may include both polygonal surfaces and UV maps. As described, a UV map is a two-dimensional surface that can be folded to lie against the polygonal surfaces.);
von Cramon does not disclose generating, for each of the one or more 3D input meshes and via a trained machine learning model, one or more rendered frames associated with the 3D, wherein each of the one or more rendered frames includes a two-dimensional (2D) representation of the object as viewed from the virtual camera position and illuminated by one or more virtual lights located at the one or more virtual lighting positions; and generating an output animation sequence based on the one or more rendered frames. However,
IWASE discloses:
generating, for each of the one or more 3D input meshes and via a trained machine learning model, one or more rendered frames associated with the 3D input (IWASE, see page 16664, Hand Modeling), wherein each of the one or more rendered frames includes a two-dimensional (2D) representation of the object as viewed from the virtual camera position and illuminated by one or more virtual lights located at the one or more virtual lighting positions (IWASE, see pages 16664-16665, Hand Modeling and Image-space Human Relighting); and
generating an output animation sequence based on the one or more rendered frames (IWASE, see page 16665, Model-based Human Relighting).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the method and apparatus of von Cramon, with generating, for each of the one or more 3D input meshes and via a trained machine learning model, one or more rendered frames associated with the 3D input, wherein each of the one or more rendered frames includes a two-dimensional (2D) representation of the object as viewed from the virtual camera position and illuminated by one or more virtual lights located at the one or more virtual lighting positions; and generating an output animation sequence based on the one or more rendered frames, as provided by IWASE. The modification provides an improved system and method for rendering animated performances based on a multiple 2D representations of a scene, thereby to achieve generalization with physics-inspired illumination features such as visibility, diffuse shading, and specular reflections computed on a coarse proxy geometry, maintaining a small computational overhead. (IWASE, see abst).
Regarding claim 2. von Cramon in view of IWASE discloses the computer-implemented method of claim 1 (as rejected above), von Cramon in view of IWASE further discloses wherein the trained machine learning model includes a trained relightable Mixture of Volumetric Primitives (MVP) model (IWASE, see page 16666, first col., first paragraph).
Regarding claim 3. von Cramon in view of IWASE discloses the computer-implemented method of claim 2 (as rejected above), von Cramon in view of IWASE further discloses further comprising generating, via the trained relightable MVP model, a 3D MVP frame including a plurality of volumetric primitives, wherein each of the plurality of volumetric primitives includes position, orientation, size, color and opacity information associated with the volumetric primitive (IWASE, see page 16667 and first paragraph).
Regarding claim 4. Von Cramon in view of IWASE discloses the computer-implemented method of claim 1 (as rejected above), von Cramon in view of IWASE further discloses wherein the object includes a human actor exhibiting a facial expression (IWASE, see page 16665, 2nd col. And first par.).
Regarding claim 8. von Cramon in view IWASE discloses the computer-implemented method of claim 1 (as rejected above), von Cramon in view of IWASE further discloses wherein the rendered frame includes a 2D raster image including a plurality of pixels each including color and opacity values (von Cramon, see par. [0146] Image data can be arranged in any order using any desired number of bits to represent data values corresponding to the electrical signal produced at a corresponding location in the image sensor at a defined location in the raster of pixels. In computer graphics, pixels encoding the RGBA color space information, where the channel defined by the letter A corresponds to opacity, are stored in computer memory or in files on disk, in well-defined formats).
Regarding claim 10. von Cramon in view of IWASE discloses the computer-implemented method of claim 1 (as rejected above), von Cramon in view of IWASE further discloses wherein the trained machine learning model calculates one or more local lighting directions associated with one of the one or more virtual lighting positions (von Cramon, see at least par. [0166] When rendering a virtual environment from a model of a real-world scene, a rendering engine takes into account various settings for controls in a virtual camera which correspond to many of the features and controls present in real cameras. As a starting point, whereby a real camera is located somewhere in a three-dimensional coordinate space and is pointed in a direction, settings for translation and rotation in the virtual camera serve a similar end and can be animated over time to mirror camera movement in the real world. In the case of video games and VR, various inputs from a user allow one to navigate and observe from any angle a virtual environment at will using real-time rendering engines.).
Regarding claim 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of claim 1. Therefore, claim 11 is further rejected based on the same rationale as claim 1 set forth above and incorporated herein.
Regarding claim 12. The one or more non-transitory computer-readable media of claim 12 performs same step of claim 2. Therefore, claim 12 is further rejected based on the same rationale as claim 2 set forth above and incorporated herein.
Regarding claim 13. The one or more non-transitory computer-readable media of claim 13 performs same step of claim 3. Therefore, claim 13 is further rejected based on the same rationale as claim 3 set forth above and incorporated herein.
Regarding claim 14. The one or more non-transitory computer-readable media of claim 14 performs same step of claim 4. Therefore, claim 14 is further rejected based on the same rationale as claim 4 set forth above and incorporated herein.
Regarding claim 18. The one or more non-transitory computer-readable media of claim 16 performs same step claim 10. Therefore, claim 18 is further rejected based on the same rationale as claim 10 set forth above and incorporated herein.
Regarding claim 19. A system comprising: one or more memories storing instructions; and one or more processors (von Cramon, see FIG. 5 and par. [0158]) for executing the instructions to performs same rationale as claim 1. Therefore, claim 19 is further rejected based on the same rationale as claim 7 set forth above and incorporated herein.
Regarding claim 20. The system of claim 20 performs same step of claim 3. Therefore, claim 20 is further rejected based on the same rationale as claim 3 set forth above and incorporated herein.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over von Cramon (US 20170154463 A1) in view of IWASE et al. “RelightableHands: Efficient Neural Relighting of Articulated Hand Models”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pages 16663-16673, as applied claim 4 above, and further in view of MOSER et al. (US 20240161407 A1).
Regarding claim 5. Von Cramon in view of IWASE discloses the computer-implemented method of claim 4 (as rejected above), but von Cramon in view of IWASE does not discloses wherein the 3D input mesh is based on a blendshape model associated with the human actor. However,
MOSER discloses:
wherein the 3D input mesh is based on a blendshape model associated with the human actor (MOSER, see at least par. [0035] performing a blendshape decomposition of the approximate actor-specific ROM to yield a blendshape basis or a plurality of blendshapes; performing a blendshape optimization to obtain a blendshape-optimized 3D mesh, the blendshape optimization comprising determining, for each frame of the HMC-captured actor performance, a vector of blendshape weights and a plurality of transformation parameters which, when applied to the blendshape basis to reconstruct the 3D mesh topology, minimize a blendshape optimization loss function which attributes loss to differences between the reconstructed 3D mesh topology and the frame of the HMC-captured actor performance; performing a mesh-deformation refinement on the blendshape-optimized 3D mesh to obtain a mesh-deformation-optimized 3D mesh, the mesh-deformation refinement comprising determining, for each frame of the HMC-captured actor performance, 3D locations of a plurality of handle vertices which, when applied to the blendshape-optimized 3D mesh using a mesh-deformation technique, minimize a mesh-deformation refinement loss function which attributes loss to differences between the deformed 3D mesh topology and the HMC-captured actor performance; and generating the plurality of frames of facial animation based on the mesh-deformation-optimized 3D mesh.).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the method and apparatus of von Cramon, with wherein the 3D input mesh is based on a blendshape model associated with the human actor, as provided by MOSER. The modification provides an improved system and method for rendering animated performances based on a multiple 2D representations of a scene, thereby to imparting the facial characteristics of an actor onto such a computer representation (MOSER, see par. [0004]).
Claims 6 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over von Cramon (US 20170154463 A1) in view of IWASE et al. “RelightableHands: Efficient Neural Relighting of Articulated Hand Models”, Procdding of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pages 16663-16673, as applied claims 1 and 11 above, and further in view of Zhou et al. (US 20210012560 A1).
Regarding claim 6. Von Cramon in view of IWASE discloses the computer-implemented method of claim 1 (as rejected above), but Von Cramon in view of IWASE does not discloses further comprising blending two or more of the rendered frames, wherein each of the two or more rendered frames is associated with a single virtual lighting position. However,
Zhou discloses:
further comprising blending two or more of the rendered frames, wherein each of the two or more rendered frames is associated with a single virtual lighting position (Zhou, see at least par. [0052] In block 222, a relighted frame is rendered by adding a virtual light. A virtual light mimics the effect that would have been produced by a corresponding real light source if the real light source was present when the video is captured. Adding the virtual light includes computing adjustments to the color values of pixels in the frame based on the type of light and position of the light source).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the method and apparatus of von Cramon, with further comprising blending two or more of the rendered frames, wherein each of the two or more rendered frames is associated with a single virtual lighting position, as provided by Zhou. The modification provides an improved system and method for rendering animated performances based on a multiple 2D representations of a scene, thereby to improve the lighting conditions in the 3D scene and can also provide different lighting effects (Zhou, see par. [0053]).
Regarding claim 15. The one or more non-transitory computer-readable media of claim 15 performs same step of claim 6. Therefore, claim 15 is further rejected based on the same rationale as claim 6 set forth above and incorporated herein.
Claims 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over von Cramon (US 20170154463 A1) in view of IWASE et al. “RelightableHands: Efficient Neural Relighting of Articulated Hand Models”, Procdding of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pages 16663-16673, further in view of Zhou et al. (US 20210012560 A1), as applied claims 6 and 15 above, and further in view of Morin et al. (US 20090202114 A1).
Regarding claim 7. von Cramon in view IWASE and further in view of Zhou discloses the computer-implemented method of claim 6 (as rejected above), but von Cramon in view IWASE and further in view of Zhou does not disclose wherein blending the two or more rendered frames is based at least on light intensity values associated with the one or more virtual lights. However,
Morin discloses:
wherein blending the two or more rendered frames is based at least on light intensity values associated with the one or more virtual lights (Morin, see at least par. [0049] Facial Mapping With Lighting: Lighting intensity may be determined for particular areas of a user's face in a video feed, and objects that have been added to the face (e.g., animated hair or glasses/goggles) may be rendered after being subjected to a comparable level of virtual light.).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the method and apparatus of von Cramon, with wherein blending the two or more rendered frames is based at least on light intensity values associated with the one or more virtual lights, as provided by Morin. The modification provides an improved system and method for rendering animated performances based on a multiple 2D representations of a scene, thereby to providing live-action image or video capture, such as capture of player faces in real time for use in interactive video games (Morin, see par. [0002]).
Regarding claim 16. The one or more non-transitory computer-readable media of claim 16 performs same step of claim 7. Therefore, claim 16 is further rejected based on the same step of claim 7 set forth above and incorporated herein.
Claims 9 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over von Cramon (US 20170154463 A1) in view of IWASE et al. “RelightableHands: Efficient Neural Relighting of Articulated Hand Models”, Procdding of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pages 16663-16673, as applied claims 1 and 11 above, and further in view of KAWAKAMI et al. (US 20220165032 A1).
Regarding claim 9. von Cramon in view of IWASE discloses the computer-implemented method of claim 1 (as rejected above), but von Cramon in view of IWASE does not disclose wherein the trained machine learning model calculates one or more local view directions associated with the virtual camera position. However,
KAWAKAMI discloses:
wherein the trained machine learning model calculates one or more local view directions associated with the virtual camera position (KAWAMI, see at least par. [0055] Alternatively, the image processing unit 211 may specify the virtual space based on the real image and include the data related to the virtual space (virtual space data) in the content image data. The virtual space data may include a position of a virtual camera set corresponding to a position of the distributor terminal 21. The virtual space data may include information about a position of each object in an optical axis direction (in other words, a z-direction or a depth direction) of the virtual camera. For example, the virtual space data may include a distance (that is, depth) from the virtual camera to each object. When the imaging unit 207 is configured by using a depth camera, the image processing unit 211 may acquire a distance to each reality object in the real image measured by the depth camera. Alternatively, the image processing unit 211 may calculate a positional relationship between objects in the optical axis direction of the virtual camera by analyzing the real image by a method such as machine learning. Alternatively, the image processing unit 211 may acquire the position or depth set for each first virtual object.).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to modify the method and apparatus of von Cramon, with wherein blending the two or more rendered frames is based at least on light intensity values associated with the one or more virtual lights, as provided by Morin. The modification provides an improved system and method for rendering animated performances based on a multiple 2D representations of a scene, thereby to providing live-action image or video capture, such as capture of player faces in real time for use in interactive video games (Morin, see par. [0002]).
Regarding claim 17. The one or more non-transitory computer-readable media of claim 17 performs same step of claim 9. Therefore, claim 17 is further rejected based on the same rationale as claim 9 set forth above and incorporated herein.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KIM THANH THI TRAN whose telephone number is (571)270-1408. The examiner can normally be reached Monday-Friday 8:00am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ALICIA HARRINGTON can be reached at 5712722330. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KIM THANH T TRAN/Examiner, Art Unit 2615
/JAMES A THOMPSON/Primary Examiner, Art Unit 2615