DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3, 12, 13, and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering”, 12/07/2023, arXiv, 2310.08528v2 in view of Lv et al. (U.S. Pub. No. 20220239844) and Kopf et al. (U.S. Doc. No. 10038894).
Regarding claim 1, Wu discloses a four-dimensional scene reconstruction method (sec 2.1, “Our method aims at constructing a highly efficient training and rendering pipeline in Fig. 2 (c), while maintaining the quality, even for sparse inputs.”; also, sec 6, “This paper proposes 4D Gaussian splatting to achieve real time dynamic scene rendering”; also, sec 6, “This paper proposes 4D Gaussian splatting to achieve real time dynamic scene rendering”; also, method of constructing a 4D gaussian scene model), comprising: video frame at an initial moment in the multi-view video (sec 5.1, “We use the points computed by SfM[32] from the first frame of each video in Neu3D’s dataset and 200 frames randomly selected in HyperNeRF’s.”; also, initial moment is also referred to as the first frame of a video); deformable network (sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian deformation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, determining deformable network is performed by the Gaussian deformation field network); three-dimensional scene model (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds”; also, sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian de formation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, sec 4.1, “As shown in Fig. 3, given a view matrix M = [R,T], times tamp t, our 4D Gaussian splatting framework includes 3D Gaussians G and Gaussian deformation field network F.”; also, three-dimensional scene model was discuses since when 4D Gaussian splatting framework request both the deformation field network and the 3D Gaussians G is represented as the scene), four-dimensional scene model (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds”; also, sec 1, “An efficient 4D Gaussian splatting framework with an efficient Gaussian deformation field is proposed by modeling both Gaussian motion and Gaussian shape changes across time.”; also, sec 4.1, “Our 4D Gaussian splatting converts the original 3D Gaussians G into another group of 3D Gaussians G′ given a timestamp t, maintaining the effectiveness of the differential splatting as referred in [46]”; also, 4D Gaussian splatting is a holistic representation for dynamic scenes) three- dimensional scene model and the deformable network (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds”; also, sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian de formation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, sec 4.1, “Then a novel-view image ˆIis rendered by differential splat ting [46] S following ˆI= S(M,G′), where G′ = ∆G +G”; also, sec 4.2, “Finally, we obtain the deformed 3D Gaussians G′ = {X′,s′,r′,σ,C}.”; also, 3D Gaussian G is referred to as a scene. 4D scene was created with both the 3D gaussians G and the Gaussian deformation field network). Wu does not disclose generating a three-dimensional scene model corresponding to the multi-view images; corresponding to the multi-view video, the multi-view video, and camera pose information corresponding to the multi-view video.
However, in a similar field of endeavor, Lv discloses obtaining a multi-view video, wherein the multi-view video comprises multi-view images (para 29, “More formally, embodiments described herein may be used to tackle the problem of reconstructing dynamic 3D scenes from video inputs from multiple cameras, i.e., {C.sup.(t)} for time index t∈T={1, 2, . . . , T} with known camera intrinsic and extrinsic parameters.”; also, para 29, “The dynamic scene may be simultaneously recorded by several cameras (e.g., 5, 10, or 20 cameras) over a period of time”; also, para 58, “We first train a NeRF model on the keyframes, which we sample equidistant from the multi-view image sequence at fixed intervals K.”), corresponding to the multi-view video (para 62, “We have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation”; also, para 48, “The number of iterations it takes in training per epoch scales linearly with the number of the pixels in the multi-view videos.”), the multi-view video, and camera pose information corresponding to the multi-view video (para 29, “More formally, embodiments described herein may be used to tackle the problem of reconstructing dynamic 3D scenes from video inputs from multiple cameras, i.e., {C.sup.(t)} for time index t∈T={1, 2, . . . , T} with known camera intrinsic and extrinsic parameters.”; also, para 42, “More formally, given a ray r(s)=o+sd (origin o and direction d defined by the specified camera pose and camera intrinsics), the rendered color of the pixel corresponding to this ray C(r) is an integral over the radiance weighted by accumulated opacity:”; also, para 41, “Conceptually, NeRF learns and encodes the radiance and opacity values of a dynamic scene over time based on video frames captured by multiple cameras.”), corresponding to the multi-view video (para 62, “We have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation.”; also, para 41, “Conceptually, NeRF learns and encodes the radiance and opacity values of a dynamic scene over time based on video frames captured by multiple cameras.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of a four-dimensional scene reconstruction method in which a Gaussian deformation field network is determined from an explicit three-dimensional scene representation of point clouds and a timestamp, and in which the four-dimensional model is that representation combined with the deformation the network predicts, with the features of Lv's invention of a dynamic scene simultaneously recorded by several cameras whose video inputs are obtained together with their known camera intrinsic and extrinsic parameters and whose recordings form a multi-view image sequence the representation is made to correspond to. The combination would have been obvious because Wu takes its input recording as already in hand and recites neither the act of obtaining that recording from a set of cameras nor the registration of those cameras that its own rendering step consumes as a view matrix, and Lv supplies both for the same class of dynamic scene reconstruction. A person of ordinary skill working from Wu had to specify where the frames and the view matrices came from before the pipeline could run at all, and would have looked to Lv, which addresses that same acquisition problem for the same kind of dynamic scene, with the predictable result that Wu's deformation network and four-dimensional model are determined for a multi-view video whose camera poses are known.
Kopf discloses generating a three-dimensional scene model corresponding to the multi-view images (col 6, “To enable better sharing and preservation of immersive experiences, a graphics system reconstructs a three-dimensional scene from a set of images of the scene taken from different vantage points”; also, col 6, “FIG. 1 illustrates a system for reconstructing a three-dimensional scene from a set of images, in accordance with one embodiment. As depicted, an image capture system 110 (e.g., a camera) is used to take a set of images 115 from different viewing positions in a scene and outputs the images 115 to a three-dimensional (3D) photo reconstruction system 120. The three-dimensional photo reconstruction system 120 processes the images 115 to generate a three-dimensional renderable panoramic image 125.”; also, col 10, “Generally, the input images include a set of images of a scene captured from different vantage points, with at least some of the input images overlapping other images of the scene.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, in which a deformation network and a four-dimensional model are determined from an explicit three-dimensional point cloud representation for a multi-view video whose camera poses are known, with the features of Kopf's invention of a reconstruction system that takes a set of images of one scene captured from different vantage points and processes those images to reconstruct a three-dimensional scene from them. The combination would have been obvious because the three-dimensional scene representation that Wu's deformation network operates on must exist before that network can be determined, and while Wu identifies structure-from-motion points as the source of its initialization it treats that starting point as given and recites no act of building a three-dimensional scene from the captured images. Kopf performs exactly that act on exactly that input, reconstructing a three-dimensional scene from overlapping images of one scene taken from different vantage points. A person of ordinary skill needed the static three-dimensional scene that the method then deforms across time, and would have adopted Kopf's image-to-scene reconstruction to obtain it, with the predictable result that the three-dimensional scene model carried into Wu's deformation step is one generated from the multi-view images the method already obtains.
Regarding claim 3, Wu as modified by Lv and Kopf discloses the method according to claim 1, wherein Lv further discloses the method further comprises: determining a specified moment and a specified view; and outputting scene information corresponding to the specified moment and the specified view (para 61, “For example, at step 750 after training completes, the computing system may render output frames for an output video of the scene. Each output frame may be rendered by querying the updated NeRF using one of the updated latent codes corresponding to a desired time associated with the output frame, a desired viewpoint for the output frame, and ray directions associated with pixels in the output frame.”; also, para 30, “For example, each frame of the new video may be generated by querying NeRF using any desired viewpoint (including view position and direction in 3D space), field of view, and/or time.”; also, para 39, “By processing a given latent code 210 and a desired view ray 220 for a pixel, NeRF 200 outputs a color 230 and opacity 240 of the pixel corresponding to the view ray 220 for the moment in time that corresponds to the latent code 210.”; also, para 39, “The output frame 250b is an example of the dynamic cooking scene at time t as viewed from the perspective of the view position p.”; also, since when there exist a database that stores information regarding every output frames of each video frame, querying for a given frame by a specific viewpoint and time indicates a moment of when a given frame was captured based on a specific cameras perspective).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of a four-dimensional scene model formed from an explicit three-dimensional point cloud representation and the deformation a network predicts for a given timestamp, with the features of Lv's invention of rendering an output frame by fixing a desired time and a desired viewpoint for that frame, querying the trained representation with them, and outputting the color and opacity of each pixel for the moment in time that query specifies. The combination would have been obvious because Wu renders a novel view for a given view matrix and timestamp but does not recite determining a specified moment and a specified view as the selections that drive that rendering, nor outputting the resulting scene information for that pair, and Lv supplies both acts for the same kind of representation. A person of ordinary skill who had trained Wu's model would have needed a way to obtain a particular view at a particular instant from it, and would have adopted Lv's query-and-render step, with the predictable result of scene information being output for whatever moment and view are specified.
Regarding claim 12, Wu discloses video frame at an initial moment in the multi-view video (sec 5.1, “We use the points computed by SfM[32] from the first frame of each video in Neu3D’s dataset and 200 frames randomly selected in HyperNeRF’s.”; also, initial moment is also referred to as the first frame of a video); deformable network (sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian deformation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, determining deformable network is performed by the Gaussian deformation field network); three-dimensional scene model (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds.”; also, sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian deformation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the de formed 3D Gaussians G′ can be introduced.”; also, sec 4.1, “As shown in Fig. 3, given a view matrix M = [R,T], times tamp t, our 4D Gaussian splatting framework includes 3D Gaussians G and Gaussian deformation field network F.”), determine a four-dimensional scene model (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds”; also, sec 1, “An efficient 4D Gaussian splatting framework with an efficient Gaussian deformation field is proposed by modeling both Gaussian motion and Gaussian shape changes across time.”; also, sec 4.1, “Our 4D Gaussian splatting converts the original 3D Gaussians G into another group of 3D Gaussians G′ given a timestamp t, maintaining the effectiveness of the differential splatting as referred in [46]”; also, 4D Gaussian splatting is a holistic representation for dynamic scenes) three- dimensional scene model and the deformable network (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds”; also, sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian de formation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, sec 4.1, “Then a novel-view image ˆIis rendered by differential splat ting [46] S following ˆI= S(M,G′), where G′ = ∆G +G”; also, sec 4.2, “Finally, we obtain the deformed 3D Gaussians G′ = {X′,s′,r′,σ,C}.”; also, 3D Gaussian G is referred to as a scene. 4D scene was created with both the 3D gaussians G and the Gaussian deformation field network). Wu does not disclose an electronic device, comprising: one or more processors; and a storage apparatus having one or more programs stored thereon, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to: obtain a multi-view video, wherein the multi-view video comprises multi-view images, generate a three-dimensional scene model corresponding to the multi-view images; corresponding to the multi-view video, the multi-view video, and camera pose information corresponding to the multi-view video; corresponding to the multi-view video.
However, in a similar field of endeavor, Lv discloses an electronic device, comprising: one or more processors; and a storage apparatus having one or more programs stored thereon, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to:
obtain a multi-view video, wherein the multi-view video comprises multi-view images (para 29, “More formally, embodiments described herein may be used to tackle the problem of reconstructing dynamic 3D scenes from video inputs from multiple cameras, i.e., {C.sup.(t)} for time index t∈T={1, 2, . . . , T} with known camera intrinsic and extrinsic parameters.”; also, para 29, “The dynamic scene may be simultaneously recorded by several cameras (e.g., 5, 10, or 20 cameras) over a period of time”; also, para 58, “We first train a NeRF model on the keyframes, which we sample equidistant from the multi-view image sequence at fixed intervals K.”), corresponding to the multi-view video (para 62, “We have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation.”; also, para 48, “The number of iterations it takes in training per epoch scales linearly with the number of the pixels in the multi-view videos.”; also, para 41, “Conceptually, NeRF learns and encodes the radiance and opacity values of a dynamic scene over time based on video frames captured by multiple cameras.”), the multi-view video, and camera pose information corresponding to the multi-view video (para 29, “More formally, embodiments described herein may be used to tackle the problem of reconstructing dynamic 3D scenes from video inputs from multiple cameras, i.e., {C.sup.(t)} for time index t∈T={1, 2, . . . , T} with known camera intrinsic and extrinsic parameters.”; also, para 42, “More formally, given a ray r(s)=o+sd (origin o and direction d defined by the specified camera pose and camera intrinsics), the rendered color of the pixel corresponding to this ray C(r) is an integral over the radiance weighted by accumulated opacity:”; also, para 41, “Conceptually, NeRF learns and encodes the radiance and opacity values of a dynamic scene over time based on video frames captured by multiple cameras.”); corresponding to the multi-view video (para 62, “We have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation.”; also, para 48, “The number of iterations it takes in training per epoch scales linearly with the number of the pixels in the multi-view videos.”; also, para 41, “Conceptually, NeRF learns and encodes the radiance and opacity values of a dynamic scene over time based on video frames captured by multiple cameras.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of determining a Gaussian deformation field network from an explicit three-dimensional scene representation of point clouds and a timestamp and determining a four-dimensional model as that representation combined with the deformation the network predicts, with the features of Lv's invention of a computer system having a processor and a memory holding the instructions making up a computer program that the processor fetches and executes to perform the described steps, operating on a dynamic scene simultaneously recorded by several cameras whose video inputs are obtained together with their known camera intrinsic and extrinsic parameters and whose recordings form a multi-view image sequence the representation is made to correspond to. The combination would have been obvious because Wu describes its pipeline as a sequence of operations without reciting the processor and storage apparatus that carry it out, and takes both its input recording and the registration of the cameras that produced it as already in hand, and Lv supplies the hardware and the acquisition for the same class of dynamic scene reconstruction. A person of ordinary skill reducing Wu's pipeline to a working device had to select hardware to run it on and a registered multi-camera capture to feed it, and would have adopted Lv's, with the predictable result of processors executing programs held in a storage apparatus to determine the deformation network and the four-dimensional model for a multi-view video whose camera poses are known.
Kopf discloses generate a three-dimensional scene model corresponding to the multi-view images (col 6, “To enable better sharing and preservation of immersive experiences, a graphics system reconstructs a three-dimensional scene from a set of images of the scene taken from different vantage points”; also, col 6, “FIG. 1 illustrates a system for reconstructing a three-dimensional scene from a set of images, in accordance with one embodiment. As depicted, an image capture system 110 (e.g., a camera) is used to take a set of images 115 from different viewing positions in a scene and outputs the images 115 to a three-dimensional (3D) photo reconstruction system 120. The three-dimensional photo reconstruction system 120 processes the images 115 to generate a three-dimensional renderable panoramic image 125.”; also, col 10, “Generally, the input images include a set of images of a scene captured from different vantage points, with at least some of the input images overlapping other images of the scene.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, in which processors executing a stored program determine a deformation network and a four-dimensional model from an explicit three-dimensional point cloud representation for a multi-view video whose camera poses are known, with the features of Kopf's invention of a reconstruction system that takes a set of images of one scene captured from different vantage points and processes those images to reconstruct a three-dimensional scene from them. The combination would have been obvious because the three-dimensional scene representation on which Wu's deformation network operates must exist before that network can be determined, and while Wu identifies structure-from-motion points as the source of its initialization it treats that starting point as given and recites no act of building a three-dimensional scene from the captured images. Kopf performs that act on that input, reconstructing a three-dimensional scene from overlapping images of one scene taken from different vantage points. A person of ordinary skill needed the static three-dimensional scene that the device then deforms across time, and would have adopted Kopf's image-to-scene reconstruction to obtain it, with the predictable result that the three-dimensional scene model the device deforms is one generated from the multi-view images it already obtains.
Regarding claim 13, Wu discloses video frame at an initial moment in the multi-view video (sec 5.1, “We use the points computed by SfM[32] from the first frame of each video in Neu3D’s dataset and 200 frames randomly selected in HyperNeRF’s.”; also, initial moment is also referred to as the first frame of a video); determine a deformable network (sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian deformation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, determining deformable network is performed by the Gaussian deformation field network); three-dimensional scene model (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds”; also, sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian de formation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, sec 4.1, “As shown in Fig. 3, given a view matrix M = [R,T], times tamp t, our 4D Gaussian splatting framework includes 3D Gaussians G and Gaussian deformation field network F.”; also, three-dimensional scene model was discuses since when 4D Gaussian splatting framework request both the deformation field network and the 3D Gaussians G is represented as the scene), determine a four-dimensional scene model (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds”; also, sec 1, “An efficient 4D Gaussian splatting framework with an efficient Gaussian deformation field is proposed by modeling both Gaussian motion and Gaussian shape changes across time.”; also, sec 4.1, “Our 4D Gaussian splatting converts the original 3D Gaussians G into another group of 3D Gaussians G′ given a timestamp t, maintaining the effectiveness of the differential splatting as referred in [46]”; also, 4D Gaussian splatting is a holistic representation for dynamic scenes) three- dimensional scene model and the deformable network (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds”; also, sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian de formation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, sec 4.1, “Then a novel-view image ˆIis rendered by differential splat ting [46] S following ˆI= S(M,G′), where G′ = ∆G +G”; also, sec 4.2, “Finally, we obtain the deformed 3D Gaussians G′ = {X′,s′,r′,σ,C}.”; also, 3D Gaussian G is referred to as a scene. 4D scene was created with both the 3D gaussians G and the Gaussian deformation field network). Wu does not disclose a non-transitory computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, causes the processor to: obtain a multi-view video, wherein the multi-view video comprises multi-view images, generate a three-dimensional scene model corresponding to the multi-view images; corresponding to the multi-view video, the multi-view video, and camera pose information corresponding to the multi-view video; corresponding to the multi-view video.
However, in a similar field of endeavor, Lv discloses obtain a multi-view video, wherein the multi-view video comprises multi-view images (para 29, “More formally, embodiments described herein may be used to tackle the problem of reconstructing dynamic 3D scenes from video inputs from multiple cameras, i.e., {C.sup.(t)} for time index t∈T={1, 2, . . . , T} with known camera intrinsic and extrinsic parameters.”; also, para 29, “The dynamic scene may be simultaneously recorded by several cameras (e.g., 5, 10, or 20 cameras) over a period of time”; also, para 58, “We first train a NeRF model on the keyframes, which we sample equidistant from the multi-view image sequence at fixed intervals K.”), corresponding to the multi-view video (para 62, “We have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation”; also, para 48, “The number of iterations it takes in training per epoch scales linearly with the number of the pixels in the multi-view videos.”), the multi-view video, and camera pose information corresponding to the multi-view video; (para 29, “More formally, embodiments described herein may be used to tackle the problem of reconstructing dynamic 3D scenes from video inputs from multiple cameras, i.e., {C.sup.(t)} for time index t∈T={1, 2, . . . , T} with known camera intrinsic and extrinsic parameters.”; also, para 42, “More formally, given a ray r(s)=o+sd (origin o and direction d defined by the specified camera pose and camera intrinsics), the rendered color of the pixel corresponding to this ray C(r) is an integral over the radiance weighted by accumulated opacity:”; also, para 41, “Conceptually, NeRF learns and encodes the radiance and opacity values of a dynamic scene over time based on video frames captured by multiple cameras.”), corresponding to the multi-view video (para 62, “We have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation.”; also, para 41, “Conceptually, NeRF learns and encodes the radiance and opacity values of a dynamic scene over time based on video frames captured by multiple cameras.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of determining a Gaussian deformation field network from an explicit three-dimensional scene representation of point clouds and a timestamp and determining a four-dimensional model as that representation combined with the deformation the network predicts, with the features of Lv's invention of a dynamic scene simultaneously recorded by several cameras whose video inputs are obtained together with their known camera intrinsic and extrinsic parameters and whose recordings form a multi-view image sequence the representation is made to correspond to. The combination would have been obvious because Wu takes its input recording as already in hand and recites neither the act of obtaining that recording from a set of cameras nor the registration of those cameras that its own rendering step consumes as a view matrix, and Lv supplies both for the same class of dynamic scene reconstruction. A person of ordinary skill working from Wu had to specify where the frames and the view matrices came from before the stored program could run at all, and would have looked to Lv, which addresses that same acquisition problem for the same kind of dynamic scene, with the predictable result that the deformation network and four-dimensional model the program determines are determined for a multi-view video whose camera poses are known.
Kopf discloses a non-transitory computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, causes the processor to (para 10, “The software may be embodied, for example, as instructions on a non-transitory computer-readable storage medium of a camera, client computer device, or on a cloud server communicatively coupled to the camera”; also, col 6, “Each of the illustrated systems 110, 120, 130 may include one or more processors and a computer-readable storage medium that stores instructions that when executed cause the respective systems to carry out the processes and functions attributed to the systems 110, 120, 130 described herein.”; also, col 16, “In one embodiment, a software module is implemented with a computer program product comprising a computer-readable medium containing computer program code, which can be executed by a computer processor for performing any or all of the steps, operations, or processes described.”), and generate a three-dimensional scene model corresponding to the multi-view images (col 6, “To enable better sharing and preservation of immersive experiences, a graphics system reconstructs a three-dimensional scene from a set of images of the scene taken from different vantage points.”; also, col 6, “FIG. 1 illustrates a system for reconstructing a three-dimensional scene from a set of images, in accordance with one embodiment. As depicted, an image capture system 110 (e.g., a camera) is used to take a set of images 115 from different viewing positions in a scene and outputs the images 115 to a three-dimensional (3D) photo reconstruction system 120. The three-dimensional photo reconstruction system 120 processes the images 115 to generate a three-dimensional renderable panoramic image 125.”; also, col 9, “Generally, the input images include a set of images of a scene captured from different vantage points, with at least some of the input images overlapping other images of the scene.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, in which a deformation network and a four-dimensional model are determined from an explicit three-dimensional point cloud representation for a multi-view video whose camera poses are known, with the features of Kopf's invention of software carried as instructions on a non-transitory computer-readable storage medium that stores instructions that when executed cause a processor to carry out the described processes, and of a reconstruction system that takes a set of images of one scene captured from different vantage points and processes those images to reconstruct a three-dimensional scene from them. The combination would have been obvious because Wu describes a pipeline of operations without reciting any article of manufacture that carries the program performing them, and because the three-dimensional scene representation on which its deformation network operates must exist before that network can be determined, while Wu treats its structure-from-motion starting point as given and recites no act of building a three-dimensional scene from the captured images. Kopf supplies both, distributing the software on a non-transitory medium whose stored instructions cause the processor to carry out the processes, and reconstructing a three-dimensional scene from overlapping images of one scene taken from different vantage points. A person of ordinary skill needed a way to distribute and execute the pipeline and a way to build the static three-dimensional scene it then deforms, and would have taken Kopf's, with the predictable result of a stored program that when executed causes a processor to generate the three-dimensional scene model from the multi-view images and to deform it across time.
Regarding claim 15, Wu as modified by Lv and Kopf discloses the non-transitory computer-readable medium according to claim 13, wherein Lv further discloses the computer program further causes the processor to: determine a specified moment and a specified view; and output scene information corresponding to the specified moment and the specified view (para 61, “For example, at step 750 after training completes, the computing system may render output frames for an output video of the scene. Each output frame may be rendered by querying the updated NeRF using one of the updated latent codes corresponding to a desired time associated with the output frame, a desired viewpoint for the output frame, and ray directions associated with pixels in the output frame.”; also, para 30, “For example, each frame of the new video may be generated by querying NeRF using any desired viewpoint (including view position and direction in 3D space), field of view, and/or time.”; also, para 39, “By processing a given latent code 210 and a desired view ray 220 for a pixel, NeRF 200 outputs a color 230 and opacity 240 of the pixel corresponding to the view ray 220 for the moment in time that corresponds to the latent code 210.”; also, para 39, “The output frame 250b is an example of the dynamic cooking scene at time t as viewed from the perspective of the view position p.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of a four-dimensional scene model formed from an explicit three-dimensional point cloud representation and the deformation a network predicts for a given timestamp, with the features of Lv's invention of rendering an output frame by fixing a desired time and a desired viewpoint for that frame, querying the trained representation with them, and outputting the color and opacity of each pixel for the moment in time that query specifies. The combination would have been obvious because Wu renders a novel view for a given view matrix and timestamp but does not recite determining a specified moment and a specified view as the selections that drive that rendering, nor outputting the resulting scene information for that pair, and Lv supplies both acts for the same kind of representation. A person of ordinary skill who had trained Wu's model would have needed a way to obtain a particular view at a particular instant from it, and would have adopted Lv's query-and-render step, with the predictable result of the stored program outputting scene information for whatever moment and view are specified.
Claim(s) 2 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering”, 12/07/2023, arXiv, 2310.08528v2 as modified by Lv et al. (U.S. Pub. No. 20220239844) and Kopf et al. (U.S. Doc. No. 10038894), further in view of Kerbl et al. " 3D Gaussian Splatting for Real-Time Radiance Field Rendering", 08/08/2023, arXiv, 2308.04079v1.
Regarding claim 2, Wu as modified by Lv and Kopf discloses method according to claim 1, wherein Wu further discloses determining a deformable network corresponding to the multi-view video based on the three-dimensional scene model, the multi-view video, and camera pose information corresponding to the multi-view video comprises: inputting a target moment into an initial deformable network, to obtain an offset corresponding to the target moment (sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian deformation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the de formed 3D Gaussians G′ can be introduced.”; also, sec 4.2, “Separate MLPs are employed to compute the deformation of position ∆X = ϕx(fd), rotation ∆r = ϕr(fd), and scaling ∆s = ϕs(fd). Then, the deformed feature (X′,r′,s′) can be addressed as: (X′,r′,s′) = (X +∆X,r +∆r,s+∆s).”; also, fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, a target moment was disclosed as the timestamp t in in Gaussian deformation field network formula and the spatial features of 3D Gaussians formula, initial deformable network is referred to as the initial frame of the video while the deformable network is the Gaussian deformation field network, the offset is shown by the formula for computing the position, rotation and scaling); obtaining a three-dimensional scene model at the target moment based on the offset corresponding to the target moment, and the three-dimensional scene model (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds.”; also, sec , “Our 4D Gaussian splatting converts the original 3D Gaussians G into another group of 3D Gaussians G′ given a timestamp t, maintaining the effectiveness of the differential splatting as referred in [46].”; also, sec 4.1, “Then a novel-view image ˆIis rendered by differential splatting [46] S following ˆI= S(M,G′), where G′ = ∆G +G”; also, sec 4.2, “Finally, we obtain the deformed 3D Gaussians G′ = {X′,s′,r′,σ,C}.”; also, the obtained 3d model scene is the deformed 3d gaussians g, target moment is still referred to as the timestamp t, the offset corresponding to G′ = ∆G +G,); and optimizing the initial deformable network using the image loss value, to obtain the deformable network (Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, para 4.3, “Similar to other reconstruction methods [6, 14, 28], we use the L1 color loss to supervise the training process.”) projecting, for each of a plurality of views, the three-dimensional scene model at the target moment from the view, comparing a projected image in the view with a multi-view image corresponding to the view at the target moment, to obtain an image loss value, and corresponding to the multi-view video.
However, in a similar field of endeavor, Lv discloses corresponding to the multi-view video (para 62, “have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation.”; also, para 48, “The number of iterations it takes in training per epoch scales linearly with the number of the pixels in the multi-view videos”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of feeding a timestamp to a Gaussian deformation field network to obtain a per-Gaussian deformation, adding that deformation to the original Gaussians to obtain the deformed representation at that timestamp, and optimizing that network under a color loss, with the features of Lv's invention of a representation made to correspond to real-world multi-view video recordings of dynamic scenes. The combination would have been obvious because Wu optimizes its deformation network against whatever recording it is given without reciting that the resulting network corresponds to a multi-view video, and Lv establishes that correspondence for the same kind of dynamic scene representation. A person of ordinary skill training Wu's network needed a source of supervision, and would have used the multi-view recordings Lv describes, with the predictable result that the deformable network obtained by the optimization is one corresponding to the multi-view video.
Kerbl discloses projecting, for each of a plurality of views, the three-dimensional scene model at the target moment from the view, comparing a projected image in the view with a multi-view image corresponding to the view at the target moment, to obtain an image loss value (sec 4, “However, we need to project our 3D Gaussians to 2D for rendering. Zwicker et al. [2001a] demonstrate how to do this projection to image space.”; also, sec 5.1, “The optimization is based on successive iterations of rendering and comparing the resulting image to the training views in the captured dataset.”; also, sec B, “𝑉, ˆ𝐼←SampleTrainingView() ⊲Camera𝑉andImage 𝐼←Rasterize(𝑀,𝑆,𝐶,𝐴,𝑉) ⊲Alg.2 𝐿←𝐿𝑜𝑠𝑠(𝐼, ˆ𝐼) ⊲Loss 𝑀,𝑆,𝐶,𝐴←Adam(∇𝐿) ⊲Backprop&Step”; also, sec 5.1, “The loss function is L1 combined with a D-SSIM term: L =(1−𝜆)L1 +𝜆LD-SSIM We use 𝜆 = 0.2inallourtests.”; also, sec 1, “Our results on previously published datasets show that we can optimize our 3D Gaussians from multi-view captures and achieve equal or better quality than the best quality previous implicit radiance field approaches.”; also, sample training view function would be the plurality of views, two methods here first indicate the camera and the image and the sample training view which is being compared, which belongs to I as a parameter of the function Loss and rasterize is the projection to the same I to that results in the image loss, and target moment has already been disclosed in the dependent claim).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, further in view of Kopf, in which a timestamp fed to a deformation network yields a per-Gaussian deformation that is added to a three-dimensional point cloud representation generated from the captured images to obtain the representation at that timestamp, and in which that network is optimized under a color loss against a multi-view video, with the features of Kerbl's invention of projecting the three-dimensional Gaussians to image space, sampling a training view and its captured image, rasterizing the Gaussians from that view to produce a rendered image, and taking the loss between the rendered image and the captured image of that view. The combination would have been obvious because Wu states only that a color loss supervises its training and never identifies the operands that loss is computed between, leaving the quantity that drives the optimization unspecified, while Kerbl sets out the render-and-compare loop that produces it and does so for the same explicit three-dimensional Gaussian representation Wu adopts and cites as the foundation of its own framework. A person of ordinary skill implementing Wu's optimization had to decide what the color loss compares, and would have taken Kerbl's per-view comparison of the rendered image against the captured image of that view, both because Wu builds directly on that work and because the deformed Gaussians Wu produces are rendered by the same differential splatting operation Kerbl's loop rasterizes, with the predictable result that the deformation network is optimized on an image loss value obtained by projecting the representation at the target moment into each of a plurality of views and comparing each projected image against the multi-view image for that view.
Regarding claim 14, Wu as modified by Lv and Kopf discloses the non-transitory computer-readable medium according to claim 13, wherein Wu further discloses the computer program for determining a deformable network corresponding to the multi-view video based on the three-dimensional scene model, the multi-view video, and camera pose information corresponding to the multi-view video further causes the processor to: input a target moment into an initial deformable network, to obtain an offset corresponding to the target moment (sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian deformation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the de formed 3D Gaussians G′ can be introduced.”; also, sec 4.2, “Separate MLPs are employed to compute the deformation of position ∆X = ϕx(fd), rotation ∆r = ϕr(fd), and scaling ∆s = ϕs(fd). Then, the deformed feature (X′,r′,s′) can be addressed as: (X′,r′,s′) = (X +∆X,r +∆r,s+∆s).”; also, fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, a target moment was disclosed as the timestamp t in in Gaussian deformation field network formula and the spatial features of 3D Gaussians formula, initial deformable network is referred to as the initial frame of the video while the deformable network is the Gaussian deformation field network, the offset is shown by the formula for computing the position, rotation and scaling); obtain a three-dimensional scene model at the target moment based on the offset corresponding to the target moment, and the three-dimensional scene model (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds.”; also, sec , “Our 4D Gaussian splatting converts the original 3D Gaussians G into another group of 3D Gaussians G′ given a timestamp t, maintaining the effectiveness of the differential splatting as referred in [46].”; also, sec 4.1, “Then a novel-view image ˆIis rendered by differential splatting [46] S following ˆI= S(M,G′), where G′ = ∆G +G”; also, sec 4.2, “Finally, we obtain the deformed 3D Gaussians G′ = {X′,s′,r′,σ,C}.”; also, the obtained 3d model scene is the deformed 3d gaussians g, target moment is still referred to as the timestamp t, the offset corresponding to G′ = ∆G +G,); and optimize the initial deformable network using the image loss value, to obtain the deformable network (Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, para 4.3, “Similar to other reconstruction methods [6, 14, 28], we use the L1 color loss to supervise the training process.”) project, for each of a plurality of views, the three-dimensional scene model at the target moment from the view, compare a projected image in the view with a multi-view image corresponding to the view at the target moment, to obtain an image loss value, corresponding to the multi-view video.
However, in a similar field of endeavor, Lv discloses corresponding to the multi-view video (para 62, “have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation.”; also, para 48, “The number of iterations it takes in training per epoch scales linearly with the number of the pixels in the multi-view videos”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of feeding a timestamp to a Gaussian deformation field network to obtain a per-Gaussian deformation, adding that deformation to the original Gaussians to obtain the deformed representation at that timestamp, and optimizing that network under a color loss, with the features of Lv's invention of a representation made to correspond to real-world multi-view video recordings of dynamic scenes. The combination would have been obvious because Wu optimizes its deformation network against whatever recording it is given without reciting that the resulting network corresponds to a multi-view video, and Lv establishes that correspondence for the same kind of dynamic scene representation. A person of ordinary skill training Wu's network needed a source of supervision, and would have used the multi-view recordings Lv describes, with the predictable result that the deformable network the stored program obtains by the optimization is one corresponding to the multi-view video.
Kerbl discloses project, for each of a plurality of views, the three-dimensional scene model at the target moment from the view, compare a projected image in the view with a multi-view image corresponding to the view at the target moment, to obtain an image loss value (sec 4, “However, we need to project our 3D Gaussians to 2D for rendering. Zwicker et al. [2001a] demonstrate how to do this projection to image space.”; also, sec 5.1, “The optimization is based on successive iterations of rendering and comparing the resulting image to the training views in the captured dataset.”; also, sec B, “𝑉, ˆ𝐼←SampleTrainingView() ⊲Camera𝑉andImage 𝐼←Rasterize(𝑀,𝑆,𝐶,𝐴,𝑉) ⊲Alg.2 𝐿←𝐿𝑜𝑠𝑠(𝐼, ˆ𝐼) ⊲Loss 𝑀,𝑆,𝐶,𝐴←Adam(∇𝐿) ⊲Backprop&Step”; also, sec 5.1, “The loss function is L1 combined with a D-SSIM term: L =(1−𝜆)L1 +𝜆LD-SSIM We use 𝜆 = 0.2inallourtests.”; also, sec 1, “Our results on previously published datasets show that we can optimize our 3D Gaussians from multi-view captures and achieve equal or better quality than the best quality previous implicit radiance field approaches.”; also, sample training view function would be the plurality of views, two methods here first indicate the camera and the image and the sample training view which is being compared, which belongs to I as a parameter of the function Loss and rasterize is the projection to the same I to that results in the image loss, and target moment has already been disclosed in the dependent claim).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, further in view of Kopf, in which a stored program feeds a timestamp to a deformation network to obtain a per-Gaussian deformation that is added to a three-dimensional point cloud representation generated from the captured images, and optimizes that network under a color loss against a multi-view video, with the features of Kerbl's invention of projecting the three-dimensional Gaussians to image space, sampling a training view and its captured image, rasterizing the Gaussians from that view to produce a rendered image, and taking the loss between the rendered image and the captured image of that view. The combination would have been obvious because Wu states only that a color loss supervises its training and never identifies the operands that loss is computed between, leaving the quantity that drives the optimization unspecified, while Kerbl sets out the render-and-compare loop that produces it and does so for the same explicit three-dimensional Gaussian representation Wu adopts and cites as the foundation of its own framework. A person of ordinary skill implementing Wu's optimization had to decide what the color loss compares, and would have taken Kerbl's per-view comparison of the rendered image against the captured image of that view, both because Wu builds directly on that work and because the deformed Gaussians Wu produces are rendered by the same differential splatting operation Kerbl's loop rasterizes, with the predictable result that the stored program optimizes the deformation network on an image loss value obtained by projecting the representation at the target moment into each of a plurality of views and comparing each projected image against the multi-view image for that view.
Claim(s) 4 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering”, 12/07/2023, arXiv, 2310.08528v2 as modified by Lv et al. (U.S. Pub. No. 20220239844) and Kopf et al. (U.S. Doc. No. 10038894), further in view of Martin-Brualla et al. " LookinGood: Enhancing Performance Capture with Real-time Neural Re-Rendering", 12/04/2018, ACM Transactions on Graphics,Vol. 37, No. 6.
Regarding claim 4, Wu as modified by Lv and Kopf discloses the method according to claim 1, wherein Wu further discloses the method, further comprises: projecting, for each of a plurality of views, the four-dimensional scene model from the view (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds.”; also, sec 4.1, “Then a novel-view image ˆIis rendered by differential splat ting [46] S following ˆI= S(M,G′), where G′ = ∆G +G.”; also, 5.4, “differential rendering [46] can project all the point clouds into viewpoints by ˆI= S(M,G′).”), target network comprises the deformable network (sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian deformation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, determining deformable network is performed by the Gaussian deformation field network). Wu does not disclose comparing a projected video in the view with a multi-view video corresponding to the view, to obtain a video loss value, and optimizing a network parameter of a target network using the video loss value.
However, in a similar field of endeavor, Martin-Brualla discloses comparing a projected video in the view with a multi-view video corresponding to the view, to obtain a video loss value, and optimizing a network parameter of a target network using the video loss value (sec 3.2, “Given an image I rendered from a volumetric reconstruction, we want to compute an enhanced version of I, that we denote by Ie.”; also, sec 3.2, “To minimize the amount of flickering between two consecutive frames, we design a temporal loss between a frame It and It−1. A simple loss minimizing the difference between It and It−1 would produce temporally blurred results, and thus we use a loss that tries to match the temporal gradient of the predicted sequence, i.e. It pred −It−1 pred , with the temporal gradient of the ground truth sequence, i.e. It дt − It−1 дt . In particular, the loss is computed as Ltemporal = ∥(It pred − It−1 pred) − (It дt − It−1 дt )∥1 .”)”; also, sec 3.1, “To optimize for F(I), we train a neural network to optimize the loss function”; also, sec 3.1, “we mount additional “witness” color cameras to the existing capture rigs, that capture higher quality images from different viewpoints.”; also, sec 4, “In the full body capturerig,wemounted8‘high’resolution(4096× 2048) witness cameras”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, further in view of Kopf, in which a deformation network is determined from an explicit three-dimensional point cloud representation generated from a multi-view video and the resulting four-dimensional model is projected into viewpoints by differential rendering, with the features of Martin-Brualla's invention of a loss that matches the temporal gradient of the predicted sequence against the temporal gradient of the ground truth sequence captured by cameras at different viewpoints, and of training the network by optimizing that loss. The combination would have been obvious because Wu evaluates its reconstruction one rendered frame at a time and identifies no term that measures the reconstruction as a sequence, so nothing in its objective penalizes a result whose individual frames are each accurate while the sequence flickers, and Martin-Brualla addresses precisely that failure for the same kind of rendered output taken from a volumetric reconstruction and compared against camera captures from different viewpoints. A person of ordinary skill reconstructing a moving scene would have recognized that a per-frame objective leaves temporal incoherence unpenalized, and would have added Martin-Brualla's sequence-level term to the objective that trains the deformation network, with the predictable result of a video loss value obtained by comparing the projected sequence for a view against the captured sequence for that view and used to optimize the network parameters.
Regarding claim 16, Wu as modified by Lv and Kopf discloses the non-transitory computer-readable medium according to claim 13, wherein Wu further discloses the computer program, further causes the processor to: project, for each of a plurality of views, the four-dimensional scene model from the view (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds.”; also, sec 4.1, “Then a novel-view image ˆIis rendered by differential splat ting [46] S following ˆI= S(M,G′), where G′ = ∆G +G.”; also, 5.4, “differential rendering [46] can project all the point clouds into viewpoints by ˆI= S(M,G′).”), target network comprises the deformable network (sec 4.1, “Specifically, the deformation of 3D Gaussians ∆G is introduced by the Gaussian deformation field network ∆G = F(G,t), in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t), and the multi-head Gaussian deformation decoder D can decode the features and predict each 3D Gaussian’s deformation ∆G = D(f), then the deformed 3D Gaussians G′ can be introduced.”; also, Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable.”; also, determining deformable network is performed by the Gaussian deformation field network). Wu does not disclose compare a projected video in the view with a multi-view video corresponding to the view, to obtain a video loss value, and optimize a network parameter of a target network using the video loss value.
However, in a similar field of endeavor, Martin-Brualla discloses compare a projected video in the view with a multi-view video corresponding to the view, to obtain a video loss value, and optimize a network parameter of a target network using the video loss value (sec 3.2, “Given an image I rendered from a volumetric reconstruction, we want to compute an enhanced version of I, that we denote by Ie.”; also, sec 3.2, “To minimize the amount of flickering between two consecutive frames, we design a temporal loss between a frame It and It−1. A simple loss minimizing the difference between It and It−1 would produce temporally blurred results, and thus we use a loss that tries to match the temporal gradient of the predicted sequence, i.e. It pred −It−1 pred , with the temporal gradient of the ground truth sequence, i.e. It дt − It−1 дt . In particular, the loss is computed as Ltemporal = ∥(It pred − It−1 pred) − (It дt − It−1 дt )∥1 .”)”; also, sec 3.1, “To optimize for F(I), we train a neural network to optimize the loss function”; also, sec 3.1, “we mount additional “witness” color cameras to the existing capture rigs, that capture higher quality images from different viewpoints.”; also, sec 4, “In the full body capturerig,wemounted8‘high’resolution(4096× 2048) witness cameras”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, further in view of Kopf, in which a stored program determines a deformation network from an explicit three-dimensional point cloud representation generated from a multi-view video and projects the resulting four-dimensional model into viewpoints by differential rendering, with the features of Martin-Brualla's invention of a loss that matches the temporal gradient of the predicted sequence against the temporal gradient of the ground truth sequence captured by cameras at different viewpoints, and of training the network by optimizing that loss. The combination would have been obvious because Wu evaluates its reconstruction one rendered frame at a time and identifies no term that measures the reconstruction as a sequence, so nothing in its objective penalizes a result whose individual frames are each accurate while the sequence flickers, and Martin-Brualla addresses precisely that failure for the same kind of rendered output taken from a volumetric reconstruction and compared against camera captures from different viewpoints. A person of ordinary skill reconstructing a moving scene would have recognized that a per-frame objective leaves temporal incoherence unpenalized, and would have added Martin-Brualla's sequence-level term to the objective that trains the deformation network, with the predictable result of a video loss value obtained by comparing the projected sequence for a view against the captured sequence for that view and used to optimize the network parameters.
Claim(s) 5 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering”, 12/07/2023, arXiv, 2310.08528v2 as modified by Lv et al. (U.S. Pub. No. 20220239844) and Kopf et al. (U.S. Doc. No. 10038894), Martin-Brualla et al. " LookinGood: Enhancing Performance Capture with Real-time Neural Re-Rendering", 12/04/2018, ACM Transactions on Graphics,Vol. 37, No. 6, further in view of Kerbl et al. " 3D Gaussian Splatting for Real-Time Radiance Field Rendering", 08/08/2023, arXiv, 2308.04079v1.
Regarding claim 5, Wu as modified by Lv, Kopf, and Martin-Brualla discloses the method according to claim 4, wherein target network further comprises a three-dimensional Gaussian radiance field, wherein the three-dimensional Gaussian radiance field is used to determine the three- dimensional scene model corresponding to the multi-view images.
However, in a similar field of endeavor, Kerbl further discloses the target network further comprises a three-dimensional Gaussian radiance field, wherein the three-dimensional Gaussian radiance field is used to determine the three- dimensional scene model corresponding to the multi-view images (sec 1, “The introduction of anisotropic 3D Gaussians as a high-quality, unstructured representation of radiance fields.”; also, sec 4, “An obvious approach would be to directly optimize the covariance matrix Σ to obtain 3D Gaussians that represent the radiance field.”; also, sec 3, “The input to our method is a set of images of a static scene, together with the corresponding cameras calibrated by SfM [Schönberger and Frahm 2016] which produces a sparse point cloud as a side effect. From these points we create a set of 3D Gaussians (Sec. 4), defined by a position (mean), covariance matrix and opacity 𝛼, that allows a very flexible optimization regime. This results in a reason ably compact representation of the 3D scene, in part because highly anisotropic volumetric splats can be used to represent fine structures compactly.”; also, sec B, “𝑀,𝑆,𝐶,𝐴←Adam(∇𝐿) ⊲Backprop&Step”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, further in view of Kopf, further in view of Martin-Brualla, in which a deformation network operating on an explicit three-dimensional point cloud representation is trained under a loss that measures the rendered sequence against the captured sequence, with the features of Kerbl's invention of anisotropic three-dimensional Gaussians introduced as a representation of radiance fields, created from a set of images of a scene together with their corresponding calibrated cameras, and optimized by backpropagation of that same loss. The combination would have been obvious because Wu adopts three-dimensional Gaussians as its scene representation but characterizes them only as point clouds and does not recite that the representation from which the model is determined is a radiance field or that its parameters are optimized alongside the deformation network. Kerbl supplies both, naming the Gaussians a representation of radiance fields and stepping their parameters on the gradient of the training loss. A person of ordinary skill working with Wu, which adopts Kerbl's representation and cites it as the foundation of its framework, would have carried across Kerbl's characterization and its joint parameter optimization, with the predictable result that the target network optimized by the loss further comprises a three-dimensional Gaussian radiance field used to determine the three-dimensional scene model.
Regarding claim 17, Wu as modified by Lv, Kopf, and Martin-Brualla discloses the non-transitory computer-readable medium according to claim 16, target network further comprises a three-dimensional Gaussian radiance field, wherein the three-dimensional Gaussian radiance field is used to determine the three-dimensional scene model corresponding to the multi-view images.
However, in a similar field of endeavor, Kerbl discloses wherein the target network further comprises a three-dimensional Gaussian radiance field, wherein the three-dimensional Gaussian radiance field is used to determine the three-dimensional scene model corresponding to the multi-view images (sec 1, “The introduction of anisotropic 3D Gaussians as a high-quality, unstructured representation of radiance fields.”; also, sec 4, “An obvious approach would be to directly optimize the covariance matrix Σ to obtain 3D Gaussians that represent the radiance field.”; also, sec 3, “The input to our method is a set of images of a static scene, together with the corresponding cameras calibrated by SfM [Schönberger and Frahm 2016] which produces a sparse point cloud as a side effect. From these points we create a set of 3D Gaussians (Sec. 4), defined by a position (mean), covariance matrix and opacity 𝛼, that allows a very flexible optimization regime. This results in a reason ably compact representation of the 3D scene, in part because highly anisotropic volumetric splats can be used to represent fine structures compactly.”; also, sec B, “𝑀,𝑆,𝐶,𝐴←Adam(∇𝐿) ⊲Backprop&Step”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, further in view of Kopf, further in view of Martin-Brualla, in which a stored program trains a deformation network operating on an explicit three-dimensional point cloud representation under a loss that measures the rendered sequence against the captured sequence, with the features of Kerbl's invention of anisotropic three-dimensional Gaussians introduced as a representation of radiance fields, created from a set of images of a scene together with their corresponding calibrated cameras, and optimized by backpropagation of that same loss. The combination would have been obvious because Wu adopts three-dimensional Gaussians as its scene representation but characterizes them only as point clouds and does not recite that the representation from which the model is determined is a radiance field or that its parameters are optimized alongside the deformation network. Kerbl supplies both, naming the Gaussians a representation of radiance fields and stepping their parameters on the gradient of the training loss. A person of ordinary skill working with Wu, which adopts Kerbl's representation and cites it as the foundation of its framework, would have carried across Kerbl's characterization and its joint parameter optimization, with the predictable result that the target network optimized by the loss further comprises a three-dimensional Gaussian radiance field used to determine the three-dimensional scene model.
Claim(s) 6-8, 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Wu et al. “4D Gaussian Splatting for Real-Time Dynamic Scene Rendering”, 12/07/2023, arXiv, 2310.08528v2 as modified by Lv et al. (U.S. Pub. No. 20220239844) and Kopf et al. (U.S. Doc. No. 10038894), and Kerbl et al. " 3D Gaussian Splatting for Real-Time Radiance Field Rendering", 08/08/2023, arXiv, 2308.04079v1, further in view of Kopanas et al. " Neural Point Catacaustics for Novel-View Synthesis of Reflections”, 11/30/2022, ACM Transactions on Graphics, Vol 41, No 6.
Regarding claim 6, Wu as modified by Lv, Kopf, and Kerbl discloses the method according to claim 2, wherein Wu discloses inputting a target moment into an initial deformable network, to obtain an offset corresponding to the target moment comprises: combining a temporal feature and a spatial feature of the target moment, to obtain combined feature information (sec 4.2, “Specifically, the spatial-temporal structure encoder H contains 6 multi-resolution plane modules Rl(i,j) and a tiny MLP ϕd, i.e. H(G,t) = {Rl(i,j),ϕd|(i,j) ∈ {(x,y),(x,z),(y,z),(x,t),(y,t),(z,t)},l ∈ {1,2}}”; also, sec 4.2, “fh = interp(Rl(i, j)), l (i, j) ∈ {(x,y),(x,z),(y,z),(x,t),(y,t),(z,t)}.”; also, 4.2, “Then a tiny MLP ϕd merges all the features by fd = ϕd(fh)”; also, sec 4.1, “in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t)”); and inputting the combined feature information into the initial deformable network, to obtain the offset corresponding to the target moment (sec 4.2, “When all the features of 3D Gaussians are encoded, we can com pute any desired variable with a multi-head Gaussian de formation decoder D = {ϕx,ϕr,ϕs}”; also, sec 4.2, “Separate MLPs are employed to compute the deformation of position ∆X = ϕx(fd), rotation ∆r = ϕr(fd), and scaling ∆s = ϕs(fd). Then, the deformed feature (X′,r′,s′) can be addressed as: (X′,r′,s′) = (X +∆X,r +∆r,s+∆s).”; also, Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable”), wherein the temporal feature of the target moment is determined based on the target moment, and the spatial feature is determined based on a point cloud position of the three-dimensional scene (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds.”; also, sec 4.2, “The position X = (x,y,z) is the mean value of 3D Gaussians G.”; also, Fig. 3, “Given a group of 3D Gaussians G, we extract the center coordinate of each 3D Gaussian X and timestamp t to compute the voxel feature by querying multi-resolution voxel planes.”; also, sec 4.2, “This entails encoding information of the 3D Gaussians within the 6 2D voxel planes while considering temporal information.”), and camera pose information corresponding to the multi-view video.
However, in a similar field of endeavor, Lv discloses corresponding to the multi-view video (para 62, “We have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation”; also, para 48, “The number of iterations it takes in training per epoch scales linearly with the number of the pixels in the multi-view videos.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of encoding a Gaussian's position and a timestamp into six plane modules whose interpolated features are merged by a small network and decoded into a per-Gaussian deformation, with the features of Lv's invention of a representation corresponding to real-world multi-view video recordings of dynamic scenes reconstructed from video inputs from multiple cameras whose intrinsic and extrinsic parameters are known. The combination would have been obvious because Wu never states what recording the camera information used in its pipeline belongs to, and Lv establishes that correspondence for the same kind of dynamic scene representation. A person of ordinary skill had to identify the recording the camera information came from, and would have taken Lv's multi-view video, with the predictable result that the camera information consumed in determining the spatial feature is that corresponding to the multi-view video.
Kopanas discloses the camera pose information (sec 4.2, “Formally, given an initial reflection point position p ∈ R3 and the current camera position c ∈ R3, we seek to compute the reflection position p′ ∈ R3 on the Neural Point Catacaustic, using the field F∈ R3 × R3 → R3 via: p′ = p+F(p,c).”; also, sec 4.2, “We realize F using an MLP with trainable parameters warp.”; also, sec 4.2, “The initial point positions p are fixed, not optimized (see Sec. 5.1 for our initialization procedure), but each point is displaced by the neural warp field. Notice that the warp field takes the initial point position p as input, so that we can use the same field for all reflection points,”; also, sec 6.3, “normalize the camera position and the reflection point cloud to the range [−1,1]”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, further in view of Kopf, further in view of Kerbl, in which a per-point feature built from a point cloud position and a timestamp is decoded into an offset that is added to the position, with the features of Kopanas's invention of a warp field realized as a network that takes a point cloud position together with the camera position and returns a displacement added to that position. The combination would have been obvious because Wu builds its per-point feature from the point position and the time alone, so the deformation it predicts cannot account for anything that varies with where the scene is observed from, and Kopanas identifies camera position as a quantity a displacement field can usefully depend on and shows a network that accepts it alongside the point cloud position and still returns a displacement of the same form Wu already adds to its Gaussians. A person of ordinary skill seeking to make the predicted displacement responsive to the observing camera had a network architecture of exactly that signature available in Kopanas, whose input is a point of an explicit point cloud and a camera position and whose output is added to the point, with the predictable result that the spatial feature supplied to the deformation network is determined from the point cloud position together with the camera pose information.
Regarding claim 7, Wu as modified by Lv, Kopf, Kerbl, and Kopanas discloses the method according to claim 6, wherein the deformable network comprises a multilayer perceptron (sec 4.2, “Separate MLPs are employed to compute the deformation of position ∆X = ϕx(fd), rotation ∆r = ϕr(fd), and scaling ∆s = ϕs(fd).”; also, sec 5.3, “Our proposed Gaussian deformation decoder D decodes the features from the spatial-temporal structure encoder H. All the changes in 3D Gaussians can be explained by separate MLPs{ϕx,ϕr,ϕs}”; also, sec 4.2, “we introduce an efficient spatial-temporal structure encoder H including a multi-resolution HexPlane R(i,j) and a tiny MLP ϕd in spired by [4, 6, 8, 34]”; also, sec A.2, “The Gaussian deformation decoder is a tiny MLP with a learning rate of 1.6 × 10−3”).
Regarding claim 8, Wu as modified by Lv, Kopf, Kerbl, and Kopanas discloses the method according to claim 6, wherein Wu further discloses the combined feature information comprises a pairwise combination of the temporal feature and spatial features in three dimensions (sec 4.2, “While the vanilla 4D neural voxel is memory-consuming, we adopt a 4D K-Planes [8] module to decompose the 4D neural voxel into 6 planes.”; also, sec 4.2, “Specifically, the spatial-temporal structure encoder H contains 6 multi-resolution plane modules Rl(i,j) and a tiny MLP ϕd, i.e. H(G,t) = {Rl(i,j),ϕd|(i,j) ∈ {(x,y),(x,z),(y,z),(x,t),(y,t),(z,t)},l ∈ {1,2}}.”; also, sec 4.2, “fh = interp(Rl(i, j)), l (i, j) ∈ {(x,y),(x,z),(y,z),(x,t),(y,t),(z,t)}.”; also, sec 4.2, “The position X = (x,y,z) is the mean value of 3D Gaussians G.”; also, sec 4.2, “Then a tiny MLP ϕd merges all the features by fd = ϕd(fh).”).
Regarding claim 18, Wu as modified by Lv, Kopf, and Kerbl discloses the non-transitory computer-readable medium according to claim 14, wherein Wu further discloses the computer program for inputting a target moment into an initial deformable network, to obtain an offset corresponding to the target moment further causes the processor to: combine a temporal feature and a spatial feature of the target moment, to obtain combined feature information (sec 4.2, “Specifically, the spatial-temporal structure encoder H contains 6 multi-resolution plane modules Rl(i,j) and a tiny MLP ϕd, i.e. H(G,t) = {Rl(i,j),ϕd|(i,j) ∈ {(x,y),(x,z),(y,z),(x,t),(y,t),(z,t)},l ∈ {1,2}}”; also, sec 4.2, “fh = interp(Rl(i, j)), l (i, j) ∈ {(x,y),(x,z),(y,z),(x,t),(y,t),(z,t)}.”; also, 4.2, “Then a tiny MLP ϕd merges all the features by fd = ϕd(fh)”; also, sec 4.1, “in which the spatial-temporal structure encoder H can encode both the temporal and spatial features of 3D Gaussians fd = H(G,t)”); and input the combined feature information into the initial deformable network, to obtain the offset corresponding to the target moment (sec 4.2, “When all the features of 3D Gaussians are encoded, we can com pute any desired variable with a multi-head Gaussian de formation decoder D = {ϕx,ϕr,ϕs}”; also, sec 4.2, “Separate MLPs are employed to compute the deformation of position ∆X = ϕx(fd), rotation ∆r = ϕr(fd), and scaling ∆s = ϕs(fd). Then, the deformed feature (X′,r′,s′) can be addressed as: (X′,r′,s′) = (X +∆X,r +∆r,s+∆s).”; also, Fig. 8c, “avoiding numeric errors in optimizing the Gaussian deformation network F and keeping the training process stable”), wherein the temporal feature of the target moment is determined based on the target moment, and the spatial feature is determined based on a point cloud position of the three-dimensional scene (sec 3.1, “3D Gaussians [14] is an explicit 3D scene representation in the form of point clouds.”; also, sec 4.2, “The position X = (x,y,z) is the mean value of 3D Gaussians G.”; also, Fig. 3, “Given a group of 3D Gaussians G, we extract the center coordinate of each 3D Gaussian X and timestamp t to compute the voxel feature by querying multi-resolution voxel planes.”; also, sec 4.2, “This entails encoding information of the 3D Gaussians within the 6 2D voxel planes while considering temporal information.”), and camera pose information corresponding to the multi-view video.
However, in a similar field of endeavor, Lv discloses corresponding to the multi-view video (para 62, “We have proposed a novel neural 3D video synthesis approach that is able to represent real-world multi-view video recordings of dynamic scenes in a compact, yet expressive representation”; also, para 48, “The number of iterations it takes in training per epoch scales linearly with the number of the pixels in the multi-view videos.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Wu's invention of encoding a Gaussian's position and a timestamp into six plane modules whose interpolated features are merged by a small network and decoded into a per-Gaussian deformation, with the features of Lv's invention of a representation corresponding to real-world multi-view video recordings of dynamic scenes reconstructed from video inputs from multiple cameras whose intrinsic and extrinsic parameters are known. The combination would have been obvious because Wu never states what recording the camera information used in its pipeline belongs to, and Lv establishes that correspondence for the same kind of dynamic scene representation. A person of ordinary skill had to identify the recording the camera information came from, and would have taken Lv's multi-view video, with the predictable result that the camera information consumed in determining the spatial feature is that corresponding to the multi-view video.
Kopanas discloses the camera pose information (sec 4.2, “Formally, given an initial reflection point position p ∈ R3 and the current camera position c ∈ R3, we seek to compute the reflection position p′ ∈ R3 on the Neural Point Catacaustic, using the field F∈ R3 × R3 → R3 via: p′ = p+F(p,c).”; also, sec 4.2, “We realize F using an MLP with trainable parameters warp.”; also, sec 4.2, “The initial point positions p are fixed, not optimized (see Sec. 5.1 for our initialization procedure), but each point is displaced by the neural warp field. Notice that the warp field takes the initial point position p as input, so that we can use the same field for all reflection points,”; also, sec 6.3, “normalize the camera position and the reflection point cloud to the range [−1,1]”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Wu in view of Lv, further in view of Kopf, further in view of Kerbl, in which a stored program builds a per-point feature from a point cloud position and a timestamp and decodes it into an offset that is added to the position, with the features of Kopanas's invention of a warp field realized as a network that takes a point cloud position together with the camera position and returns a displacement added to that position. The combination would have been obvious because Wu builds its per-point feature from the point position and the time alone, so the deformation it predicts cannot account for anything that varies with where the scene is observed from, and Kopanas identifies camera position as a quantity a displacement field can usefully depend on and shows a network that accepts it alongside the point cloud position and still returns a displacement of the same form Wu already adds to its Gaussians. A person of ordinary skill seeking to make the predicted displacement responsive to the observing camera had a network architecture of exactly that signature available in Kopanas, whose input is a point of an explicit point cloud and a camera position and whose output is added to the point, with the predictable result that the spatial feature supplied to the deformation network is determined from the point cloud position together with the camera pose information.
Regarding claim 19, Wu as modified by Lv, Kopf, Kerbl, and Kopanas discloses the non-transitory computer-readable medium according to claim 18, wherein Wu further discloses the deformable network comprises a multilayer perceptron (sec 4.2, “Separate MLPs are employed to compute the deformation of position ∆X = ϕx(fd), rotation ∆r = ϕr(fd), and scaling ∆s = ϕs(fd).”; also, sec 5.3, “Our proposed Gaussian deformation decoder D decodes the features from the spatial-temporal structure encoder H. All the changes in 3D Gaussians can be explained by separate MLPs{ϕx,ϕr,ϕs}”; also, sec 4.2, “we introduce an efficient spatial-temporal structure encoder H including a multi-resolution HexPlane R(i,j) and a tiny MLP ϕd in spired by [4, 6, 8, 34]”; also, sec A.2, “The Gaussian deformation decoder is a tiny MLP with a learning rate of 1.6 × 10−3”).
Regarding claim 20, Wu as modified by Lv, Kopf, Kerbl, and Kopanas discloses the non-transitory computer-readable medium according to claim 18, wherein Wu further discloses the combined feature information comprises a pairwise combination of the temporal feature and spatial features in three dimensions (sec 4.2, “While the vanilla 4D neural voxel is memory-consuming, we adopt a 4D K-Planes [8] module to decompose the 4D neural voxel into 6 planes.”; also, sec 4.2, “Specifically, the spatial-temporal structure encoder H contains 6 multi-resolution plane modules Rl(i,j) and a tiny MLP ϕd, i.e. H(G,t) = {Rl(i,j),ϕd|(i,j) ∈ {(x,y),(x,z),(y,z),(x,t),(y,t),(z,t)},l ∈ {1,2}}.”; also, sec 4.2, “fh = interp(Rl(i, j)), l (i, j) ∈ {(x,y),(x,z),(y,z),(x,t),(y,t),(z,t)}.”; also, sec 4.2, “The position X = (x,y,z) is the mean value of 3D Gaussians G.”; also, sec 4.2, “Then a tiny MLP ϕd merges all the features by fd = ϕd(fh).”).
Allowable Subject Matter
Claim 9-11 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is an examiner’s statement of reasons for allowance:
The element of claim 9 that the prior art does not reach is "performing depth estimation on the multi-view images, to determine sparse point clouds corresponding to the multi-view images." The art produces sparse point clouds and it produces point clouds by depth estimation, but never the one by the other, and it draws that line expressly. Where the art recites a sparse point cloud, that cloud arises during camera registration, described as produced for free as a side effect of the structure-from-motion process, so it is a single scene-level set of points obtained from feature correspondence rather than a per-image quantity determined by estimating depth. Where the art instead performs depth estimation on multi-view images, it warps image features from neighboring viewpoints into a plane-swept cost volume, regresses a depth probability volume, combines the per-plane depth values into a depth map, and unprojects that map into space, which returns a point for every pixel of the view and therefore a dense cloud. Art in this field states the division in a single sentence, performing sparse point cloud reconstruction through an incremental structure-from-motion algorithm to obtain the sparse point cloud and the camera pose corresponding to each image, and then separately performing depth estimation on each image to obtain a depth image from which a dense point cloud is reconstructed. The claimed element requires both properties of one quantity, a cloud that is per-image, that is sparse, and that is the output of depth estimation performed on the multi-view images, and the art that reaches sparsity reaches it by a different operation while the art that performs the recited operation reaches the opposite density. Claim 10 adds that "the three-dimensional scene model comprises a three-dimensional Gaussian radiance field," and claim 11 adds projecting that field from each of a plurality of views and comparing a projected image with the multi-view image corresponding to the view to obtain an image loss value and optimizing a parameter of the field with it. Both additions are reached by the art, which represents scenes as three-dimensional Gaussians characterized as a representation of radiance fields and optimizes their parameters against a loss taken between a rasterized image and the captured image of the sampled training view, so the element on which claims 10 and 11 turn is the sparse point cloud determined by depth estimation carried forward from claim 9, which is the input from which the claimed three-dimensional scene model is generated.
Any comments considered necessary by applicant must be submitted no later than the payment of the issue fee and, to avoid processing delays, should preferably accompany the issue fee. Such submissions should be clearly labeled “Comments on Statement of Reasons for Allowance.”
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jai Li whose telephone number is (571)272-1170. The examiner can normally be reached Mon-Thu between 06:00-16:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571)272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAI W LI/Junior Patent Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613