DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Response to Amendment
2. Applicant’s amendments filed on 06/24/2026 have been entered. Claims 1, 4, 5, 6, 13, 14, 17 and 18 have been amended. Claims 3 and 16 have been canceled. Claims 1-2 and 4-15 and 17-20 are pending in this application, with claims 1, 13 and 14 being independent.
Response to Arguments
3. Applicant’s arguments, see page 9, filed 06/24/2026, with respect to the claim objections have been fully considered and are persuasive. The amendments to the claims are sufficient to overcome the informalities of the previous claims; thus the objections to these claims have been withdrawn.
4. Applicant's arguments filed on 06/24/2026, with respect to the 103 rejection have been fully considered but are moot in view of the new grounds of rejection.
Examiner notes that independent claims 1 and 13-14 have been amended to include new limitation. Examiner finds these limitations to be unpatentable as can be found in below detail action.
In light of the current Office Action, the Examiner respectfully submits that independent claims 1 and 13-14 are rejected in view of newly discovered reference(s) to Jiang et al. (US-2019/0138889-A1).
Examiner notes that independent claims 1 and 13-14 have been amended to include new limitation. Examiner finds these limitations to be unpatentable as can be found in above detail action.
On page 14 of Applicant's Remarks, the Applicant argues that the dependent claims are not taught by the prior art, insomuch as they depend from claims that are not taught by the prior art. Examiner respectfully disagrees with these arguments, for the reasons discussed below.
Claim Rejections - 35 USC § 103
5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
6. Claims 1-2, 4-6, 13-15 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al., (“Chen”) [US-2009/0110239-A1] in view of Kosiorek et al. (“Kosiorek”) [US-2024/0070972-A1], further in view of Jiang et al. (“Jiang”) [US-2019/0138889-A1]
Regarding claim 1, Chen discloses a rendering method (Chen- Abstract, at least discloses system and method for identifying objects in an image dataset that occlude other objects and for transforming the image dataset to reveal the occluded objects) comprising:
obtaining a target image corresponding to a target view by inputting first parameter information corresponding to the target view (Chen- Fig. 1A shows a photograph of a street locality showing a shop sign partially occluded by a street sign [a target image] with the first viewpoint; ¶0020, at least discloses a photograph of street scene with buildings adjoining the street. In the view of FIG. 1A, a street sign 100 partially blocks the view of a sign 102 on one of the shops next to the street [obtaining a target image corresponding to a target view]; Fig 2A and ¶0014, at least disclose an overhead schematic of a vehicle collecting images in a street locality; in this view, a doorway is partially occluded by a signpost [target view]; Fig 2A and ¶0026, at least disclose 206 shows the point-of-view [first parameter information] of just one camera in the image-capture system 204. As the vehicle 200 passes the building 208, this camera (and others, not shown) image various areas of the building 208. Located in front of the building 208 are “occluding objects,” here represented by signposts 210. These are called “occluding objects” because they hide (or “occlude”) whatever is behind them. In the example of FIG. 2A, from the illustrated point-of-view 206 [first parameter information], the doorway 212 of the building 208 [target view] is partially occluded by one of the signposts 210);
obtaining an adjacent view that satisfies a predetermined condition with respect to the target view (Chen- Fig. 1B shows a photograph of the same locality as shown FIG. 1A but taken from a different viewpoint [an adjacent view] where the street sign does not occlude the same portion of the shop sign as occluded in FIG. 1A [a target image]; Fig 1B and ¶0021, at least disclose FIG. 1B is another photograph [an adjacent view], taken from a slightly different point of view from that of FIG. 1A [a predetermined condition with respect to the target view]. Comparing FIGS. 1A and 1B, the foreground street sign 100 has “moved” relative to the background shop sign 102. Because of this parallax “movement,” the portion of the shop sign 102 that is blocked in FIG. 1A is now clearly visible in FIG. 1B. (Also, a portion of the shop sign 102 that is visible in FIG. 1A is blocked in FIG. 1B.));
obtaining an adjacent image corresponding to the adjacent view by inputting second parameter information (Chen- Fig. 1B shows a photograph of the same locality as shown FIG. 1A but taken from a different viewpoint [an adjacent view] where the street sign does not occlude the same portion of the shop sign as occluded in FIG. 1A [a target image]; Fig 1B and ¶0021, at least disclose FIG. 1B is another photograph [an adjacent view], taken from a slightly different point of view [second parameter information] from that of FIG. 1A. Comparing FIGS. 1A and 1B, the foreground street sign 100 has “moved” relative to the background shop sign 102. Because of this parallax “movement,” the portion of the shop sign 102 that is blocked in FIG. 1A is now clearly visible in FIG. 1B. (Also, a portion of the shop sign 102 that is visible in FIG. 1A is blocked in FIG. 1B.); ); and
obtaining a final image by correcting the target image based on the adjacent image (Chen- Fig. 1C shows a photograph showing the same view as in FIG. 1A but post-processed to remove the portion of the street sign that occludes the shop sign and replace the vacant space with the missing portion of the shop sign; Fig. 1C and ¶0022-0023, at least disclose the viewpoint is the same as in FIG. 1A, but the processed image of FIG. 1C shows the entire shop sign 102. When the resulting image dataset is viewed, the processed image of FIG. 1C effectively replaces the image of FIG. 1A, thus rendering the shop sign 102 completely visible […] only the portion of the street sign 100 that blocks the shop sign 102 is removed: Most of the street sign 100 is left in place. This is meant to clearly illustrate the removal of the blocking portion of the street sign 100 […] the blocking portion of the street sign 100 is left in place but is visually “de-emphasized” (e.g., rendered semi-transparent) so that the shop sign 102 is revealed behind it [Wingdings font/0xE0] suggests the partial portion of shop sign 102 being blocked by street sign in Fig. 1A in the target image (the target image) being corrected by removing of the blocking portion of the street sign 100 which blocked the right portion of the shop sign 102 shown in Fig. 1B (the adjacent image)).
Chen does not explicitly disclose a neural scene representation (NSR) model; inputting second parameter information corresponding to the adjacent view to the NSR model; wherein the obtaining the final image comprises: obtaining a visibility map based on the target image and the adjacent image, wherein the obtaining the visibility map comprises obtaining a first warped image by warping the adjacent image to the target view, and obtaining the visibility map based on a difference between the first warped image and the target image; and correcting the target image based on the visibility map.
However, Kosiorek discloses
a neural scene representation (NSR) model (Kosiorek- Fig. 2 and ¶0070, at least disclose FIG. 2 illustrates an example of volumetric rendering of an image 215 of a scene 250 using a radiance field. The volumetric rendering can be used by an image rendering system (e.g., the system 100 in FIG. 1 ) to render an image 215 depicting the scene 250 from a perspective of a camera at a new camera location 216. In particular, the image rendering system can use a scene representation neural network 240 [a neural scene representation (NSR) model], defining a geometric model of the scene as a three-dimensional radiance field 260, to render the image 215).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen to incorporate the teachings of Kosiorek, and apply the geometric model of the scene as a three-dimensional radiance field into Chen’s teachings for obtaining a target image corresponding to a target view based on by inputting first parameter information corresponding to the target view to a neural scene representation (NSR) model; obtaining an adjacent view that satisfies a predetermined condition with respect to the target view; obtaining an adjacent image corresponding to the adjacent view by inputting second parameter information corresponding to the adjacent view to the NSR model; and obtaining a final image by correcting the target image based on the adjacent image.
Doing so would reduce the amount of computation and training needed as the scene representation neural network does not also have to learn how to render an image.
Kosiorek further discloses
obtaining a target image corresponding to a target view based on by inputting first parameter information corresponding to the target view to a neural scene representation (NSR) model (Kosiorek- Fig. 2 and ¶0070, at least disclose volumetric rendering of an image 215 of a scene 250 [obtaining a target image corresponding to a target view] using a radiance field. The volumetric rendering can be used by an image rendering system (e.g., the system 100 in FIG. 1 ) to render an image 215 depicting the scene 250 from a perspective of a camera at a new camera location 216 [inputting first parameter information]. In particular, the image rendering system can use a scene representation neural network 240 [a neural scene representation model], defining a geometric model of the scene as a three-dimensional radiance field 260, to render the image 215);
obtaining an adjacent image corresponding to the adjacent view by inputting second parameter information corresponding to the adjacent view to the NSR model (Kosiorek- Fig. 2 and ¶0077, at least disclose the system can use the neural network 240 conditioned on the latent variable representing the scene 250 to render another new image 220 of the scene 250 [adjacent image corresponding to the adjacent view] from the perspective of the camera at a completely different camera location 225 (e.g., illustrated as being perpendicular to the camera location 216) [second parameter information]).
The prior art does not explicitly disclose, but Jiang discloses
wherein the obtaining the final image (Jiang- Fig. 1D and ¶0042, at least disclose The high-quality variable-length multi-frame interpolation technique performed by the frame interpolation system 100 predicts a frame at any arbitrary time step between two frames by warping the input two images to the specific time step and then adaptively fusing the two warped images to generate the intermediate image […] When the two intermediate bi-directional optical flows ({circumflex over (F)}t→0, {circumflex over (F)}t→1) are known, the intermediate image It may be synthesized) comprises:
obtaining a visibility map based on the target image and the adjacent image (Jiang- ¶0006, at least discloses A second neural network model refines the optical flow data and predicts visibility maps for each timestep; ¶0046, at least discloses To account for occlusion, the refinement neural network models 145-0 and 145-1 predict soft visibility maps Vt←0 and Vt←1, respectively, for each timestep. Vt←0(p)∈[0,1] denotes whether the pixel p remains visible (0 is fully occluded and therefore, not visible) when moving from T=0 to T=t. The two visibility maps are forced to satisfy the following constraint V t←0=1−V t←1; Fig. 1F and ¶0055, at least disclose At step 185, the visibility maps are predicted by the flow interpolation neural network model 122 based on the inputs to the flow interpolation neural network model 122 […] the flow refinement neural network model 145-0 predicts the visibility map Vt→0 based on the first input frame I0, the intermediate forward optical flow data {circumflex over (F)}t→0, and the first frame warped according to the intermediate forward optical flow data (Î0→t). In an embodiment, the flow refinement neural network model 145-1 predicts the visibility map Vt→1 based on the second input frame I1, the intermediate backward optical flow data {circumflex over (F)}t→1, and the second frame warped according to the intermediate backward optical flow data (Î1→t)), wherein the obtaining the visibility map comprises obtaining a first warped image by warping the adjacent image to the target view (Jiang- ¶0006-0007, at least disclose A second neural network model refines the optical flow data and predicts visibility maps for each timestep […] Backward optical flow data computed starting from the second frame to the first frame and intermediate backward optical flow data for the time is received. A flow interpolation neural network model generates an intermediate frame at the time based on the first frame and the second frame, the intermediate forward optical flow data, the intermediate backward optical flow data, the first frame warped according to the intermediate forward optical flow data, and the second frame warped according to the intermediate backward optical flow data; Fig. 1D and ¶0042, at least disclose The high-quality variable-length multi-frame interpolation technique performed by the frame interpolation system 100 predicts a frame at any arbitrary time step between two frames by warping the input two images to the specific time step and then adaptively fusing the two warped images to generate the intermediate image [obtaining a first warped image]. In an embodiment, instead of performing a forward warping operation, a backward warping operation is applied to the first and second input frames. For the backward warp, each pixel in the intermediate (target) image finds a correspondence in the input (source) image; Fig. 1F and ¶0055, at least disclose At step 185, the visibility maps are predicted by the flow interpolation neural network model 122 based on the inputs to the flow interpolation neural network model 122 […] the flow refinement neural network model 145-0 predicts the visibility map Vt→0 based on the first input frame I0, the intermediate forward optical flow data {circumflex over (F)}t→0, and the first frame warped according to the intermediate forward optical flow data (Î0→t). In an embodiment, the flow refinement neural network model 145-1 predicts the visibility map Vt→1 based on the second input frame I1, the intermediate backward optical flow data {circumflex over (F)}t→1, and the second frame warped according to the intermediate backward optical flow data (Î1→t)), and obtaining the visibility map based on a difference between the first warped image and the target image (Jiang- Fig. 2B and ¶0068-0069, at least disclose the training loss function 210 compares the intermediate frame with a ground truth frame. At step 230, based on differences between the predicted frame and the ground truth frame, training may be complete. A loss function, such as the loss function defined by Equation (7), may be computed to measure distances (i.e., differences or gradients) between the ground truth frame and the predicted intermediate frame […] at step 235, the parameters of the intermediate optical flow neural network model 142 and the flow interpolation neural network model 102 or 122 are adjusted to reduce differences between the intermediate frame and the ground truth frame); and
correcting the target image based on the visibility map (Jiang- ¶0006, at least discloses Artifacts caused by motion boundaries and occlusions are reduced in the predicted intermediate frames; ¶0025, at least discloses A neural network model may be trained to interpolate one or more in-between (intermediate) frames for a video sequence […] The intermediate frames that are generated by the neural network model are both spatially and temporally coherent within the video sequence. A successful solution should not only correctly interpret the motion between two input images (implicitly or explicitly), but also understand occlusions to reduce artifacts in the interpolated frames, especially around motion boundaries; ¶0050, at least discloses By applying the visibility maps to the warped images before fusion, the contribution of occluded pixels is excluded from the interpolated intermediate frame, thereby avoiding or reducing artifacts).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen/Kosiorek to incorporate the teachings of Jiang, and apply generating the intermediate image and the visibility maps are predicted into Chen/Kosiorek’s teachings for obtaining a visibility map based on the target image and the adjacent image, wherein the obtaining the visibility map comprises obtaining a first warped image by warping the adjacent image to the target view, and obtaining the visibility map based on a difference between the first warped image and the target image; and correcting the target image based on the visibility map.
Doing so artifacts caused by motion boundaries and occlusions are reduced in the predicted intermediate frames.
Regarding claim 2, Chen in view of Kosiorek and Jiang, discloses the rendering method of claim 1, and further discloses wherein the obtaining the final image (see Claim 1 rejection for detailed analysis) comprises:
detecting an occlusion area in the target image (Chen- Fig. 1A shows a photograph of a street locality showing a shop sign partially occluded by a street sign [a target image] with the first viewpoint; ¶0020, at least discloses a photograph of street scene with buildings adjoining the street. In the view of FIG. 1A, a street sign 100 partially blocks the view of a sign 102 on one of the shops next to the street [obtaining a target image corresponding to a target view]; Fig 2A and ¶0014, at least disclose an overhead schematic of a vehicle collecting images in a street locality; in this view, a doorway is partially occluded by a signpost [target view]; Fig 2A and ¶0026, at least disclose 206 shows the point-of-view [first parameter information] of just one camera in the image-capture system 204. As the vehicle 200 passes the building 208, this camera (and others, not shown) image various areas of the building 208. Located in front of the building 208 are “occluding objects,” here represented by signposts 210. These are called “occluding objects” because they hide (or “occlude”) whatever is behind them. In the example of FIG. 2A, from the illustrated point-of-view 206 [first parameter information], the doorway 212 of the building 208 [target view] is partially occluded by one of the signposts 210); Jiang- ¶0070, at least discloses In the input frames the arms of the football player move downwards from T=0 to T=1. The area directly above the arm at T=0 is visible at t, but the same area is occluded (i.e., invisible) at T=1); and
correcting the occlusion area in the target image based on the adjacent image (Chen- Fig. 1C shows a photograph showing the same view as in FIG. 1A but post-processed to remove the portion of the street sign that occludes the shop sign and replace the vacant space with the missing portion of the shop sign; Fig. 1C and ¶0022-0023, at least disclose the viewpoint is the same as in FIG. 1A, but the processed image of FIG. 1C shows the entire shop sign 102. When the resulting image dataset is viewed, the processed image of FIG. 1C effectively replaces the image of FIG. 1A, thus rendering the shop sign 102 completely visible […] only the portion of the street sign 100 that blocks the shop sign 102 is removed: Most of the street sign 100 is left in place. This is meant to clearly illustrate the removal of the blocking portion of the street sign 100 […] the blocking portion of the street sign 100 is left in place but is visually “de-emphasized” (e.g., rendered semi-transparent) so that the shop sign 102 is revealed behind it [Wingdings font/0xE0] suggests the partial portion of shop sign 102 being blocked by street sign in Fig. 1A in the target image (the target image) being corrected by removing of the blocking portion of the street sign 100 which blocked the right portion of the shop sign 102 shown in Fig. 1B (the adjacent image)).
Regarding claim 4, Chen in view of Kosiorek and Jiang, discloses the rendering method of claim 1, and further discloses wherein the warping the adjacent image (see Claim 1 rejection for detailed analysis) comprises
backward-warping the adjacent image to the target view (Jiang- Fig. 1D and ¶0042, at least disclose instead of performing a forward warping operation, a backward warping operation is applied to the first and second input frames. For the backward warp, each pixel in the intermediate (target) image finds a correspondence in the input (source) image; Fig. 1F and ¶0055, at least disclose At step 185, the visibility maps are predicted by the flow interpolation neural network model 122 based on the inputs to the flow interpolation neural network model 122 […] the flow refinement neural network model 145-0 predicts the visibility map Vt→0 based on the first input frame I0, the intermediate forward optical flow data {circumflex over (F)}t→0, and the first frame warped according to the intermediate forward optical flow data (Î0→t). In an embodiment, the flow refinement neural network model 145-1 predicts the visibility map Vt→1 based on the second input frame I1, the intermediate backward optical flow data {circumflex over (F)}t→1, and the second frame warped according to the intermediate backward optical flow data (Î1→t)).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen/Kosiorek to incorporate the teachings of Jiang, and apply backward warping operation into Chen/Kosiorek’s teachings in order the warping the adjacent image comprises backward-warping the adjacent image to the target view.
The same motivation that was utilized in the rejection of claim 1 applies equally to this claim.
Regarding claim 5, Chen in view of Kosiorek and Jiang, discloses the the rendering method of claim 1, and further discloses wherein the obtaining the visibility map based on the difference (see Claim 1 rejection for detailed analysis) comprises:
obtaining a visibility value for a first pixel of the first warped image based on a difference between the first pixel of the first warped image and a second pixel of the target image corresponding to the first pixel (Jiang- ¶0039, at least discloses The intermediate forward optical flow data are for a time T=t in the sequence of frames that is between the first frame and the second frame. In an embodiment, color data for each frame may be represented as red, green, and blue (RGB) color components, YUV components, or the like […] a specific color of a particular pixel or edge of an object in the first frame may move to a different pixel in the second frame. The color may be associated with an object in a scene that moves between the first and second frames; Fig. 2D and ¶0071, at least disclose visibility maps and predicted intermediate frames, in accordance with an embodiment. Areas that are not visible are indicated by black in the visibility maps and areas that are visible are indicated by white. The visibility maps in the top row clearly show that the area directly above the arm at T=0 is visible and is occluded at T=1. The white area around the arms in Vt←1 indicates pixels in I1 that contribute most to the predicted It while the corresponding occluded pixels (in black areas) in I0 have little contribution).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen/Kosiorek to incorporate the teachings of Jiang, and apply different pixel in the second frame into Chen/Kosiorek’s teachings for obtaining a visibility value for a first pixel of the first warped image based on a difference between the first pixel of the first warped image and a second pixel of the target image corresponding to the first pixel.
The same motivation that was utilized in the rejection of claim 1 applies equally to this claim.
Regarding claim 6, Chen in view of Kosiorek and Jiang, discloses the rendering method of claim 1, and further discloses wherein the correcting the target image (see Claim 1 rejection for detailed analysis) comprises:
detecting an occlusion area in the target image based on the visibility map (Chen- Fig. 1A shows a photograph of a street locality showing a shop sign partially occluded by a street sign [a target image] with the first viewpoint; ¶0020, at least discloses a photograph of street scene with buildings adjoining the street. In the view of FIG. 1A, a street sign 100 partially blocks the view of a sign 102 on one of the shops next to the street [detecting an occlusion area in the target image]; Fig 2A and ¶0014, at least disclose an overhead schematic of a vehicle collecting images in a street locality; in this view, a doorway is partially occluded by a signpost [target view]; Jiang - Fig 2D and ¶0071, at least disclose visibility maps and predicted intermediate frames, in accordance with an embodiment. Areas that are not visible are indicated by black in the visibility maps and areas that are visible are indicated by white. The visibility maps in the top row clearly show that the area directly above the arm at T=0 is visible and is occluded at T=1. The white area around the arms in Vt←1 indicates pixels in I1 that contribute most to the predicted It while the corresponding occluded pixels (in black areas) in I0 have little contribution); and
correcting the occlusion area in the target image based on the adjacent image (Chen- Fig. 1C shows a photograph showing the same view as in FIG. 1A but post-processed to remove the portion of the street sign that occludes the shop sign and replace the vacant space with the missing portion of the shop sign; Fig. 1C and ¶0022-0023, at least disclose the viewpoint is the same as in FIG. 1A, but the processed image of FIG. 1C shows the entire shop sign 102. When the resulting image dataset is viewed, the processed image of FIG. 1C effectively replaces the image of FIG. 1A, thus rendering the shop sign 102 completely visible […] only the portion of the street sign 100 that blocks the shop sign 102 is removed: Most of the street sign 100 is left in place. This is meant to clearly illustrate the removal of the blocking portion of the street sign 100 […] the blocking portion of the street sign 100 is left in place but is visually “de-emphasized” (e.g., rendered semi-transparent) so that the shop sign 102 is revealed behind it [Wingdings font/0xE0] suggests the partial portion of shop sign 102 being blocked by street sign in Fig. 1A in the target image (the target image) being corrected by removing of the blocking portion of the street sign 100 which blocked the right portion of the shop sign 102 shown in Fig. 1B (the adjacent image)).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen/Kosiorek to incorporate the teachings of Jiang, and apply the occluded pixels into Chen/Kosiorek’s teachings for detecting an occlusion area in the target image based on the visibility map; and correcting the occlusion area in the target image based on the adjacent image.
The same motivation that was utilized in the rejection of claim 1 applies equally to this claim.
Regarding claim 13, Chen in view of Kosiorek and Jiang, discloses a non-transitory computer-readable storage medium storing instructions (Chen- Fig. 3 and ¶0028-0029, at least disclose When the image dataset 300 is transformed or processed in some way, the resulting product is stored on the same computer-readable medium or on another one. The transformed image dataset 300, in whole or in part, may be distributed to users on a tangible medium, such as a CD, or may be transmitted over a network such as the Internet or over a wireless link to a user's mobile navigation system […] Many different applications can use this same image dataset 300. As one example, an image-based navigation application 302 allows a user to virtually walk down the street 202; Claim 17 at least cites “A computer-readable medium containing computer-executable instructions for a method for transforming an image dataset,”) that, when executed by a processor(Chen- Fig. 5 and ¶0040, at least disclose platform 500 runs one or more applications 502 including, for example, the image-based navigation application 302 as discussed above. For this application 302, the user navigates using tools provided by an interface 506 (such as a keyboard, mouse, microphone, voice recognition software, and the like); Kosiorek- ¶0135, at least discloses The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers […] he apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them) , cause the processor to perform a method comprising:
obtaining a target image corresponding to a target view by inputting first parameter information corresponding to the target view to a neural scene representation (NSR) model (see Claim 1 rejection for detailed analysis);
obtaining an adjacent view corresponding to the target view (Chen- Fig. 1B shows a photograph of the same locality as shown FIG. 1A but taken from a different viewpoint [an adjacent view] where the street sign does not occlude the same portion of the shop sign as occluded in FIG. 1A [target view]; Fig 1B and ¶0021, at least disclose FIG. 1B is another photograph [an adjacent view], taken from a slightly different point of view from that of FIG. 1A [the target view]);
obtaining an adjacent image corresponding to the adjacent view by inputting second parameter information corresponding to the adjacent view to the NSR model (see Claim 1 rejection for detailed analysis); and
obtaining a final image by correcting the target image based on the adjacent image (see Claim 1 rejection for detailed analysis), wherein the obtaining the final image comprises:
obtaining a visibility map based on the target image and the adjacent image (see Claim 1 rejection for detailed analysis), wherein the obtaining the visibility map comprises obtaining a first warped image by warping the adjacent image to the target view (see Claim 1 rejection for detailed analysis), and obtaining the visibility map based on a difference between the first warped image and the target image (see Claim 1 rejection for detailed analysis); and
correcting the target image based on the visibility map (see Claim 1 rejection for detailed analysis).
The rendering device of claims 14-15, 17-18 are similar in scope to the functions performed by the method of claims 1-2, 4, 6 and therefore claims 14-15, 17-18 are rejected under the same rationale.
Regarding claim 14, Chen in view of Kosiorek and Jiang, discloses a rendering device (Chen- Fig. 3 and ¶0028, at least disclose the image-capture system 204 delivers its images to an image dataset 300) comprising:
a memory configured to store instructions (Chen- Fig. 3 and ¶0028-0029, at least disclose When the image dataset 300 is transformed or processed in some way, the resulting product is stored on the same computer-readable medium or on another one. The transformed image dataset 300, in whole or in part, may be distributed to users on a tangible medium, such as a CD, or may be transmitted over a network such as the Internet or over a wireless link to a user's mobile navigation system […] Many different applications can use this same image dataset 300. As one example, an image-based navigation application 302 allows a user to virtually walk down the street 202; Claim 17 at least cites “A computer-readable medium containing computer-executable instructions for a method for transforming an image dataset,”),
at least one processor configured to execute the instructions (Chen- Fig. 5 and ¶0040, at least disclose platform 500 runs one or more applications 502 including, for example, the image-based navigation application 302 as discussed above. For this application 302, the user navigates using tools provided by an interface 506 (such as a keyboard, mouse, microphone, voice recognition software, and the like); Kosiorek- ¶0135, at least discloses The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers […] he apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them) to perform the method of claim 1.
7. Claims 7-9 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Kosiorek, further in view of Jiang, still further in view of Ciurea et al. (“Ciurea”) [US-2015/0049917-A1]
Regarding claim 7, Chen in view of Kosiorek and Jiang, discloses the rendering method of claim 6, and further discloses wherein the detecting the occlusion area (see Claim 6 rejection for detailed analysis), and the prior art does not explicitly disclose, but Ciurea discloses the rendering method comprises:
detecting an occluded pixel having a visibility value that is greater than or equal to a threshold value in the target image (Ciurea- ¶0249, at least discloses the photometric distance of the pixels is utilized as a measure of similarity and a threshold used to determine pixels that are likely visible and pixels that are likely occluded; ¶0250, at least discloses When the photometric distance of the selected pixel from the reference image and one of the corresponding pixels is less than the threshold, then the corresponding pixel is determined (1112) to be visible. When the photometric distance of the selected pixel from the reference image and one of the corresponding pixels exceeds the threshold, then the corresponding pixel is determined (1114) to be occluded).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen/Kosiorek/Jiang to incorporate the teachings of Ciurea, and apply the pixels exceeds the threshold into Chen/Kosiorek/Jiang’s teachings for detecting an occluded pixel having a visibility value that is greater than or equal to a threshold value in the target image.
Doing so would provide an accurate account of the pixel disparity as a result of parallax between the different cameras in the array, so that appropriate scene-dependent geometric shifts can be applied to the pixels of the captured images when performing super-resolution processing.
Regarding claim 8, Chen in view of Kosiorek, Jiang and Ciurea, discloses the rendering method of claim 7, and further discloses wherein the correcting the occlusion area (see Claim 2 rejection for detailed analysis) comprises:
replacing the occluded pixel in the target image with a pixel of the adjacent image corresponding to a position of the occluded pixel (Chen- Claim 13, at least cites “replacing 3D pixels occluded by the first object in a first image with 3D pixels corresponding to a same location in a second image and not occluded by the first object”).
Regarding claim 9, Chen in view of Kosiorek, Jiang and Ciurea, discloses the rendering method of claim 7, and further discloses wherein the obtaining the adjacent view (see Claim 2 rejection for detailed analysis) comprises:
obtaining a plurality of adjacent views corresponding to the target view (Chen- Fig. 1B shows a photograph of the same locality as shown FIG. 1A but taken from a different viewpoint [adjacent view] where the street sign does not occlude the same portion of the shop sign as occluded in FIG. 1A [a target image]; Fig 1B and ¶0021, at least disclose FIG. 1B is another photograph [an adjacent view], taken from a slightly different point of view from that of FIG. 1A. Comparing FIGS. 1A and 1B, the foreground street sign 100 has “moved” relative to the background shop sign 102. Because of this parallax “movement,” the portion of the shop sign 102 that is blocked in FIG. 1A is now clearly visible in FIG. 1B. (Also, a portion of the shop sign 102 that is visible in FIG. 1A is blocked in FIG. 1B); Fig. 2 shows an overhead schematic of a vehicle collecting images in a street locality), wherein the obtaining of the adjacent image comprises:
obtaining a plurality of adjacent images, each of the plurality of adjacent images corresponding to one of the plurality of adjacent views (Chen- Fig. 1B shows a photograph of the same locality as shown FIG. 1A but taken from a different viewpoint [adjacent view] where the street sign does not occlude the same portion of the shop sign as occluded in FIG. 1A; Fig 2A and ¶0024, at least disclose The system in these figures collects images of a locality in the real world. In FIG. 2A, a vehicle 200 is driving down a road 202 […] As the vehicle 200 proceeds down the road 202, the captured images are stored and are associated with the geographical location at which each picture was taken), and wherein the correcting of the occlusion area (see Claim 2 rejection for detailed analysis) comprises:
obtaining a pixel of each of the plurality of adjacent images corresponding to a position of the occluded pixel in the target image (Chen- Fig 1B and ¶0021, at least disclose FIG. 1B is another photograph [an adjacent view], taken from a slightly different point of view from that of FIG. 1A. Comparing FIGS. 1A and 1B, the foreground street sign 100 has “moved” relative to the background shop sign 102. Because of this parallax “movement,” the portion of the shop sign 102 that is blocked in FIG. 1A is now clearly visible in FIG. 1B. (Also, a portion of the shop sign 102 that is visible in FIG. 1A is blocked in FIG. 1B); ¶0036, at least discloses Turning back to FIG. 2A, the data representing the signposts 210 are deleted, but that leaves a visual “hole” in the doorway 212 as seen from the point-of-view 206. Several techniques can be applied to fill this visual hole. In FIG. 2B, the point-of-view 214 includes the entire doorway 212 with no occlusion. The pixels taken in FIG. 2B can be used to fill in the visual hole; Claim 13, at least cites “replacing 3D pixels occluded by the first object in a first image with 3D pixels corresponding to a same location in a second image and not occluded by the first object”); and
correcting the occluded pixel in the target image based on the pixel of each of the plurality of adjacent images (Chen- Claim 13, at least cites “replacing 3D pixels occluded by the first object in a first image with 3D pixels corresponding to a same location in a second image and not occluded by the first object”).
The rendering device of claims 19-20 are similar in scope to the functions performed by the method of claims 7-8 and therefore claims 19-20 are rejected under the same rationale.
8. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Kosiorek, further in view of Jiang, still further in view of “A survey on image-based rendering—representation, sampling and compression” by Cha Zhang, Tsuhan Chen (“Zhang”)
Regarding claim 10, Chen in view of Kosiorek and Jiang, discloses the rendering method of claim 1, and does not explicitly disclose, but Zhang discloses wherein the obtaining the adjacent view comprises:
obtaining the adjacent view by sampling views within a preset camera rotation angle based on the target view (Zhang- Fig. 4 and page 7, section 2.2.4. 3D—concentric mosaics and panoramic video, 1st and 2nd paragraphs, at least disclose In concentric mosaics, the scene is captured by mounting a camera at the end of a level beam, and shooting images at regular intervals as the beam rotates, as is shown in Fig. 4. The light rays are then indexed by the camera position or the beam rotation angle a; and the pixel locations ðu; vÞ […] This parameterization is equivalent to having many slit cameras rotating around a common center and taking images along the tangent direction […] During the rendering, the viewer may move freely inside a rendering circle (Fig. 4) with radius R sinðFOV=2Þ; where R is the camera path radius and FOV is the field of view of the cameras).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen/Kosiorek/Jiang to incorporate the teachings of Zhang, and apply the beam rotation angle into Chen/Kosiorek/Jiang’s teachings for obtaining the adjacent view by sampling views within a preset camera rotation angle based on the target view.
Doing so would reproduce the scene correctly at an arbitrary viewpoint, with unknown or limited amount of geometry.
9. Claims 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Chen in view of Kosiorek, further in view of Jiang, still further in view of Rong et al. (“Rong”) [US-2021/0383616-A1]
Regarding claim 11, Chen in view of Kosiorek and Jiang, discloses the rendering method of claim 1, and further discloses wherein the obtaining the target image comprises:
obtaining a rendered image corresponding to the target view (Chen- Fig. 1A shows a photograph of a street locality showing a shop sign partially occluded by a street sign with the first viewpoint [a target image]; ¶0020, at least discloses a photograph of street scene with buildings adjoining the street. In the view of FIG. 1A, a street sign 100 partially blocks the view of a sign 102 on one of the shops next to the street [obtaining a rendered image corresponding to the target view]) .
The prior art does not explicitly disclose, but Rong discloses
a depth map corresponding to the target view (Rong- ¶0047, at least discloses view warping may begin by rendering the selected object data's 3D mesh model at selected target viewpoint to generate the corresponding target depth. The rendered depth map of the object data set along with the source camera images may be used to generate the object's 2D texture map using an inverse warping operation).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen/Kosiorek/Jiang to incorporate the teachings of Rong, and apply the depth map into Chen/Kosiore/Jiangk’s teachings for obtaining a rendered image corresponding to the target view and a depth map corresponding to the target view.
Doing so would provide augmented data that used to test safety features of software for a self-driving vehicle.
Regarding claim 12, Chen in view of Kosiorek and Jiang, discloses the rendering method of claim 1, and further discloses wherein the obtaining the adjacent image comprises:
obtaining a rendered image corresponding to the adjacent view (Chen- Fig. 1B shows a photograph of the same locality as shown FIG. 1A but taken from a different viewpoint [an adjacent view] where the street sign does not occlude the same portion of the shop sign as occluded in FIG. 1A [a target image]; Fig 1B and ¶0021, at least disclose FIG. 1B is another photograph [an adjacent view], taken from a slightly different point of view from that of FIG. 1A). .
The prior art does not explicitly disclose, but Rong discloses
a depth map corresponding to the view (Rong- ¶0049, at least discloses The interpolation may be used to obtain the estimated depths to generate an estimated depth map of the image. The object data set may be processed to render the depth of the object).
It would have been obvious to one of ordinary in the art before the effective filing date of the claimed invention to have modified Chen/Kosiorek/Jiang to incorporate the teachings of Rong, and apply the depth map into Chen/Kosiorek/Jiang’s teachings for obtaining a rendered image corresponding to the adjacent view and a depth map corresponding to the adjacent view.
Doing so would provide augmented data that used to test safety features of software for a self-driving vehicle.
Conclusion
10. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
11. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL LE whose telephone number is (571)272-5330. The examiner can normally be reached 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571) 272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL LE/Primary Examiner, Art Unit 2614